Skip to content

tour · 3 workflows · 15 screenshots

A guided tour

Short recordings of the three main workflows, then a screenshot of every key feature. Each step in a list below jumps to that moment in its video, and the list doubles as the transcript.

The third recording shows the bring-your-own-key settings with a placeholder key, and an LLM run answered by a mocked provider inside the browser. It is labelled “Mocked AI response for illustration” on screen, its model id says so, and its numbers are not results for any real model.

// workflow 1 · 0:48 · light theme

Clean a message

Pick a sample input and step through the five stages of the original program, with the animated token diff and the exact stdout.

steps · transcript

MP4, H.264, 1280 × 800, 1.2 MB · captions as WebVTT · open the video file

// workflow 2 · 0:37 · dark theme

Under the hood

Open the original C behind the current stage, then the differential and property-based testing that checks the port against the compiled program.

steps · transcript

MP4, H.264, 1280 × 800, 1.2 MB · captions as WebVTT · open the video file

// workflow 3 · 1:05 · light theme

Rules vs LLM

Run the no-key baseline, open the bring-your-own-key settings, and walk through a clearly labelled mocked LLM run, the paired comparison and the AI audit log.

steps · transcript

MP4, H.264, 1280 × 800, 2.1 MB · captions as WebVTT · open the video file

// key features

Screenshots

Desktop captures at 1440 × 900 and phone captures at 390 × 844. Select one for a larger view; the arrow keys move between them.

On a phone

How these were made

Everything on this page comes from one Playwright script, web/e2e/showcase.spec.ts, run with pnpm showcase in Google Chrome against the deployed site or a local production build. It uses the built-in samples and the default evaluation seed (2020, N = 20), so a rerun shows the same inputs and the same baseline scores. Every step also asserts what should be on screen, so the tour doubles as an end-to-end test.

No real API key is used: requests to the AI providers are blocked during the tour, and the one LLM run is answered by a stand-in that applies the published procedure with two deliberate error modes, so the scoring and review screens have something to show. The captions, the cursor highlight and the mocked banner are drawn by the script, not by the site.

The methods page has the testing and evaluation design, and the README has the same walkthrough as GIFs.