Skip to content

methods · evidence · decisions

How I know the port is faithful

This page sets out where every input comes from, how the TypeScript port is tested against the compiled 2020 C program, how the optional Rules vs LLM evaluation is designed, and the decisions behind both. Every number below is computed from the repository when the site is built, with its sample size and a 95% confidence interval.

481/481seeded inputs where the port prints exactly what the C binary printed95% CI 99.2% to 100% (Wilson)
19/19inputs that crash the C binary, flagged by the port instead95% CI 83.2% to 100% (Wilson)
9tokenisation invariants, each checked on 300 generated inputsfast-check, seed 10002

// data provenance

Where every input comes from

No real messages or personal data are used anywhere on this site. The inputs fall into five groups, and only the first came from the University.

SourceInputsProvenance
coursework/tests2The official sample inputs and expected outputs distributed with the assignment in 2020. Kept in the repository for testing and not served on the site.
c-reference.json41Edge cases I chose while porting (line endings, limits, ties, NUL bytes, Latin-1), with the stdout the compiled program printed for each.
differential-corpus.json500Generated by scripts/generate_differential_corpus.py with seed 10002, then run through the compiled program. Anyone can regenerate the same inputs from the seed.
samples.json6Example inputs I wrote for the visualiser.
Evaluation items20 (up to 100)Synthetic messages generated in the browser from a displayed seed (default 2020), with a fixed 10-line dictionary. These are the only inputs ever sent to an AI provider.
Reference outputs were recorded by compiling coursework/src/program.c with Apple clang version 21.0.0 (clang-2100.3.34.2) on macOS arm64, using the 2020 flags (-Wall -std=c99).

// method

One reference, three kinds of test

The compiled C program is the reference for everything. The TypeScript port is checked against it, and once it passes, the port becomes the reference for the LLM evaluation.

  1. step 1

    Record the original

    Compile program.c with the 2020 flags, run it on every input, and store its exact stdout and exit status as fixtures.

  2. step 2

    Differential tests

    Replay every recorded input through the port as raw bytes and require byte-for-byte identical output (DR-002).

  3. step 3

    Property-based tests

    Generate thousands of further inputs and check invariants that must hold for any input, not just recorded ones.

  4. step 4

    Rules vs LLM

    Score an LLM, given the same procedure in plain language, against the port on seeded synthetic messages (DR-004).

// differential testing

Parity with the compiled program

The seeded corpus has 500 inputs in 10 families. Most families sit on one of the program's limits, with inputs on both sides of it. For each input the port must print the same bytes and exit with the same status. When the C binary crashes, which happens when stage 5 overflows its 2,500-byte buffer, the port must raise its overflow warning instead.

The port matched all 481 inputs that ran to completion in C (100%, 95% CI 99.2% to 100%) and flagged all 19 that crashed. Its overflow warning fired on none of the inputs the binary survived. A perfect observed rate is still an estimate: with 481 inputs, a true mismatch rate of up to 0.8% is consistent with what was seen.

FamilyInputsMatched / completedParity (95% CI)bar axis 80% to 100%Crashes flagged
typicalChat-style messages and dictionaries drawn from a fixed vocabulary200197/19798.1% to 100%3/343.9% to 100%
first-line-limitFirst line of 45 to 55 bytes (the program's limit is 50)4040/4091.2% to 100%none
message-limitA later message of 275 to 285 bytes (limit 280)4040/4091.2% to 100%none
dictionary-line-limitA dictionary line of 47 to 53 bytes (limit 50)3030/3088.6% to 100%none
slot-capacity22 to 28 punctuation tokens, around stage 5's 50-slot buffer4024/2486.2% to 100%16/1680.6% to 100%
long-tokenA single token of 45 to 55 bytes (stage-5 rows are 50 bytes wide)2525/2586.7% to 100%none
message-count95 to 102 messages around the 99-read loop in stage 22525/2586.7% to 100%none
line-endingsCRLF, bare CR, blank lines and a missing final newline3030/3088.6% to 100%none
structureMissing separator, empty input or dictionary, separator variants3030/3088.6% to 100%none
raw-bytesNon-ASCII UTF-8, Latin-1, NUL, byte-order mark, whitespace and control bytes4040/4091.2% to 100%none
Parity is the share of completed runs reproduced byte for byte, with a Wilson 95% interval; crashes flagged also carry a Wilson 95% interval. Families with few runs have wide intervals.

Alongside the corpus, the port reproduces both official sample tests (2/2, 95% CI 34.2% to 100%) and all 39 completed hand-picked cases (39/39, 95% CI 91.0% to 100%), and flags the other 2 (2/2, 95% CI 34.2% to 100%). Perfect scores on so few cases are weak evidence on their own, which is why the seeded corpus exists. Re-run the generator with uv run scripts/generate_differential_corpus.py and the tests with pnpm test.

// divergences

Where the port differs on purpose, and where it does not

The port departs from the C program in three places. Each one is visible on screen when it applies.

  • Stage-5 buffer overflow

    When the tokens need more than 50 slots, the C program writes past char list[50][50] and is killed by the stack protector, losing its output. The port keeps going with a large-enough buffer and shows a warning. This covers 19 corpus inputs and 2 hand-picked ones.

  • Windows line endings in uploaded files

    The upload button converts \r\n to \n before running and says so. The engine itself still treats \r exactly as C does (DR-003).

  • Memory the C code never wrote

    Stages 1, 2 and 4 rely on stack memory being zero: count_tokens reads a fixed 50 bytes past the end of the first line, and the stage-2 and stage-4 copies are never NUL-terminated. The port assumes zeros, which matched every recorded run on the reference platform but is not guaranteed by C.

Several behaviours look like bugs but are reproduced exactly, because they are what the submitted program does. They are not divergences.

  • The first message is read with a 50-byte limit instead of 280 (try it).
  • A first message that starts with a comma reports one token too few.
  • A message whose last token is dropped in stage 5 runs into the next one (see it).
  • With an empty dictionary, stage 5 drops commas as well as emoticons.
  • At most 99 messages follow the first, and lengths are counted in bytes rather than characters.

// property-based tests

Invariants that hold for any input

Recorded cases only cover inputs someone thought of. Property-based tests generate 300 inputs per invariant (fast-check, fixed seed 10002 so failures are reproducible) and shrink any failure to a minimal counter-example. To check that the properties can fail at all, a mutation test in invariants.test.ts runs the stage-3 properties against 3 broken copies of the comma filter, each with one condition deleted, which keep a trailing comma, a leading comma and repeated commas respectively. Each must be caught within the same 300 generated inputs that the properties normally run on.

  • stage 1count_tokens returns the number of commas, plus one unless the line starts with a comma.
  • stage 2Stage 2 output is exactly the input's ispunct() bytes, in their original order.
  • stage 3The comma filter equals splitting on commas, dropping empty pieces and re-joining with single commas, so there are no leading, trailing or doubled commas.
  • stage 3Running the stage-3 comma filter twice changes nothing.
  • stage 4A dictionary emoticon is the text before the line's first comma.
  • stage 5The hand-written emotionexisitence() agrees with ordinary string equality.
  • stage 5In stage 5 a token is kept if and only if it is in the dictionary (inputs within the program's buffers), and kept tokens keep their order.
  • stage 5Stage 5 uses exactly two slots of its char list[50][50] buffer per token.
  • whole programFor any byte input the port terminates without throwing, exits with 0 or 1, and prints the stage headers in order up to where it stopped.

// LLM evaluation

Rules vs LLM: evaluation design

The evaluation asks whether a language model, given the procedure for stages 3 to 5 in plain language, can replace the rule engine. Each method receives the same seeded messages and the same 10-line dictionary, and returns two lists, the candidate tokens and the emoticons kept. The port supplies the reference answer.

Primary metric
Exact match of the kept list per message (order and duplicates included), with a Wilson 95% interval.
Secondary metrics
Token-level precision and recall of the kept list, micro-averaged, with percentile bootstrap intervals that resample whole messages (2,000 resamples, seed 2020). Exact match of the candidate list measures detection.
Paired comparison
Two runs of competing methods on the same seed and N are compared item by item with an exact McNemar test, plus paired bootstrap intervals for the differences in exact match and F1. The rule engine is the answer key, so it is never one of the two.
Failures
Rate limits, overloads and network errors are retried up to 2 times with the same policy for both providers, honouring retry-after. A call that still fails, is refused, is cut off or breaks the JSON schema counts as wrong and contributes no tokens. Exact match on the answered items is shown alongside, and paired comparisons say how many discordant items were failed calls.
Models
Claude Haiku 4.5 by default (temperature 0), Claude Sonnet 5.5 as an option (low effort, default sampling), or any OpenAI model id (default gpt-5-mini, default settings). Prompt version cleanse-v1, one call per message, and the request settings are saved with every run and audit entry.
Traps
The dictionary contains <3 and :D, which can never match because stage 2 deletes digits and letters. Messages also include emoticons glued to words, such as hi:), which the program keeps.

The two non-AI methods run without a key and anchor the scale. I have not published LLM results, because I have not spent money on API calls for this project. Visitors can run the LLM with their own key, and their results stay in their browser.

MethodExact match (95% CI)PrecisionRecallDetection
Rule engine against itselfN = 2020/2083.9% to 100%1.00 [0.91, 1.00]† Wilson, 37/37 tokens1.00 [0.91, 1.00]† Wilson, 37/37 tokens20/2083.9% to 100%
Naive split, no AIN = 2011/2034.2% to 74.2%0.91 [0.78, 1.00]0.81 [0.67, 0.92]0/200% to 16.1%
Naive split, no AIN = 10052/10042.3% to 61.5%0.89 [0.84, 0.94]0.81 [0.76, 0.86]0/1000% to 3.7%
LLM with your keyN: you chooseRun it in your browser
Seed 2020. Exact match and detection (exact match of the candidate list) use Wilson intervals; precision and recall use an item-level percentile bootstrap. The rule engine scored against itself only confirms that the harness works. † Every bootstrap resample gave the same value, so the percentile bootstrap cannot show any uncertainty here. The interval shown instead is a Wilson interval on the pooled token counts, which ignores clustering within messages and so understates the uncertainty somewhat.

At the default N of 20, the intervals are wide. A perfect 20/20 only supports a true exact-match rate above 83.9%, so the default run is a demonstration rather than a benchmark. Raising N to 100 narrows the naive baseline's interval from 34.2% to 74.2% to 42.3% to 61.5%.

The items themselves move the number too. The naive baseline is deterministic, yet across seeds 2020 to 2024 at N = 20 its exact match ranged from 40.0% to 85.0% (mean 54.0%; per seed 11/20, 9/20, 9/20, 17/20, 8/20). For LLMs, the evaluation page also groups repeated runs with identical settings and reports how much they differ.

// AI use statement

What AI does here, and what it never does

What it does

One optional feature uses a language model. The Rules vs LLM evaluation sends each synthetic message to the model you choose and scores its answer against the rule engine.

What it never does

AI never changes the visualiser, the C source, the parity results or any number on this page. Nothing you type or upload is sent to a model, and the site works fully without a key.

Data sent to the provider

Each request carries the fixed instructions, the 10-line dictionary and one synthetic message, together with your key in a request header. Requests go straight from your browser to api.anthropic.com or api.openai.com, and the site has no server of its own. Its Content Security Policy only loads scripts, styles, images and fonts from this site, and only lets the page send requests (fetch, XHR, beacons) to this site and those two APIs. It cannot restrict top-level navigation or code injected by a browser extension, so use a separate key with a low spending limit.

Human in the loop

Every AI output is labelled “AI-generated”. Every call is logged in your browser with its prompts, request settings, attempts, answer, latency and token use, and you can mark each output accepted, edited or rejected in the AI audit log, then export it as JSON or CSV. The log is the single record of these decisions, so run exports read them from it. The key is never logged.

These practices are informed by the Australian Government's policy for the responsible use of AI in government (DTA), the EU AI Act transparency principles and the NIST AI Risk Management Framework. This is not a claim of compliance with any of them. The key handling is explained in DR-005.

// assumptions and limitations

What these results do not show

  • One compiler, one platform. Every fixture was recorded with Apple clang version 21.0.0 (clang-2100.3.34.2) on macOS arm64. The C code reads stack memory it never wrote, and a different compiler could lay that memory out differently and print something else.
  • My generator, my blind spots. The corpus follows families I designed. Perfect parity on them says nothing about inputs the generator never produces.
  • An unverifiable starting point. The submitted 2020 file no longer exists, so the recovered source is shown to be consistent with the official samples, not identical to the submission (DR-001).
  • Agreement, not truth. The LLM evaluation measures agreement with this program's specification, which is not the same as being right about what an emoticon is.
  • Small default samples. At N = 20, intervals are wide, the seed alone moves the naive baseline by 45 percentage points, and model answers can change between runs. The intervals cover sampling over items only; the run-to-run spread on the evaluation page covers the rest.

// what I'd change

What I would do with more time

  • Record the corpus with gcc on Linux as well, and run both in CI, so platform differences show up.
  • Add coverage-guided fuzzing of the C program with AddressSanitizer to find inputs that my hand-designed families miss.
  • Publish a small set of LLM runs once there is a budget, with repeated runs per model and a pre-registered threshold for what would count as an acceptable replacement.
  • Stratify evaluation items by trap type, with enough items per trap to estimate each error rate.
  • Offer a “run the exact bytes” toggle for uploaded files with Windows line endings.

// decision records

Decision records

Each record states the decision first, the options I weighed, and what happened afterwards, including the weak spots. Records are never edited after the fact. A changed decision gets a new record that supersedes the old one.

  1. DR-001Accepted2026-10-05Recover the 2020 solution from a Notability noteI recovered the full solution from the Notability note and transcribed it into coursework/src/program.c exactly as written, without reformatting, renaming or fixing anything. The file in the repository is byte-identical to that transcription, and it is never edited. Any later analysis is added around it.
  2. DR-002Accepted2026-10-05Byte-for-byte parity with the compiled C program as the acceptance criterionThe port is accepted only if, for every recorded input, its stdout is byte-for-byte identical to what the compiled C binary printed and its exit status matches. Where the C program has undefined behaviour that crashes it (a stage-5 buffer overflow), the port cannot reproduce a crash meaningfully, so it must flag the input with a visible warning instead.
  3. DR-003Accepted2026-10-05Convert Windows line endings at the upload boundary, not in the engineThe engine stays faithful. Given \r\n it behaves exactly as the C program does, and the parity tests check this. The upload path converts each \r\n pair to \n before running, and the visualiser shows a notice saying that it did so and why. A bare \r is left alone. Text typed or pasted into the editor never contains \r, because browsers normalise line breaks in a textarea to \n.
  4. DR-004Accepted2026-10-06Evaluate LLMs against the C-faithful rule engine on seeded synthetic messagesThe evaluation asks a model to perform stages 3 to 5 of the program on one message at a time. It returns the candidate tokens left after deleting letters and digits, and the candidates that exactly match the dictionary. The reference answer for every message comes from the TypeScript port, whose parity with the compiled C program is established in DR-002.
  5. DR-005Accepted2026-10-06Optional AI with the visitor's own key, called directly from the browserAI features are optional. Every page works without a key. A visitor who wants to run an LLM evaluation pastes their own Anthropic or OpenAI key into the AI settings dialog.

// model card

Model card

The model card covers both systems that make predictions here, the rule engine and the LLM cleanser. It records their intended use, data provenance, evaluation with intervals, known failure modes and ethical considerations.