Context
The visualiser runs a TypeScript port of coursework/src/program.c in the browser. Every explanation on the site depends on the port behaving like the original, so I needed a test of "faithful" that could not be argued with.
The original program has behaviour that only shows at the level of individual bytes. Stage 5 suppresses a newline when a message's last token is dropped, so two messages run together. The first line is read with a 50-byte limit instead of 280. Lengths are counted in UTF-8 bytes, so non-ASCII emoticons report longer lengths than they look.
Decision
The port is accepted only if, for every recorded input, its stdout is byte-for-byte identical to what the compiled C binary printed and its exit status matches. Where the C program has undefined behaviour that crashes it (a stage-5 buffer overflow), the port cannot reproduce a crash meaningfully, so it must flag the input with a visible warning instead.
The recorded inputs are the two official sample tests, 41 hand-picked edge cases, and a seeded corpus of 500 generated inputs (seed 10002) that concentrates on the program's limits.
Options considered
- Semantic equivalence. Accept the port if it keeps the same emoticons per message. This ignores the newline and comma behaviour that the marker's tests would have seen.
- Normalised comparison. Compare after trimming whitespace and blank lines. This would hide the stage-5 line-merging quirk, which is one of the behaviours the site exists to explain.
- Byte-for-byte parity on recorded runs. This is what I chose.
- A clean reimplementation of the specification. This is easier to read but would show a program I never wrote.
Why
The original was marked by comparing stdout with expected output, so byte equality is the criterion the program was built for. It also forces the port to model C strings as bytes rather than JavaScript strings, which is why lengths, limits and NUL handling come out right.
A seeded random corpus adds something hand-picked cases cannot. I chose the edge cases knowing the code, so they share my blind spots. The generator produces inputs near every limit on both sides, and fixing the seed means anyone can regenerate exactly the same 500 inputs.
What happened
- Both official sample tests match byte for byte.
- Of the 41 hand-picked cases, 39 match exactly and 2 crash the C binary, which the port flags.
- Of the 500 seeded cases, 481 ran to completion in C and 19 crashed. The port matched 481/481 of the completed runs (100%, Wilson 95% CI 99.2% to 100%) and flagged 19/19 of the crashes (95% CI 83.2% to 100%). The overflow warning fired on exactly those 19 inputs and on none that the binary survived.
These numbers have limits. A perfect observed rate on 481 cases still allows a true mismatch rate of up to about 0.8% at 95% confidence. All fixtures were recorded with one compiler on one platform (Apple clang 21 on macOS arm64). The port assumes that stack memory the C code reads without writing is zero-filled, which held for every recorded run but is not guaranteed by the C standard, so a different compiler could print something else. Finally, the corpus follows a generator I wrote, so 100% parity says nothing about inputs it never produces.
What I'd change
Record the corpus with gcc on Linux as well and run both in CI, so platform-dependent behaviour becomes visible instead of assumed. Add coverage-guided fuzzing of the C program (for example libFuzzer with AddressSanitizer) to find inputs that my generator's families miss.