Context
An obvious question for a 2026 revival of a text-processing assignment is whether a large language model could do the same job. I wanted to answer it with a measurement rather than an opinion, and I wanted the measurement to be repeatable by anyone who visits the site.
There is no budget for API calls in this project, so any LLM run has to use the visitor's own key (DR-005). That shapes the design: runs must be cheap, small by default, and fully reproducible from a seed.
Decision
The evaluation asks a model to perform stages 3 to 5 of the program on one message at a time. It returns the candidate tokens left after deleting letters and digits, and the candidates that exactly match the dictionary. The reference answer for every message comes from the TypeScript port, whose parity with the compiled C program is established in DR-002.
- Items. Synthetic messages generated from a displayed seed (default seed 2020, default N = 20, at most 100), with a fixed 10-line dictionary that includes two traps,
<3and:D, which can never match because stage 2 deletes their digit and letter from messages. - Primary metric. Exact match of the kept list (order and duplicates included), with a Wilson 95% interval.
- Secondary metrics. Token-level precision and recall of the kept list, micro-averaged, with percentile bootstrap intervals that resample whole messages (2,000 resamples, seed 2020). When every resample gives the same value, as it does for a method with no false positives, the bootstrap interval has zero width, so a Wilson interval on the pooled token counts is shown instead and marked as ignoring clustering. Exact match of the candidate list, with a Wilson interval, measures detection separately from cleansing.
- Comparisons. Any two complete runs of competing methods on the same seed and N can be compared item by item with an exact McNemar test and paired bootstrap intervals for the differences in exact match and F1. The rule engine is never one of the two, because it is the answer key and comparing against it only restates the other run's accuracy.
- Failures. Rate limits, overloads and network errors are retried up to twice with one policy for both providers (honouring
retry-after), so neither provider loses items to infrastructure the other would have recovered from. A call that still fails or returns invalid JSON counts as wrong and contributes no predicted tokens. Errors are reported, not dropped: exact match on the answered items is shown next to the headline figure, and the paired comparison says how many discordant items were failed calls. - Repeats. Runs with the same model, request settings, prompt version, seed and N are grouped, with the spread of exact match and the share of items answered identically, because the item-level intervals do not cover run-to-run variation.
- Anchors. The rule engine scored against itself (a check that the harness works) and a naive baseline that splits on commas and looks tokens up without stripping letters or digits. Both run without a key.
The prompt states the procedure in plain language and asks the model to apply it mechanically. It is versioned (cleanse-v1) and recorded with every run, together with the request settings (max tokens, temperature or effort, retry limit).
Options considered
- A human-labelled gold standard. I have no labelling budget, and "what counts as an emoticon" is a judgement call. The C program is the specification the assignment was marked against, so it is the right reference.
- An LLM as judge. This would measure agreement between models, not correctness.
- F1 as the headline. F1 rewards partial answers. A downstream system that consumes cleansed messages needs the whole list right, so exact match leads and F1 shows how close the misses are.
- One batched prompt for all items. This is cheaper, but one failure ruins the run, items can influence each other, and the audit log loses per-item detail. I chose one call per item.
- Real social-media messages. Sending real people's text to a third-party API would raise privacy questions that synthetic messages avoid entirely.
Why
Item-level resampling matters because tokens inside one message are not independent. If a model misreads a glued emoticon such as hi:), it tends to make the same mistake on every glued token in that message, so a token-level bootstrap would understate the uncertainty. Wilson intervals stay sensible at 0/N and N/N, which is where small evaluations often land.
The traps are deliberate. A model that "knows" <3 is a heart will keep it, while the program deletes the 3 in stage 2 and drops what is left. Those are the cases where following a specification and following intuition disagree, and they are what the evaluation is designed to expose.
What happened
Without a key, the harness runs both anchors. On the default items (seed 2020, N = 20) the rule engine scores 20/20 against itself, which only confirms the plumbing. Its precision and recall are 1.00 in every bootstrap resample, which is exactly the zero-width case above: the pooled Wilson interval on its 37 tokens is 90.6% to 100%. The naive baseline gets 11/20 messages exactly right (55.0%, Wilson 95% CI 34.2% to 74.2%), with precision 0.91 (95% CI 0.78 to 1.00) and recall 0.81 (95% CI 0.67 to 0.92). It never gets the candidate list right (0/20, 95% CI 0% to 16.1%), because it does not delete letters and digits.
The items matter as much as the method. The naive baseline is deterministic, yet on seeds 2020 to 2024 at N = 20 its exact match ranged from 40.0% to 85.0% (8/20 to 17/20), which is what the wide intervals predict.
I have not published any LLM results. I have not spent money on API calls for this project, so LLM numbers exist only in the browsers of visitors who run the evaluation with their own key, and they stay there unless exported.
N = 20 is small. Even a perfect 20/20 only supports a true exact-match rate above 83.9% at 95% confidence, so the default run is a demonstration, not a benchmark. The interface lets a visitor raise N to 100 when they want a tighter interval.
What I'd change
Once there is a budget, publish a small set of LLM runs with repeated runs per model, so run-to-run spread can be reported next to the sampling interval. Stratify the generated items by trap type, with enough items in each stratum to estimate per-trap error rates. Decide in advance what exact-match rate would make an LLM an acceptable replacement, and write that threshold down before running.