Skip to content
methods

model card

The emoticon cleanser and its LLM comparison

docs/model-card.md

This card covers the two systems that make predictions on this site. The first is the rule engine, a TypeScript port of my 2020 C program for COMP10002. The second is an LLM used as a cleanser, which a visitor can run with their own key and compare with the rule engine. Nothing here is a trained statistical model of my own. The card exists so that the intended use, the evidence and the known failure modes are written down in one place.

  • Owner: Sunchuangyu (Rin) Huang
  • Last updated: 2026-10-06
  • Related records: DR-002 (parity), DR-004 (evaluation design), DR-005 (bring-your-own-key AI)

System details

AspectRule engineLLM cleanser
What it isLine-for-line port of coursework/src/program.c (stages 1 to 5)A general-purpose model given the stage 3 to 5 procedure in plain language (prompt cleanse-v1)
VersionsThe 2020 submission, recovered in 2026 (DR-001)Chosen by the visitor: Claude Haiku 4.5 (default) or Claude Sonnet 5.5, or an OpenAI model id they type
OutputExactly what the C program printsJSON with two lists, the candidate tokens and the kept emoticons
DeterminismDeterministicNot guaranteed. Haiku runs at temperature 0, Sonnet 5.5 at low effort with default sampling, and OpenAI models with their defaults. The settings are saved with every run and audit entry
Where it runsIn the browserAt the provider, called from the browser with the visitor's key

Intended use

  • Teaching: showing how each stage of a small C program transforms its input, including the quirks of the submitted code.
  • Evaluation: measuring how closely an LLM follows a precise, deterministic text-processing specification.

Out of scope: moderating or analysing real social-media content, inferring anyone's sentiment, and any decision about a person. Neither system was built or tested for those uses.

Data provenance

  • Rule engine. No training data. The logic was written by hand in 2020 on top of a staff skeleton. The emoticon dictionary is part of each input.
  • LLM cleanser. Pretrained by its provider. I have no visibility into its training data, and this project does no fine-tuning. Prompts sent to a provider are subject to that provider's API data policy.
  • Evaluation inputs. All synthetic. The differential corpus is 500 inputs generated with seed 10002 by scripts/generate_differential_corpus.py. The LLM evaluation uses messages generated in the browser from a displayed seed (default 2020) and a fixed 10-line dictionary. No real messages or personal data are used anywhere.

Evaluation

Rule engine against the compiled C program

The reference is the stdout of program.c compiled with Apple clang 21 (-Wall -std=c99) on macOS arm64.

EvidenceResult
Official sample tests2/2 byte-for-byte (95% CI 34.2% to 100%)
Hand-picked edge cases39/39 completed runs match (95% CI 91.0% to 100%), 2/2 crashes flagged (95% CI 34.2% to 100%)
Seeded corpus (seed 10002, 500 inputs)481/481 completed runs match (100%, Wilson 95% CI 99.2% to 100%)
Crashes in the seeded corpus19/19 flagged by the port (95% CI 83.2% to 100%)
Tokenisation invariants9 properties, 300 generated cases each (fast-check, seed 10002), all hold

LLM cleanser against the rule engine

Primary metric: exact match of the kept list per message, with a Wilson interval. Secondary metrics: token-level precision and recall with item-level bootstrap intervals, and exact match of the candidate list (detection) with a Wilson interval. Paired runs of competing methods on the same items are compared with an exact McNemar test.

Method (seed 2020, N = 20)Exact matchPrecisionRecallDetection
Rule engine against itself (harness check)20/20 (100%, 95% CI 83.9% to 100%)1.00 (95% CI 0.91 to 1.00)*1.00 (95% CI 0.91 to 1.00)*20/20 (95% CI 83.9% to 100%)
Naive split baseline, no AI11/20 (55.0%, 95% CI 34.2% to 74.2%)0.91 (95% CI 0.78 to 1.00)0.81 (95% CI 0.67 to 0.92)0/20 (95% CI 0% to 16.1%)
LLMNot run by the author (no API budget). Visitors produce their own results.

* Every bootstrap resample gives 1.00, so the percentile bootstrap shows no uncertainty. The interval given is a Wilson interval on the 37 pooled tokens, which ignores clustering within messages.

With N = 20 the intervals are wide. Even 20/20 only supports a true exact-match rate above 83.9% at 95% confidence. The items matter too: on seeds 2020 to 2024 the naive baseline's exact match ranged from 40.0% to 85.0%. The intervals above cover sampling over items only. For LLMs, the evaluation page also reports the spread across repeated runs with identical settings.

Known failure modes

Rule engine (inherited from the 2020 code and kept on purpose)

  • The first message is read with a 50-byte limit instead of 280. A longer first line ends the program.
  • A first message that starts with a comma reports one token too few.
  • When a message's last token is dropped in stage 5, its newline is not printed and the next message runs on.
  • With an empty dictionary, stage 5 drops commas as well as emoticons.
  • Stage 5 stores tokens in a 2,500-byte buffer. More than 50 slots (about 25 tokens) overflows it, and the compiled program crashes. The port warns instead.
  • Tokens of 50 bytes or more can never match the dictionary.
  • Lengths are counted in bytes, so non-ASCII emoticons report longer lengths than they look.
  • Windows line endings produce empty lines and an empty dictionary (DR-003).
  • The port assumes zero-filled stack memory where the C code reads bytes it never wrote. This held on the reference platform but is not guaranteed elsewhere.

LLM cleanser (failure modes the evaluation is designed to expose)

  • Keeping emoticons that contain letters or digits, such as <3 and :D, which the program can never match.
  • Missing emoticons glued to words, such as hi:), which the program keeps after stripping the letters.
  • Dropping duplicates or changing their order.
  • Returning candidates or kept tokens that are not in the message at all.
  • Refusals, truncated answers and answers that do not fit the JSON schema. These are counted as wrong, not removed. Rate limits and outages are retried up to twice with the same policy for both providers, and calls that still fail are reported separately from wrong answers.
  • Different answers on repeated runs.

Ethical considerations

  • Privacy. Only synthetic messages are sent to a provider. The evaluation never sends anything a visitor typed.
  • Keys. The visitor's key stays in their browser and is sent only to the chosen provider (DR-005).
  • Transparency. AI outputs are labelled "AI-generated". Every call is logged in the visitor's browser without the key, can be reviewed with a human decision, and can be exported.
  • Cost. Each LLM run costs the visitor a small amount of money. The default of 20 one-message calls keeps it low.
  • Governance. These practices are informed by the Australian Government's policy for the responsible use of AI in government (DTA), the EU AI Act transparency principles and the NIST AI Risk Management Framework. This is not a claim of compliance with any of them.

Caveats

The parity evidence comes from one compiler on one platform, and from inputs produced by a generator I wrote. The LLM comparison measures agreement with this program's specification, which is not the same as being right about what an emoticon is.