Context
The Rules vs LLM evaluation (DR-004) needs calls to a model provider. The site is static and has no server, there is no budget for API usage, and API keys are secrets that must not end up anywhere they could be read or replayed.
Decision
AI features are optional. Every page works without a key. A visitor who wants to run an LLM evaluation pastes their own Anthropic or OpenAI key into the AI settings dialog.
- The key is stored in
sessionStorageby default, so it disappears when the tab closes. "Remember on this device" moves it tolocalStorage. "Forget key" removes keys for every provider from both. - Calls go directly from the browser to the provider. The key is attached only to that request, and this site has no server that could receive it. The page's Content Security Policy loads scripts, styles, images and fonts only from this site, and limits fetch, XHR and beacon requests (
connect-src) to this site,api.anthropic.comandapi.openai.com. It cannot restrict top-level navigation or code injected by a browser extension. - Every call is written to an audit log in the browser's IndexedDB with its feature, provider, model, the system prompt and message sent, the other request settings (max tokens, temperature or effort, how the JSON schema was requested, the retry limit), the number of attempts, the answer, latency, token usage and a human decision. The key is never stored in the log. The log is the only place decisions are stored, so evaluation exports read them from it. It can be reviewed and exported as JSON or CSV at
/ai-log. - Every AI output on screen is labelled "AI-generated".
- Only synthetic, seeded messages are sent. The evaluation never sends anything the visitor typed.
Options considered
- A server-side proxy with my key. This costs money I do not have and would need abuse controls.
- A server-side proxy with the visitor's key. The key would pass through my server, which is exactly the exposure I want to avoid.
- Browser-direct calls with the visitor's key. This is what I chose.
- No AI at all. This is the safest choice, but it would leave the evaluation question unanswered.
Why
Browser-direct calls keep the key between the visitor and their provider. The design is informed by public guidance on responsible and transparent AI use, namely the Australian Government's policy for the responsible use of AI in government (DTA), the transparency principles for AI-generated content in the EU AI Act, and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them. In practice that guidance shows up as labelled outputs, a record of every call, a human decision on each output, and a plain statement of what is sent where.
What happened
Tests check that the key never appears in a request body, an audit entry, an exported file or an error message, including when a provider echoes it back. Provider calls in tests use a mocked fetch, so the test suite never touches the network.
The risk that remains is the one the Anthropic SDK names in its dangerouslyAllowBrowser option. Any script running on the page can read a key held by the page, and that includes browser extensions. The site loads no third-party scripts, and the Content Security Policy limits where the page can load code from and send requests to, but neither can stop a malicious extension. The settings dialog therefore recommends a separate key with a low spending limit.
What I'd change
Use short-lived, narrowly scoped credentials for browser calls if providers offer them, so a leaked key would expire on its own.