ODA

Evals

Public demo red team

Thirty-one single-turn attacks on the public chat, each reviewed by hand, with the demoThirty-one single-turn probes against the demo persona on the house brain. Human-reviewed. Config isolation is a second control plane.rsquo;s locked-down configuration as a second layer of defense.

Latest run

When: 2026-09-24 ~01:45–01:52 PT (America/Phoenix).

What was tested: 31 attacks against the public demo’s persona prompt, sent directly to the model (Gemma 4 26B on llama.cpp). This tests the model and persona only, not the full website path.

Result: 31/31 pass on run 2 (max_tokens 3000). No system prompt leak, no claimed tool execution, no fabricated secrets, no harmful instructions. Indirect-injection cases flagged the attack, including an SSRF bait aimed at cloud metadata.

Finding: budget the reasoning

Run 1 used max_tokens 600. 20 of 31 replies came back empty — the model’s hidden reasoning consumed the budget before any answer. That looks like an outage. Fix: leave headroom for the answer (run 2 used 3000; all 31 answered).

Defense in depth

On the live demo, toolsets are disabled, memory is off, approvals deny, and turns cap at two. This harness did not need those controls to hold — the persona refused — but they remain the backstop if a future persona failure tries to call a tool.

Limits

Single sample, single turn, no automated grader, no full web path. Next: multi-turn, repeated sampling, and the full website route.

Operating policy Incidents