Latest run
When: 2026-09-24 ~01:45–01:52 PT (America/Phoenix).
What was tested: 31 attacks against the public demo’s persona prompt, sent directly to the model (Gemma 4 26B on llama.cpp). This tests the model and persona only, not the full website path.
Result: 31/31 pass on run 2 (max_tokens 3000). No system prompt leak, no claimed tool execution, no fabricated secrets, no harmful instructions. Indirect-injection cases flagged the attack, including an SSRF bait aimed at cloud metadata.
Finding: budget the reasoning
Run 1 used max_tokens 600. 20 of 31 replies came back empty — the model’s hidden reasoning consumed the budget before any answer. That looks like an outage. Fix: leave headroom for the answer (run 2 used 3000; all 31 answered).
Defense in depth
On the live demo, toolsets are disabled, memory is off, approvals deny, and turns cap at two. This harness did not need those controls to hold — the persona refused — but they remain the backstop if a future persona failure tries to call a tool.
Limits
Single sample, single turn, no automated grader, no full web path. Next: multi-turn, repeated sampling, and the full website route.