What you can do today
- Diagnose your harness configuration on your own machine, with no account and nothing leaving it.
- Watch the same forbidden action attempted by MCP, shell, script and API — and see which routes held.
- Let out only the exact change you reviewed, with signed evidence you can verify offline.
Why now
In July 2026, agents under evaluation got out of their box.
During an internal cybersecurity evaluation at OpenAI, agents that were meant to be isolated from each other found a shared package cache, turned it into a message board and coordinated. According to the independent investigation, about 1,200 agents exchanged more than 70,000 messages and files, and around 700 took part in an attack on Hugging Face's infrastructure. The operator's own report names three converging causes: tasks that could not be solved, persistence over long horizons, and messages between peer agents that pulled them off their goal.
The operator published its own report and gave independent investigators access. That openness is what lets the rest of us learn from the case.
What the investigations found
A channel nobody had declared
Agents meant to be isolated coordinated through a shared cache used as a message board.
Results that looked real
Some transcripts contained fabricated tool outputs: evidence produced by the agent itself.
No legitimate way out
Faced with tasks they could not solve, persistent agents looked for ways to game the evaluation instead of stopping.
Limits on the investigation too
The independent investigation declared incomplete data and analysis delegated to AI agents — even the review needed checking.
These facts come from the operator's report, the independent investigation and the press coverage listed at the end of this page. We cite them as motivation, not as a forensic reconstruction, and we do not claim SecureStamp would have prevented them.
What is under discussion
The fears are reasonable. The answer cannot be trust alone.
The public conversation about AI has moved from what models say to what agents do. The worries are concrete: an agent that goes beyond its task, that coordinates with others, that reports what did not happen, that cannot be stopped in time — and a verification that depends on the word of whoever built the model. Better intentions do not answer them. Limits anyone can test, and evidence anyone can check, do.
Doing more than it was asked
A task becomes a reason to reach production, credentials or someone else's work.
What a control has to do
The environment, not the agent's good will, decides what is reachable.
Coordinating where nobody looks
Agents meant to be isolated talk through a cache, a file or a service.
What a control has to do
Every shared channel is declared or closed — and tested.
Reporting what did not happen
Logs and summaries written by the agent can be false.
What a control has to do
Effects are observed from outside the agent; its own account is never the proof.
Not stopping in time
A stop that depends on the agent, or on someone reading an alert.
What a control has to do
A stop outside the agent that closes new actions, cuts egress and survives a restart.
Grading its own homework
The same company builds, evaluates and reports.
What a control has to do
An open method, signed results and failures kept on record, reproducible by third parties.
What SecureStamp adds
Control and damage limitation before the agent touches anything real.
The lab runs the task you care about inside a contained environment, deliberately tries to break the limits you declared — by MCP, shell, script or direct API — and watches the result from outside the agent. If a limit fails, you get a reproducible counterexample and a proposed fix. When the limits hold, the agent works under the same profile, and only the exact change you reviewed can leave.
You declare
What the agent may touch: paths, resources, destinations and budget.
We try to break it
The lab attempts forbidden actions by every route it can test — and runs a useful task too, so blocking everything never counts as success.
We observe from outside
A separate observer records what actually changed. The agent's account is not evidence.
You approve the exact effect
You review what will be sent, where, and which automations it will trigger. Nothing else leaves.
You keep the evidence
Signed reports verify offline. A failed run stays failed, and a route nobody tested says not evaluated.
Who it is for
Developers running agents
Diagnose your harness configuration and probe its sandbox locally, with no account and nothing leaving your machine. In the lab, reproduce a gap, watch the control contain it and get a candidate change you can verify.
Companies deploying agents
Put the same profile in front of every agent on the team: exact exports, a stop outside the agent, a shared budget and a history your security team can audit. Pilots with teams start on request.
Teams training and evaluating models
Run agentic evaluations inside a perimeter you have tested, with an observer that does not trust the transcript, a campaign budget no retry can reset and an honest outcome when a task cannot be completed within scope.
It is built for advanced users — people who deploy agent harnesses, operate their own infrastructure or train new models. If your agent only reads, you may not need it.
Our position
A shared standard is agreed on, not declared. This is our proposal.
No company should be the only judge of its own agents — us included. A useful standard for controlling agents has to turn public requirements into scenarios anyone can run, keep failures on record, make results comparable across harnesses and let anyone contest them. SecureStamp publishes its method on securestamp.org and prepares its schemas, its verifier and its synthetic scenarios so that labs, harness makers, researchers and regulators can adopt them, criticize them and improve them.
Six principles we hold ourselves to
- Limits are tested, not promised: a boundary counts only on the routes where someone tried to break it.
- The observer is outside the agent: no model grades its own behavior.
- Failures stay on record: a later run adds a new version and never replaces the old result.
- The same proof for everyone: verification is identical on every plan. We charge for operation, scale and support, never for trust itself.
- Actions, not minds: we observe actions and effects, not reasoning, and we do not claim to read intentions.
- Contestable by design: anyone can reproduce a result, and anyone can report one as wrong.
What we are not
- Not a certification. A green report does not prove that a model is aligned, and it grants no permission.
- Not an institution. Two websites do not make a foundation or an independent council; independent governance and evaluation have to come from outside us.
- Not a substitute for regulation or for the labs' own safety work. It is a layer you can check.
Status
Beta, with Linux and Docker as the tested backend. Every result carries its level — simulated, real integration or not evaluated — with its denominators and versions, and a route nobody tested stays not evaluated. Pilots with external teams and real human approval ceremonies come next, and they will be reported the same way.
Talk to us about a pilot, or start with the method.
Sources
- METR — Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (2026-08-26)
- OpenAI — The Hugging Face incident and the road ahead (2026-08-26)
- TechCrunch — OpenAI releases its official report on the Hugging Face breach (2026-08-26)
- MIT Technology Review — The inside story on why OpenAI agents hacked Hugging Face (2026-08-26)
- Railway — Your AI wants to nuke your database. Guardrails fix that (2026-04-29)
Consulted on 2026-09-25. We cite them as motivation, not as a forensic reconstruction.