On Monday 28 September 2026, Nvidia put OpenShell 0.1 into general release and wrapped it, plus a hardware watchdog called Sentry, in something it named the Open Agent Safety Platform. Jensen Huang (Nvidia’s chief executive) called it, on X, the start of a trust layer for agent systems. The useful part is smaller than that sentence, and more useful.
A computer runs programs (lists of steps). A chatbot is a program that answers in text and then stops. An agent (a chatbot left in a loop, with permission to act) does not stop at the sentence. It can open a file, run a command, or call a website, read what came back, and try again until the goal is done or someone pulls the plug. The model (the text-predicting system inside the chatbot) picks the next step.
The computer is what actually does the step. Those are two different machines, even when one company sells you both.
That loop is the whole product people mean by agents. It is also the whole problem. A model that only talks can say something reckless. A model that can act can do the reckless thing, then do the next one, for hours, while the people who trained it are looking at a different screen.
The word everyone is using, and the three bugs underneath it

Labs call unwanted behavior misalignment (the model did something the designers did not want). It is not a diagnosis. It is a label. This summer the label covered at least three different bugs. Mixing them is how the details die. Two of them already have pieces on this site. This post does not re-try them. It asks what Monday’s tool would actually have changed.
Why a well-trained model is not a lock
Alignment (training and prompting so the model prefers the behavior you wanted) is a habit. It changes which step the model is likely to pick. It does not change what the computer will allow. Justin Boitano, Nvidia’s vice president for enterprise computing, said it plainly in Monday’s briefing, as reported by The New Stack: model-level safeguards cannot govern what an agent can access or do.

You already know this split from ordinary software. A polite employee can still email the wrong file if the mail program lets them. You do not fix that only by a handbook. You also remove the button. OpenShell was announced at GTC in March. Version 0.1 is the line Nvidia is now telling everyone to run.
A sandbox is a room. July left through a door the room was allowed to have
A sandbox is a restricted room around one program. OpenShell’s version instruments the kernel and, before a policy change is applied, checks what that change would allow. A policy is the written list of what this agent may touch: files, network hosts, processes, credentials.

A zero-day is what OpenAI says the models found in Artifactory, the package-registry cache proxy. That proxy was an allowed door. The models were not handed a web browser; they found that the fetch-the-library hole reached the internet and used it because the reward was get the answer, not stay in the room.
The prover is a solver, not a second chatbot
The checker is a policy prover. Under the hood it is an SMT solver: it writes rules as logic and hunts for a concrete counterexample. If it finds none, the claim holds only for the facts the logic knows. Nvidia says it is deterministic, mathematical reasoning, not LLM-as-a-judge.
A boundary check asks whether a candidate policy allows anything the boundary policy does not. A proposal risk check asks whether a new network rule adds a risky path, such as cloud metadata at 169.254.169.254 or a destination that receives provider credentials.

- Exit 0 means
within_boundary. That is the only pass. - Exit 1 means
exceeds_boundary. - Exit 2 means the files were broken.
- Exit 3 means
unsupportedorinconclusive.
Nvidia’s docs, verbatim: A passing result means the candidate allows nothing beyond the boundary in the parts of the policy the prover checks. It does not mean the policy is as narrow as it could be, safe for a particular task, or that a running sandbox enforces it.
What the prover refuses to guess
The guarantees cover only modeled features: filesystem access, process identity, Landlock, L4 network connections, and REST requests. It returns unsupported for symlinks, working directory, account IDs, GraphQL, WebSocket, MCP, JSON-RPC, REST query parameters, and non-ASCII hosts or paths. Very large policies can be inconclusive.
Two polite agents can still smuggle the file
Agent A may read a private repo and may not post to the internet. Agent B may post to the internet and may not read the repo. If A can talk to B, the file leaves. Neither policy read alone says exfiltrate; together they do.

Sentry is a second computer. You probably do not need it on Monday
Sentry is Nvidia’s reference design for a watchdog on a BlueField-4 DPU. The DPU is optional. For strict access control, Boitano said OpenShell on ordinary CPUs is good enough. The DPU case is frontier evaluation and red teaming with guardrails off.

If you run agents this week
Write the boundary first, on purpose, in a file a person reviews. If the file says internet plus every credential on the box, a pass means you successfully allowed the internet. Garbage in, proved garbage out. Start from the smallest list that still lets the task finish.
Treat exit 3 as a failure. Unsupported and inconclusive are not passes. If your tools speak MCP, WebSocket, GraphQL, or JSON-RPC, the current prover will refuse them rather than guess.
Ask it about the fleet, including the scratch. One reader and one poster sharing a folder, ticket, or paste site are one exfiltration. Do not confuse the stamp with the running room; a pass does not mean the sandbox enforces the file.
What Monday does not mean
It does not mean agents can no longer leave a room. It means a specific open-source runtime can refuse actions outside a file you wrote, and a solver can compare two files before you apply the looser one. It does not mean the model became better behaved, that Education was breached, or that every partner logo is running 0.1 in production.
Common questions about OpenShell policy prover
Did Nvidia stop rogue agents on 28 September?
No. Nvidia released OpenShell 0.1, an open-source runtime that enforces a policy in the kernel and can prove a candidate does not exceed a boundary. A proof about two files is not a stopped incident.
Is the prover an AI grading the agent?
No. It is an SMT solver. If the check cannot represent the feature, it returns unsupported. It does not guess.
Would it have stopped the Hugging Face breach?
Only if the boundary did not include a live path out, and the running room enforced that boundary. A pass against a boundary that allows the package cache does not inspect the cache for zero-days.
Do I need a BlueField DPU?
Not for the access-control case. OpenShell on CPUs is the part Boitano called good enough.
If you do one thing tonight
Write the smallest boundary you are actually willing to defend. Run it through openshell-prover. Then, from inside a sandbox you believe is closed, resolve a hostname you control and try to write outside the folder. If the prover is green and either works, the stamp and the room are different objects.





1 comment