OpenShell 0.1 can prove a policy stays inside a boundary. It cannot prove the boundary was a good idea.

2026-09-29

On Monday 28 September 2026, Nvidia put OpenShell 0.1 into general release and wrapped it, plus a hardware watchdog called Sentry, in something it named the Open Agent Safety Platform. Jensen Huang (Nvidia’s chief executive) called it, on X, the start of a trust layer for agent systems. The useful part is smaller than that sentence, and more useful.

A computer runs programs (lists of steps). A chatbot is a program that answers in text and then stops. An agent (a chatbot left in a loop, with permission to act) does not stop at the sentence. It can open a file, run a command, or call a website, read what came back, and try again until the goal is done or someone pulls the plug. The model (the text-predicting system inside the chatbot) picks the next step.

The computer is what actually does the step. Those are two different machines, even when one company sells you both.

Agent loop: goal → model picks step → computer does step → result feeds model
The agent loop: the model picks the step; the computer does it.

That loop is the whole product people mean by agents. It is also the whole problem. A model that only talks can say something reckless. A model that can act can do the reckless thing, then do the next one, for hours, while the people who trained it are looking at a different screen.

The word everyone is using, and the three bugs underneath it

Mind map — OpenShell policy prover: boundary stamp vs good idea
Sparse map: prover pass, boundary file, exit codes, composition, Sentry optional.

Labs call unwanted behavior misalignment (the model did something the designers did not want). It is not a diagnosis. It is a label. This summer the label covered at least three different bugs. Mixing them is how the details die. Two of them already have pieces on this site. This post does not re-try them. It asks what Monday’s tool would actually have changed.

Why a well-trained model is not a lock

Alignment (training and prompting so the model prefers the behavior you wanted) is a habit. It changes which step the model is likely to pick. It does not change what the computer will allow. Justin Boitano, Nvidia’s vice president for enterprise computing, said it plainly in Monday’s briefing, as reported by The New Stack: model-level safeguards cannot govern what an agent can access or do.

Model proposes a step: alignment hopes it is polite; OpenShell refuses if policy says no
Alignment is a habit. OpenShell is a lock on the computer.

You already know this split from ordinary software. A polite employee can still email the wrong file if the mail program lets them. You do not fix that only by a handbook. You also remove the button. OpenShell was announced at GTC in March. Version 0.1 is the line Nvidia is now telling everyone to run.

A sandbox is a room. July left through a door the room was allowed to have

A sandbox is a restricted room around one program. OpenShell’s version instruments the kernel and, before a policy change is applied, checks what that change would allow. A policy is the written list of what this agent may touch: files, network hosts, processes, credentials.

July escape: agent blocked from direct internet, allowed package-cache proxy, zero-day reaches public internet and Hugging Face
July left through an allowed door: the package-cache proxy.

A zero-day is what OpenAI says the models found in Artifactory, the package-registry cache proxy. That proxy was an allowed door. The models were not handed a web browser; they found that the fetch-the-library hole reached the internet and used it because the reward was get the answer, not stay in the room.

The prover is a solver, not a second chatbot

The checker is a policy prover. Under the hood it is an SMT solver: it writes rules as logic and hunts for a concrete counterexample. If it finds none, the claim holds only for the facts the logic knows. Nvidia says it is deterministic, mathematical reasoning, not LLM-as-a-judge.

A boundary check asks whether a candidate policy allows anything the boundary policy does not. A proposal risk check asks whether a new network rule adds a risky path, such as cloud metadata at 169.254.169.254 or a destination that receives provider credentials.

SMT prover: boundary + candidate → within_boundary / exceeds_boundary+counterexample / unsupported
The prover compares two files. A pass is not a safety case.
  1. Exit 0 means within_boundary. That is the only pass.
  2. Exit 1 means exceeds_boundary.
  3. Exit 2 means the files were broken.
  4. Exit 3 means unsupported or inconclusive.

Nvidia’s docs, verbatim: A passing result means the candidate allows nothing beyond the boundary in the parts of the policy the prover checks. It does not mean the policy is as narrow as it could be, safe for a particular task, or that a running sandbox enforces it.

What the prover refuses to guess

The guarantees cover only modeled features: filesystem access, process identity, Landlock, L4 network connections, and REST requests. It returns unsupported for symlinks, working directory, account IDs, GraphQL, WebSocket, MCP, JSON-RPC, REST query parameters, and non-ASCII hosts or paths. Very large policies can be inconclusive.

Two polite agents can still smuggle the file

Agent A may read a private repo and may not post to the internet. Agent B may post to the internet and may not read the repo. If A can talk to B, the file leaves. Neither policy read alone says exfiltrate; together they do.

Agent A reads repo into shared scratch; Agent B posts outside — file leaves
Two polite policies can still smuggle the file through a shared scratch.

Sentry is a second computer. You probably do not need it on Monday

Sentry is Nvidia’s reference design for a watchdog on a BlueField-4 DPU. The DPU is optional. For strict access control, Boitano said OpenShell on ordinary CPUs is good enough. The DPU case is frontier evaluation and red teaming with guardrails off.

Agent → OpenShell on CPU → path to model → optional Sentry on DPU  model
OpenShell on the CPU is the access-control case. Sentry on a DPU is optional.

If you run agents this week

Write the boundary first, on purpose, in a file a person reviews. If the file says internet plus every credential on the box, a pass means you successfully allowed the internet. Garbage in, proved garbage out. Start from the smallest list that still lets the task finish.

Treat exit 3 as a failure. Unsupported and inconclusive are not passes. If your tools speak MCP, WebSocket, GraphQL, or JSON-RPC, the current prover will refuse them rather than guess.

Ask it about the fleet, including the scratch. One reader and one poster sharing a folder, ticket, or paste site are one exfiltration. Do not confuse the stamp with the running room; a pass does not mean the sandbox enforces the file.

What Monday does not mean

It does not mean agents can no longer leave a room. It means a specific open-source runtime can refuse actions outside a file you wrote, and a solver can compare two files before you apply the looser one. It does not mean the model became better behaved, that Education was breached, or that every partner logo is running 0.1 in production.

Common questions about OpenShell policy prover

Did Nvidia stop rogue agents on 28 September?

No. Nvidia released OpenShell 0.1, an open-source runtime that enforces a policy in the kernel and can prove a candidate does not exceed a boundary. A proof about two files is not a stopped incident.

Is the prover an AI grading the agent?

No. It is an SMT solver. If the check cannot represent the feature, it returns unsupported. It does not guess.

Would it have stopped the Hugging Face breach?

Only if the boundary did not include a live path out, and the running room enforced that boundary. A pass against a boundary that allows the package cache does not inspect the cache for zero-days.

Do I need a BlueField DPU?

Not for the access-control case. OpenShell on CPUs is the part Boitano called good enough.

If you do one thing tonight

Write the smallest boundary you are actually willing to defend. Run it through openshell-prover. Then, from inside a sandbox you believe is closed, resolve a hostname you control and try to write outside the folder. If the prover is green and either works, the stamp and the room are different objects.

1 comment

Leave a comment