The 1.0 label is real. The useful part is a small computer that runs a script, and a second package that is allowed to remember an unfinished step.
I opened the Pi 1.0 note the day it shipped, 1 October 2026, then the codemode doc and the Pi Durable note. The announcement is a list. The list hides the only design choice that changes how the tool behaves.
The sharp edge, before the tour. A script that fails does not undo tool calls it already made. In the other package, a tool reruns after a crash only if it says that is safe. A deploy is the example they refuse to rerun.
Start from the model, not the product

A large language model (a program that continues text by guessing the next chunk) does not open files. It does not run commands. It does not remember you after the request ends. A token (the chunk it counts, often a word or a piece of a word) is both the bill and the memory limit. The context window (how many tokens it can see at once) is that limit.
Something else has to send the text, run the side effects, and paste the results back. That something is the harness (the program wrapped around the model). A tool (a named function the harness runs when the model asks) is how the model touches your machine. Read a file. Edit a file. Run a command. Those are tools.
Pi, in its own words on pi.dev, calls itself "a minimal terminal coding harness." Terminal means the text window, not a website. Minimal is a product decision, not a compliment. The same page says Pi "skips features like sub-agents and plan mode." A sub-agent (a second agent the first one starts for a side job) is the feature a coding app grows when it tries to be everything. The terminal app still refuses it.


What 1.0 actually adds
Earendil shipped Pi 1.0 on 1 October 2026. Earendil is the company Armin Ronacher and Colin Daymond Hanna formed. A licensing note dated 30 March 2026 already says they had just acquired Pi, and that Mario Zechner, who wrote it, was joining. The Register, on 2 October, dates that acquisition to April. I am not going to pretend those two sentences are the same day. Both put the purchase in the spring. The version bump is October.
The 1.0 note lists what went in.
- Codemode (native support for MCP, and for non-LLM models such as Jev and image models).
- Extension support for virtual models.
- Deferred tool loading.
- Cache warming for Anthropic models.
- Mid-conversation system messages, meaning transcript-aware prompt and tool changes.
- A new terminal theme, and full-screen mode by default.
Both packages are MIT licensed (you can use, change, and ship the code, including in a commercial product, if you keep the license notice). Source is github.com/earendil-works/pi. Docs are pi.dev.
They also say hundreds of thousands of people use Pi every week. That is their figure. I did not see a public count behind it. Treat it as a claim.
A virtual model (an extension that looks like one model to you, while something else picks which real model runs each step) is the demo in the announcement: plan on Claude Opus, implement on GPT, let Jev decide the handoff. Jev, in their words, is a non-LLM model (a smaller model that classifies or chooses, rather than writing the long answer). I am not reviewing Jev here.
Cache warming did not appear for the first time in the 1.0 paragraph. The 0.86.0 notes, dated 19 September 2026, already describe cost-aware prompt-cache warming during long tool runs. The 1.0 list is what the stable release will stand behind, not a diary of the first commit.
The prompt is a bill, so they stopped stapling every tool to it
MCP (Model Context Protocol, a shared way for a program to offer tools to an agent) sounds like a checklist item. The expensive part is not the letters. The expensive part is the prompt (the text you send the model).
Each tool comes with a card: its name, what it does, and the shape of its arguments. Paste forty cards into every request and you pay for those tokens before the model has read your question. You also push your own code toward the edge of the context window. A long tool list is a bill even on the turns you never call the tool.

Codemode is one tool. The model writes a script (a short program, here JavaScript). The harness runs it. Only the script’s output comes back into the conversation. Calls the script made along the way stay out of the model’s view. That is the trick. A script can call ten tools, throw away nine results, and hand the model a paragraph.
The coding-agent doc is specific. The script is raw JavaScript, not a fenced block, and it runs as the body of an async function (a function that can wait for a result). It may start with a line the doc does publish:
// @options: {"max_output_tokens": 2000, "timeout_ms": 60000}
max_output_tokens defaults to 10000. Longer output keeps the start and the end, and the full text goes to a temp file whose path comes back with the result. timeout_ms has no default. Set it if the script might call something slow, such as an image model.
Declared tool cards share a budget of 3000 estimated tokens, the setting named codemode.inlineBudget. Tools marked deferred (held back, not printed in the opening prompt) are not in that list. MCP tools with the default codemode exposure are deferred. The script finds them with searchTools(), which ranks by BM25 (a search score that cares more about rare words than about common ones). The default limit is 8 names. describeTool() and ALL_TOOLS are the other doors.
Two modes. on, the default, still lets the model call ordinary tools by name, and also explains the script. only hides those tools from the model, so the script is the door.
The script does not get your computer
The script runs in QuickJS (a small JavaScript engine), compiled to WebAssembly (a compact program format that can run without being handed the operating system). The doc’s wall is short. No Node APIs. No filesystem. No network. No timers. The VM (the little computer running that script) gets 256 MB. Past that, it throws out of memory.
The only way out is a tool the harness injected, or a non-LLM model the harness exposed. That is a sandbox (a box that can compute, and can ask the harness to act, and cannot act by itself).
Here is the line I would not skip. Tool calls are real. If the script fails halfway, calls that already finished are not undone. Calls still running are cancelled. Promises nobody waited for are dropped. A script is not a transaction (a bundle of steps that all commit, or all roll back).

Earendil’s own account of the reversal is "You Said No MCP!", dated 29 September 2026. Mario Zechner had argued you might not need MCP. The Register calls the 1.0 support a U-turn. The note does not say they lost an argument. It says the MCP of late 2026 is not the MCP they waved off, and that the change they needed was useful for more than MCP. Their sentence: "Ultimately what Pi needs is quite similar to what MCP needs: a sandbox to play with in the form of an interpreter."
An interpreter (a program that runs another program, here the script) is that sandbox. MCP tools in Pi are tools the script can call. The note says codemode is loaded when MCP is configured. Composition (using two tools in one breath, and keeping only the useful part) is what a pile of cards is bad at. A script is how you compose without paying to reread both full results.
Deferred loading is the other half. A deferred tool’s card is not in the opening prompt. On hosts that support it, the card can show up later in the transcript (the saved conversation), at the moment something loaded it. That keeps the cached opening stable. The 0.86 notes already describe this pattern for some hosts. In 1.0 it is part of the MCP path, not a side experiment.
A cache is a photocopy of the opening
A prompt cache (a copy the model host keeps of the unchanged beginning of your request) exists because rereading a long opening is expensive. The copy expires if the session goes quiet. The next turn pays again.
Warming, in the package notes, means a small replay of the last real request, with a tiny output limit, before that expiry. It does not rebuild the chat. It does not run tools. For Anthropic routes the note describes a probe about every 4 minutes, or about every 48 minutes when the request already uses a 1-hour cache. The 1.0 announcement names Anthropic. The mechanism is wider than that one line. If a probe would fire while the agent is busy, it waits.

Mid-conversation system messages are the matching idea for instructions. A system message (text the model treats as the rules, not as your latest question) used to live in the opening. Change the rules halfway and you either break the cache or you forget the change when you resume. Transcript-aware updates store the change in the conversation itself, so a resume still sees it and the old opening can stay cached. That is also in the 19 September notes. The 1.0 list will make it sound newer than it is.
The terminal agent is allowed to die
Pi the coding agent runs in a terminal, for one person. The Durable note says the quiet part. If the process dies, you look at what happened and you tell it to continue. That is the design. They did not want the terminal app to grow a second life.
That is why Durable is a different package. Experimental. Their words: "Pi Durable is experimental, and the API might still change." It does not replace the coding agent. An API (the set of function names a programmer calls) that might still change means do not build a company workflow on it this week and expect next month’s upgrade to be silent.
The terminal app, if you only want that:
curl -fsSL https://pi.dev/install.sh | sh
That line downloads a script from the internet and runs it. Read install.sh first if you care what it executes. Windows is a PowerShell one-liner of the same kind of trust: irm https://pi.dev/install.ps1 | iex.
If you are building the long-running kind:
npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord

The Doom picture sat in the press kit on purpose. A small harness can host a silly full-screen program because the silly program is an extension (code you add, which the core does not have to know about). The press kit’s line is "ask Pi to build what you want, or install a package." I believe the picture more than I believe the adjective supermalleable (their word for "you can bend this into another shape without forking the whole product").
Durable means a checkpoint, not a slogan
A run in Pi Durable is a chain of tasks. Each task writes a checkpoint (a saved note that this step finished) before it moves on. Storage can be memory, SQLite (a database that lives in one file, with no separate server), or JSONL (a text file with one JSON object per line; JSON is a text format for structured data). You can write your own backend. The SQLite and JSONL code avoids Node-only APIs, so a small adapter can run it on Bun (another JavaScript runtime) or inside a Cloudflare Durable Object (a small server that keeps state for one named thing). One process owns a storage at a time. Other clients attach to that process.
If the process dies, a new process opens the same file, finds unfinished tasks, and continues from the last checkpoint. A model request that was cut off is sent again. The partial answer stays in the transcript, marked aborted. A tool call that was cut off reruns only if the tool says that is safe. Otherwise the model is told the call was interrupted, with whatever output was stored, and the model decides. Their example is the one I would tape to the monitor. A deploy interrupted by a crash is reported, never repeated. A search can be marked safe. A deploy must not be.
A request id makes a submission exactly-once (a retry after a crash returns the original submission instead of starting a second one). Queued messages stay queued.
There are no built-in sub-agents. A sub-agent, if you write one, is another conversation. It has its own checkpoint, so it also resumes. A tool marked safe can find that conversation again and wait. The vacation-planner demo in the post is about 1,300 lines, most of them the terminal UI, and it borrows those UI pieces from the coding agent. If it looks like Pi, that is why.
Compaction (replacing older messages with a shorter summary so the next request fits) is itself a task. It can run in the background while the conversation continues. The summary lands at the next turn boundary (the seam before the next request to the model). The conversation only waits if the next request would not fit. If the host still says the request is too long, the harness compacts and retries once. The older messages stay in storage. What the model sees gets shorter. What you can audit does not.
Active transcripts stay in memory. Everything else stays on disk until needed. They argue the in-memory part stays bounded because compaction keeps it inside the context window. A conversation with tens of thousands of messages is a storage claim, not a claim that the model reads all of them.
Any number of clients can attach, because the screen is reading committed state (facts already written down, not a half-drawn animation). A late client gets the current view first, then the changes. thread.watch() sends the small operations of each commit. That is the "more than one human can steer" part. It is real in the design note. It is not a hosted product you log into today.

The sample in the post opens a file named agent.sqlite and sets the model id to gpt-6.1-sol. That is their example, not a benchmark I ran.
What I would do with it
If you want a terminal agent you can reshape, Pi 1.0 is the thing that shipped, and the part to learn is Pi 1.0 codemode, not the theme. Read the codemode doc before you set the mode to only. Remember that a failed script does not unwind the edits it already made.
If you want a job that outlives the process, that is Pi Durable, and it is marked experimental. Do not point it at a deploy, a payment, or a delete until that tool is explicitly not safe to rerun, and you have watched one crash in a scratch repo. The default, if you do not mark the tool safe, is the cautious one: do not repeat the call, tell the model. The danger is marking the wrong tool safe because the happy path is faster.
I would not switch a working setup on the strength of "hundreds of thousands." I would switch a scratch repo, ask it to boil a week of commits down through a script, kill the process on purpose, and see which package actually resumes. The terminal one will not. That is not a bug. It is the line they spent the release protecting.
Common questions about Pi 1.0 codemode
Does a failed codemode script undo its tool calls?
No. Tool calls that already finished stay finished. Calls still running are cancelled. A script is not a transaction.
Is Pi Durable the same as the terminal Pi agent?
No. Durable is a separate experimental package with checkpoints. The terminal agent is allowed to die; you tell it to continue.
What does Pi 1.0 codemode actually change about MCP?
It puts a QuickJS sandbox in front of tools. The model writes one script; MCP tools can be deferred and searched instead of stapled into every prompt. You still pay for real side effects the script triggers.
Should I point Durable at a deploy this week?
Not until the deploy tool is explicitly not safe to rerun, and you have watched one crash in a scratch repo. The cautious default is: do not repeat the call; tell the model.




