Why Ernie exists
Most agent interfaces begin with chat. Chat accepts almost any request, but it also funnels tools, code, artifacts, and delegated work through one transcript.
Consider a coding task that creates a worktree, a diff, a test run, and several delegated tasks. A chat window can report each result, but it does not show how those objects relate. The reader has to reconstruct the structure from messages.
That is the gap Ernie explores. It is a desktop interface for Prime Agent. Prime Agent can divide a task into smaller parts, delegate those parts, and repeat the process as needed. Ernie asks how the interface can show that changing structure instead of flattening it back into messages.
Omar Khattab names the technical distinction behind this change: compaction revisits a running state, while recursion lets an agent select and work on parts of its context.
The recursive part matters. A Recursive Language Model, or RLM, does not have to hold the whole problem inside one model call. It can inspect the available context with code, split it into smaller pieces, and call models on those pieces. If one piece remains too large, it can split that piece again. The model chooses how to divide the work instead of following only steps fixed by the application in advance.
Prime Agent gives an agent this kind of working environment: tools, code execution, artifacts, and other agents it can delegate to. Ernie still has to present the resulting work through a familiar shell.

Ernie v0.1.0 connects to Prime Agent’s daemon, groups sessions with repositories and Git worktrees, exposes narrow UI controls to agents, and hosts lifecycle-managed plugins. The public Apple silicon prerelease was published on Aug 16, 2026. The task-shaped interface described below is a research direction, not a v0.1.0 feature.


Ernie v0.1.1
Keep each Prime Agent session beside its repository and Git worktree in one macOS workspace.
- Apple silicon
- Public prerelease
Prime Agent can change how it organizes the work, but Ernie v0.1.0 can only present that work through fixed panels. That mismatch produced a question: if the RLM can change the shape of the work, why can’t it change the shape of the interface?
I want Ernie to become an RLM-able interface: an interface the agent can use during recursive decomposition. Remote control alone is too narrow. The interface itself should become a trusted language the agent can use while organizing the work.
If the task becomes coding, the interface should bring out the worktree, diff, terminal, tests, and review state. If it becomes research, it should bring out the notebook, sources, graphs, datasets, and experiment runs. The useful objects are different because the work is different.
Instead, we tend to pour both jobs into a transcript and ask the rectangle to remain optimistic.
I call the alternative jellyware: software whose interface waits to see what the work becomes before settling into shape. It avoids constant motion and arbitrary model-generated pixels. The jelly is in the composition. The shell can stay stable while the surfaces inside it form around the task.
Ernie v0.1.0 already exposes a smaller mechanism: a UI CLI. An agent can inspect a live, versioned capability manifest, then issue typed commands to focus the window, choose a light or dark theme, show or hide the sidebar, and set its width within a declared range.
The agent speaks a small, explicit language for changing the surface. Coordinate clicking and guesswork fall outside that contract.
This offers interface control. Interface orchestration remains the next step. The available nouns and verbs come from the host: window, theme, sidebar; focus, show, hide, resize. Jellyware asks for a larger, trusted language in which task decomposition can select and compose useful surfaces. Operating predetermined controls is only the starting point.
Why chat is not neutral
A chat box looks neutral because it accepts almost any request. It is not neutral. It assumes that work begins with a message and ends with a response. Branches, files, tests, experiments, failures, and child agents must return as text before the interface knows what to do with them.

The model may have changed. The interface is still waiting for the answer.
The same problem exists below the interface, in the agent harness. A harness is the software around model calls: the machinery that decides how a task is decomposed, which tools run, what gets observed, and whether the process continues. RLMs widen that decomposition language instead of leaving every strategy frozen inside the harness.
The Mismanaged Geniuses Hypothesis proposes that some limits we attribute to models may instead come from the systems managing them. One experiment tested this idea with a task that hides pieces of information, called “needles,” inside a long input and asks the model to retrieve them. On this benchmark, called MRCRv2, RLM(Qwen3-4B-Instruct) performed near 0% at one million tokens with eight needles. Researchers then trained the RLM with reinforcement learning on a much easier version: 32,000 tokens with one needle. After that training, it reached 100% on the larger eight-needle version. This result does not establish a universal capability. It shows that a strategy learned on an easier retrieval problem can transfer to a much larger one.
MGH concerns composing model calls. My extension concerns interfaces: if decomposition changes what an agent can accomplish, perhaps the visible environment should preserve and participate in that decomposition. A research subtask that creates a source set, notebook, and graph should retain that structure instead of melting back into prose before the person can work with it.
Existing workbenches are becoming more flexible
Several tools have already loosened parts of the workbench. Neovim makes it programmable. Herdr makes terminal panes agent-aware. bb lets agents extend an IDE through threads, worktrees, panels, commands, and plugins. Mitchell Hashimoto’s announced direction with Superlogical points toward durable sessions that outlive any one local window.
These systems move different boundaries: programmability, agent awareness, self-extension, and session durability. Jellyware asks the remaining question: can the task decomposition itself determine why a trusted surface exists, what makes it valid, and when it should leave?
A changing interface needs lifecycle rules
A task-shaped interface is not only a layout. Its panels can open files, start processes, register commands, and hold native resources. When a panel changes or disappears, the runtime must also clean up those effects.
An earlier Ernie plugin host could leak a native resource when activation failed. The plugin never appeared, but the resource stayed in the host process. The interface looked clean while the runtime was not.
This is the problem that Cordis helped me name. It models three questions: what capabilities are available here, how long is this arrangement valid, and who owns the consequences when it ends?
On Aug 16, 2026, the DeepSeek Harness developer-preview page described models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and even the UI as plugins in a shared runtime. It offers a broader example of the same idea: a harness can be assembled from capabilities instead of built as one fixed loop.

In Cordis, contexts, fibers, and effects connect availability to lifetime and cleanup. A component appears where its requirements are satisfied, remains while that situation is valid, and cleans up its changes when it leaves. A task-shaped notebook, diff, or browser needs that discipline because the host must know why it exists and when it should disappear.
Effect ownership still depends on honest cleanup functions and complete registration. Recursively composed task interfaces also remain unproven in DeepSeek Harness. Together, Cordis and the harness show that the runtime underneath such an interface can be assembled from accountable parts.
How Ernie applies lifecycle ownership
Ernie began without Cordis and still uses its own lifecycle model. The local resource leak revealed the same design problem: a temporary interface can disappear cleanly while leaving consequences behind.
Ernie now registers the cleanup for each successful acquisition immediately through an EffectScope. If activation fails halfway through, Ernie still cleans up everything acquired before the failure. Each interface piece has an owner, dependencies, and a lifetime.
Ernie’s API v3 plugin host also activates providers before consumers, reverses that order during teardown, and restores demanded consumers when providers return. It still does not implement Cordis’s broader context model.
Ernie v0.1.0 stops at interface control. Its UI CLI can focus the window, choose a light or dark theme, show or hide the sidebar, and set its width from 192 through 384 pixels. It cannot compose a diff, notebook, dataset, or browser around a task. That composition is the next experiment.
A flexible interface still needs design judgment
Composable capabilities create a second design problem: a workbench that can change shape needs judgment about which shape to take.
The useful interface may become clear only after the task’s objects and relationships become visible. That does not mean every dataset wants to become a rotating donut. It means the fixed screen is one answer, not a law of nature.
Charles Eames gives the experiment a necessary constraint. Eames, the American designer, architect, and filmmaker who worked with Ray Eames, described design as arranging elements toward a purpose and as a method of action within limits. The designer’s skill lies partly in recognizing the constraints.
Jellyware must work within many constraints: attention, permission, accessibility, latency, screen space, resource cost, reversibility, and lifetime. A surface may be available but distracting. It may fit the task and fail the person using a screen reader. It may be useful for thirty seconds and expensive for thirty minutes. A flexible interface creates more decisions, not fewer.
Wilson Miner explains why those constraints matter beyond one session. As digital tools become part of where people live and work, interface decisions make some actions easy, teach habits, preserve or hide evidence, and set expectations about how work should proceed. We shape tools, then continue living inside the behavior they encourage.
I only make Electron apps sounds modest. Even an Electron app decides what becomes prominent, what can be ignored, and which actions feel easy. Those local choices are enough responsibility to take seriously.
An RLM-able interface will make some evidence prominent and other evidence easy to miss. It will make some actions feel natural, decide what counts as finished, and teach people what deserves attention. Those are design decisions even when a model selects the panel and the panel lasts thirty seconds.
That responsibility remains when a surface is temporary or an agent requested it. A generated research view can privilege one source. A coding surface can make shipping easier than reviewing. A disappearing task panel can leave behind files, processes, permissions, or learned habits. If an agent composes the workbench, the host remains responsible for what may appear, what may persist, and how it leaves.
Cordis turns that design responsibility into a runtime rule: whoever creates a temporary surface must also clean up its effects.
The interface should follow the work

I want Ernie to become a place where the task can bring its own working shape without gaining permission to redecorate the universe.
Each task should surface the objects it creates and the capabilities it needs. The interface can follow those objects, capabilities, and lifetimes without guessing a profession from a dropdown. Stable chrome can preserve orientation while the central workspace changes. The agent can compose inside declared limits without redesigning the entire app.
Building these environments means choosing what people and agents notice, trust, consider finished, and leave behind. Those decisions begin with ordinary details such as evidence visibility, permissions, review friction, and cleanup.
Familiar interfaces also have histories. These four artifacts are a reminder that the forms we now treat as settled once had to be invented.




Chat is a useful first shape, not the final one. Jellyware is one attempt to discover what should come next and when it should disappear. Ernie is where I am testing that idea.
Huge thanks to Azwan Zuharimi for the ideas that helped shape this essay.