
Building an Agentic Harness with Jev: A Builder's Guide
Avid's 3,500-word build log: Jev as a decision layer inside keel, a local-first Rust coding workspace. One principle carries it: a model can suggest the next move, the host still owns the move — the host prepares the menu, Jev picks from it, the host checks again; pinned routes and live sessions bypass automatic selection; selection grants no permission. Plus how decision receipts become replayable evaluation scenarios.
This post is based on the Builder's Guide "How to Build Agentic Harness using Jev" published by Avid (@Av1dlive) on Sep 23, 2026, together with the code and docs of the companion open-source project keel 0.2.0. It is a rare, practical record of "boundary engineering": instead of showing off, the author keeps asking the same question — where can the model make a choice that the host code can still verify?
TL;DR: if you'd rather skip the 3,500-word essay, go straight to github.com/codejunkie99/keel — hand the repo to your agent. To understand why each layer is designed this way, read on.
Before the Code: Give the Decisions a Shape
The starting point is modest: I want a coding agent that gets better, and I want to know what changed when it does. The system still needs to choose a model and limit its tools. Then it needs to check the choice and judge the result. That's where Jev engineering comes in — give the system a choice you can inspect, then keep enough evidence to judge whether that choice helped.
Each useful decision needs three things:
- Named options;
- A result the code can check;
- A record of what happened afterward.
The last one matters: an extra model call still has to earn its place. The article walks through four parts: how Jev is used as a decision layer, why a coding harness is a useful place to put it, what was built in keel, and what "self-improvement" should mean (and what it doesn't prove yet). The sentence that runs through everything:
A model can suggest the next move. The host still owns the move. That distinction sounds small. It changes the architecture.
Figure 1: The host prepares the menu; Jev picks from it; the host checks the pick again.
01 / How Jev Is Used: A Small, Clear Job
In this build, Jev chooses from a list the host prepares. "Host" means the application code around the model — that code decides which options are available before Jev gets a say. It returns a typed result in a fixed, checkable format meaning "pick this one" or "I can't choose." That's the whole contract.
The host prepares the menu; Jev picks from it; the host checks the pick again.
Keeping the interface narrow has a practical reason: let a model return any instruction it likes, and every part of the app that calls it has to work out what that instruction means. A fixed format makes the job pleasantly dull — read the result, check it against the current state, accept it or use the fallback. A good analogy: a dispatcher at a busy workshop can choose which bench gets a job, but can't quietly rewrite the safety rules for the saw.
Figure 2: The two moments the selector can act — new-task routing and in-loop next-step focus.
Two Places Where the Selector Acts
1) Choose a route for a new task. When a fresh task has no pinned provider or model, the host prepares eligible candidates; Jev or local Laya chooses among them or abstains. Pinned routes, live sessions, and resumed sessions are preserved untouched. "Pick the best model" needs a boundary: the host must be ready to run every option it offers, so filter the options before Jev sees them — removing providers not installed or enabled, models that can't be found, unsupported reasoning levels. (The provider still checks its own login at session start, which can fail even after a route passes the host's checks.)
2) Choose the next step inside DeepSeek. The loop offers four focus choices:
- inspect
- implement
- verify
- answer
Figure 3: inspect / implement / verify / answer each map to a host-defined tool bundle.
Each choice maps to a host-defined tool bundle. The selector returns a focus id; the host finds its allowed tools, then checks them against current tool definitions before use. "Answer" is a good example — it gets no tools. If the selector says answer, the system doesn't hand it a shell and hope it behaves; the allowed bundle is empty because that's what answer means here.
A small example clarifies the split: ask the app to fix a parser bug. A pinned provider stays pinned; for a fresh unpinned task with automatic selection enabled, the host offers eligible routes — the selector chooses within that list; it can't summon a missing provider. Later, inside the embedded DeepSeek loop, the decision is smaller: choosing verify changes the admitted tool bundle — which doesn't prove the patch is correct; a check still has to run and its result still has to mean something. This is where the design clicks: "which worker?" and "which tools next?" deserve separate boundaries, because they're different jobs with different failure modes.
Figure 4: Host, decision layer, and provider-owned inner loops — ownership should be visible on the diagram itself.
The Important Boundary: External Agents Keep Their Own Loops
Keel can host several coding providers through the Agent Client Protocol (ACP). Those providers run their own tool loops and may expose controls for tools, context size, or handoffs. But in this build, Jev and Laya don't take over those internals. The reason is practical: we can read a provider's capability descriptions, but a description alone doesn't give the host a way to run that action — a listed slash command doesn't automatically become an action Jev can call (in the current path, a slash command may just become prompt text for the provider).
Drawing another arrow on the architecture sheet doesn't create an API. The boundary follows the code we control.
In practice there are three layers: a host, a decision layer, and possibly a provider-owned agent loop. The host can route a new task to a provider — that says nothing about who controls the provider's next tool call. Keel's capability audit (listing commands, skills, modes, and tools) helps draw that line.
One adjacent surface is computer use: keel computer-use decide lets Laya or Jev select a host-prepared, low-risk action id or abstain — but the command doesn't execute desktop actions. Choosing a proposed click and carrying out a click are separate responsibilities. Compatible ACP agents can use installed desktop tools through MCP; those agents still own their internal loops.
Choice and Permission Are Different Things
Jev can select a candidate. It cannot approve a dangerous action. Tools needing permission go through the normal permission path; user approval still applies where the app requires it. Model selection doesn't grant authority, and neither does a high confidence value.
Check the State Again Before Acting
Before applying a choice, the host checks that the action still matches the current task state; if the provider list changed or the route is stale, it can reject and fall back. Inside the DeepSeek loop, the named tool is checked again before running. Both checks serve the same purpose: keep an old decision from acting on a changed system. "That's the boring part. Boring is what you want from safety checks."
Figure 5: The route-selection paths — none of them grants permission to run a tool.
Give Uncertainty a Clear Path
"I can't choose" is a valid result. The host still needs to know what to do next. Define the fallback before calling the selector, then record when it was used — so a fallback doesn't get counted as a selector win. After validating the result and observing the outcome, the host writes a decision event or receipt, letting us ask ordinary engineering questions: What were the candidates? Did the selector choose or abstain? Did validation accept it? Did a fallback happen? What tool outcome followed? — These records are not private chain-of-thought. A model's explanation doesn't prove why something happened.
02 / Why Put It in a Coding Harness?
People often talk about a coding agent as if it were one call: prompt in, patch out. Real work is more like a series of forks:
- Which model should handle this task?
- Should the agent inspect files first or start editing?
- Which tools are relevant right now?
- Did the repo change while the decision was in flight?
- Should the system retry, abstain, or ask the person?
- Did the change compile, pass the focused checks, and preserve the user's intent?
These decisions have different costs: a slow model on a tiny rename wastes time; and a wrong choice can cost more than a slow one — a fast model may need more retries on a large migration; a tool choice harmless in one workspace can be consequential in another. The harness holds the details that shape each action: current repository state, available providers and tool definitions, permissions and task status, results of checks. The model is one part of that system. It isn't the system.
A standalone decision demo can show a model choosing between three labels, but it doesn't tell you whether the choice was usable. Inside a coding harness, the host can answer concrete questions: were these choices actually available? was the task new, pinned, or already running? did the selected tool still exist when execution began? did the permission layer allow the action? what result came back? did verification find a regression? The author's stance is blunt: show me the options, the choice, and the result. Without those, a claim that the agent improved is mostly vibes.
Figure 6: Laya, Jev, and Normal — each path's requirements and host checkpoints.
Three Modes, Three Different Choices
Laya and Jev are separate selectors; Normal mode skips selection. Core ML is Apple's local model runtime; Jev uses a protected credential already on the machine, with no key-entry screen in this build. "Always use the smartest model" ignores latency, privacy, credentials, and user intent — the mode choice gives those tradeoffs a place in the app.
The Extra Call Has to Earn Its Place
An important tangent: routing can become its own little bureaucracy. If choosing a worker takes longer than the task, we've made a very elaborate waiting room. The comparison to want is total task cost: selector time, worker time, retries, and review effort.
A cheaper decision call is useful only if the whole job benefits. This article hasn't demonstrated that payoff with a coding benchmark. There's another limit hiding in plain sight: the selector can't choose a good option the host forgot to offer.
A good harness draws a hard line between deciding and doing: for route selection the model receives candidate ids — labels referring to host-prepared options; those ids don't let the model invent a new command or provider. Inside the DeepSeek loop, the selector returns one focus id, which the host turns into a specific set of allowed tools, then checks again. That means the decision layer is replaceable — the caller doesn't need to trust its prose, only the host's checks, which need to be clear enough to inspect. This is why the small interface is likable: you can point to the exact choice made and the exact tools it admitted. Add more freedom and you've also added more behavior to explain when a run goes sideways.
At the same time, the harness should preserve the messy parts that matter: show abstentions, show rejections of stale actions, record fallbacks to the user's original route. Those details let us trace: the host offered these options → the selector returned this id → the host accepted or rejected it under these rules. Honesty requires the counterpart: putting a decision layer inside a harness doesn't make every result correct — it makes the boundaries easier to specify and the outcomes easier to inspect. It doesn't prove Jev chose the best model, and a successful build isn't caused by the selector. Also, "local Laya" describes the selector — your coding worker may still call a hosted model, and so can a locally running ACP client. Trace each hop before describing the whole workflow as local or private.
03 / What We Built: keel 0.2.0
The build is keel 0.2.0, a local-first Mac coding app:
- Base: an Avid-derived coding workspace
- Language: Rust
- Interface framework: GPUI
- Platform: Apple Silicon, macOS 15+
The detail matters because there are two kinds of work in the story: the application — a real coding workspace with sessions, provider connections, tools, and local decision modes — and the Jev engineering research and evaluation material. They're related, but not the same deliverable. The work started from a working coding environment (tasks, repositories, provider sessions, tools), giving it a practical center: how decisions behave when attached to tasks, repositories, provider sessions, and tool execution.
The Design Process: Keep the Choices Legible
The useful design question was simple: where can Jev make a choice that the application can still check? That question both locates where it belongs and explains stopping where the host couldn't enforce the result. Five steps:
- Identify the decision boundary. New-task routing and a bounded next-step focus were both plausible; provider-internal actions weren't, because the host doesn't own those loops.
- Make the options explicit. The host knows which providers and models are installed, enabled, and valid. The selector receives candidates; it doesn't make up candidates.
- Preserve user intent. A pinned route stays pinned; a running session isn't quietly rerouted. Fresh unpinned work is the narrow place where automatic route choice belongs.
- Build the ordinary path. Local Laya is the default; Jev is a deliberate opt-in; Normal mode remains for the existing route behavior.
- Give failure a shape. Badly formed output, invalid ids, old task details, and abstention each get a clear outcome: reject, fall back, or ask the user. No one needs to pretend every model response is usable.
Conservative? Sure.
I'd take a narrow choice I can debug over an impressive promise I can't trace through the code.
Figure 7: From task arrival to execution, Jev appears only between steps 4 and 6.
Where Jev Sits in the Request Path
The route path in plain language, eight steps:
- A new task arrives without an explicit pinned route;
- The host reads current provider and model availability;
- It constructs a finite list of eligible candidates;
- The selected decision mode calls local Laya or hosted Jev, if enabled;
- The selector returns a candidate id or abstains;
- The host checks that the choice remains eligible and current;
- The application starts the route or falls back according to policy;
- A decision record captures what happened.
Pinned routes and existing sessions bypass the automatic choice — that's not a footnote; it's a central product rule. The step selector has its own checks: the host offers focus ids and candidate tool sets, the selector returns a focus, the host prepares the bundle against current tool schemas, and checks the named tool again at dispatch. Same principle; different lifecycle.
04 / What "Self-Improving" Should Mean
The Shipped App Records Decisions; It Doesn't Train Itself
This is the line to keep clear: keel records selector activity and outcomes that can support evaluation, but it doesn't automatically train Laya from chat history, and it doesn't silently rewrite its own policy after a failed run. The release doesn't prove a self-improving agent in the sci-fi sense.
What it gives is the start of a feedback loop: a receipt can become a scenario; a scenario can be replayed against a changed prompt, policy, or model; a human can review the difference and decide whether to keep the change. Less cinematic, yes — but workable: a changed rule, an old result, and a new result. "The agent learned something" leaves you guessing.
Turn a Failure into a Case You Can Run Again
Suppose a route selector chooses a provider that became unavailable a moment later. Don't just say "Jev got confused" — save the details: the offered options and task state, the chosen id, the host's check result, the fallback used. Then create a replayable scenario: reproduce the same candidate set and task flags → run the baseline selector → run the candidate change under the same conditions → compare valid selection, abstention, fallback, latency, and downstream result → have a person review whether the change improved the intended behavior. For tool-focus choices the case differs: the model selects a bundle whose schema changed before dispatch — check that rebuilding the tool set filters out the old tool, and that the host rejects it if requested anyway. The value of a structured record is that it gives you the conditions to reconstruct the decision. "The agent seemed confused" gives you a mood.
Compare Against the Ordinary Harness: Credit Assignment
A passed task doesn't tell us which component deserves credit — maybe the worker would have solved it on the original route. Keep the task set and scoring rules fixed; changes to the worker, prompt, repository, or reviewer affect results too. Use a separate task set for the final comparison — don't tune on every case, then call the familiar ones proof. The awkward run can reveal a missing rule that clean demos hide. Record failures, fallback rate, total time, and review effort. The author is explicit: this is a proposed evaluation design; no measured coding-quality gains are claimed.
Keep a Human Gate on Production Changes
An improvement loop can propose changes to: candidate construction rules, selector prompts or typed schemas, fallback policy, the local model or hosted route, and the tool bundle attached to a focus choice. For each change: keep the old version, replay the relevant scenarios, inspect regressions, then let a human approve the version that becomes the new baseline. No silent self-training; no "it changed because it learned" with no diff and no rollback path.
Where to Start
Start with one choice the host can check. Keep six rules in mind:
- Jev can choose among host-prepared options; it doesn't own the host;
- Laya and Jev are separate modes, not two names for one model;
- Provider-owned inner loops stay provider-owned;
- Selection doesn't grant permission;
- Records support a future evaluation loop; they don't mean the app trains itself;
- The public Jev engineering repo contains the framework and examples; keel is the application build discussed here.
If you're building a coding harness, write down four things:
- One decision it's allowed to make;
- The exact choices it can see;
- What happens when it abstains;
- How you'll know whether the outcome helped.
Then build the narrowest path that proves those rules.
Start small enough that a bad choice has nowhere to hide. A better agent needs a system that can remember what went wrong.
That's the likable part of this kind of engineering: the model can be clever, and the surrounding system can still be clear.
Source:X / The Avid Builder's Guidehttps://x.com/Av1dlive/article/2102802621664985241