probearc
Tech blog · Essay · 8 October 2026

AI agent harnesses demystified: what they are, how one works, and who offers them

Every agent is a model plus a harness. One real run step by step, the vendor and open-source landscape, and the evidence that the harness moves results on its own.

Over the past several months we have been bombarded with the word "Harness". Internet is flooded with tutorials and short youtube videos explaining what it is. I am writing this articles for those who might be still little confused about it.

Anthropic uses "agentic harness" in its documentation. OpenAI wrote a post called "Harness engineering". Amazon offers "AgentCore Harness", and Microsoft has a "Harness Agent". xAI says Grok 4.7 was trained to understand the Grok Bot harness natively.

One thing to settle before anything else: this applies to every AI agent, not just coding agents. A customer support agent, a research agent, a flight booking agent, a personal agent like OpenClaw and a coding agent like Claude Code are all built the same way. Each one is a model plus a harness. The model is the same kind of thing in all of them. What makes a support agent different from a coding agent is the harness: the tools it is given, the rules about what it may do, and the information put in front of the model. The first half of this article uses coding agents for its examples, because coding agent vendors publish the most detailed traces of what their harnesses do, but everything in it carries over to the other kinds.

I'll explain the term through a real run, then go through the companies and projects offering harnesses in October 2026. After that, I'll look at the evidence that the harness affects performance on its own, and at the argument that harnesses are a temporary layer that models will absorb.

I haven't tried the hosted harness products or most of the open-source projects listed here. For those sections, I've read vendor documentation, project repositories and a few engineering blogs. I've flagged claims where the support is weak. Names change monthly in this area. The names here reflect 8 October 2026.

What the term means

Take the model out of an AI agent, any agent, and everything left is the harness. It's a program. When a developer builds an agent, their application is the harness: it calls the model's API, executes the tools the model requests and keeps state between calls. Running Claude Code, Codex CLI or OpenClaw means using a harness built by someone else.

Only the model thinks. The harness controls the information sent to it. It checks whether requested actions are permitted, carries them out and decides when to end the loop. It also chooses what to discard when memory runs out.

It has five recurring parts:

  1. A loop. The program sends a request to the model and reads the response. It carries out the model's requests, sends the results back and repeats until the task is finished.
  2. A set of tools. These are the actions available to the model. Examples include file reads, shell commands and web searches. Each tool comes with a description for the model.
  3. Gates. Before executing an action, code checks whether it can proceed. Permission rules, sandboxes and hooks belong here.
  4. Context management. This prepares each model request from the system prompt, project instruction files, memory and conversation history. It cuts or summarises the conversation when it grows too long.
  5. Stopping rules and records. These set the conditions for ending a run, including turn and dollar limits. A transcript saved to disk allows the session to resume.

It helps to separate a harness from three related things. Frameworks such as LangGraph, Google's ADK and Microsoft's Agent Framework are libraries used to build harnesses. The harness is the running program around one model. Prompt engineering deals with one input to that program; the harness handles much more. Fine-tuning changes model weights, while changing a harness leaves those weights alone. Vendors are beginning to blur that last distinction. I'll return to it later.

Designing this surrounding layer is called "harness engineering". The phrase became common in early 2026, after OpenAI published a post with that title in February and other writers picked it up.

Before 2026, people called the same thing a "scaffold" or an "agent-computer interface". "Harness" comes from older test-harness terminology. The term was popularised for agents, rather than coined by one person.

Following one run

A run makes the division of work easier to see. I've chosen Claude Code because Anthropic's SDK documentation includes a real turn-by-turn trace. It specifies both the permission order and the condition that ends the run. Other harnesses perform the same jobs in different places. I'll compare two of them afterwards.

The quoted material comes from Anthropic's documentation as it stood on 8 October 2026. The docs explicitly use "Claude" for the model and "Claude Code" for the harness. Below, I use "the model" and "the harness", paraphrasing some quoted wording to keep that distinction clear. This describes documented behaviour. It isn't a source-code walkthrough, since the Claude Code binary is closed source.

The developer enters one request in the terminal, asking for the failing tests in auth.ts to be fixed. A few minutes later, the agent reports that the auth bug is fixed and all three tests pass. Between those two messages, the work proceeds like this.

  1. The harness prepares the model's input. It puts together the developer's request, system prompt, tool definitions and conversation history before calling the model. It includes material the developer didn't enter in the terminal. The project's CLAUDE.md or AGENTS.md instruction file is added again with every request. Skill descriptions go in too, though their full text loads "only when invoked". The input includes the first 200 lines of the agent's memory file. It also includes "system reminders", such as notice that a previously read file has changed on disk. Done by the harness.
  2. The harness supplies the available tools. Each tool arrives with a schema. The list includes Read, Edit and Write for files; Glob and Grep for searches; Bash for commands; WebSearch and WebFetch; and orchestration tools including Agent and AskUserQuestion. Done by the harness.
  3. The model chooses its next move. It considers the current state and returns text, one or more tool requests, or a combination of the two. For turn 1 of this run, it requests npm test through the Bash tool. Done by the model.
  4. The harness checks the requested action. Nothing runs until the request has passed six checks, applied in a fixed order. First, hooks: small scripts the user has configured, which can reject the call outright or let it continue. Second, deny rules: a list of actions that are never allowed, such as a rule against git push. A matching deny rule blocks the tool even in bypassPermissions mode. Third, ask rules: a list of actions that always need the user's confirmation. Fourth, the permission mode the user chose for the session, such as the default mode that asks before most commands, a plan mode that allows no changes at all, or an auto mode. Fifth, allow rules: a list of actions that are pre-approved, and the harness also pre-approves file reads inside the project and a built-in set of read-only commands without any rule. Sixth, if none of the above has decided, the harness asks the program it is embedded in, which in the terminal means showing the user a prompt. In this run, under the default mode, npm test reaches that sixth step and the developer is asked. The newer auto mode replaces the prompt with a separate second model that classifies shell commands and network requests. Done by the harness, with a second model involved in auto mode.
  5. The harness runs the tools and returns their output. It executes the requested tools, gathers each set of results and sends them to the model for another decision. Here, the model receives a message containing test output with three failures. Read-only tools may execute in parallel. Tools that change state, including Edit, Write and Bash, execute sequentially so they don't conflict. The harness takes a file snapshot before editing, allowing the user to rewind. Done by the harness.
  6. Another pass through the loop follows. On turn 2, the model requests reads of auth.ts and auth.test.ts. On turn 3, it requests an edit to auth.ts and another run of npm test. This time, all three tests pass. Every tool call repeats steps 4 and 5. During a larger task, the harness also compacts the conversation if it becomes too long. Older tool output is removed first. If that isn't enough, the conversation is summarised. Architectural decisions, unresolved bugs and implementation details are retained; redundant tool output is discarded. The model chooses; the harness checks, executes and compacts.
  7. The harness ends the run and returns the result. The ending condition is a mechanical one. The loop continues while the model requests tools and receives their output. It stops when the model replies without a tool call. This run finishes with a text-only response, making four turns in total: three containing tool calls and one containing only the final text. Along with that text, the harness returns token usage, cost and a session ID. There are hard limits for turns and dollar spending, although both default to unlimited. The documentation points out that a two-turn limit would have ended this run before the edit. Each message is also saved in a plain text file on disk, allowing the session to resume. Done by the harness.

There is no thinking in steps 1, 2, 4, 5 or 7. Those steps are plumbing. At each call, the model answers one question: what should happen next, given this input? A team can change the run by adding a hook to block rm -rf, a deny rule for git push, or an instruction in CLAUDE.md to run the linter before saying the work is done. None of those changes touches a model weight. That work is harness engineering.

The run above went well. The parts of a harness that matter in production are the ones that handle a run that does not. Here is what the same harness does when things go wrong, again from Anthropic's documentation.

When a tool fails or is refused. A failed command, a file that does not exist or a denied permission does not stop the loop. The error text is sent back to the model as the tool result, the same way the three test failures were in step 5. The docs say that when a tool is denied, the model "receives a rejection message as the tool result and typically attempts a different approach or reports that it couldn't proceed". A hook can also add context to a failed tool's output before the model sees it. The model is expected to recover, and the harness only guarantees that it hears about the problem.

When the model makes a mistake. The harness snapshots every file before the model edits it, so the user can rewind to any earlier state. That covers files only. The docs are explicit that "Actions that affect remote systems (databases, APIs, deployments) can't be checkpointed", and that the only protection for those is the permission gate in step 4. The harness also gives the model the means to check its own work (the test runner it used in turns 1 and 3) but does not force it to use them. A team that wants the check enforced adds a Stop hook: a script that runs when the model says it is finished, and if the tests fail, sends it back to work with the failure as feedback. The docs' own example of a Stop hook is exactly this case. During the run, the user can interrupt at any time with Esc, or type a correction that the model reads before its next step.

When the run will not end. Nothing in the model guarantees that it stops. The harness offers two caps: a maximum number of tool-use turns and a maximum spend in dollars, which also covers any subagents the run spawned. Both are unlimited by default, and the docs call a budget "a good default for production agents". When a cap is hit, the harness ends the run and returns a result marked as cut short rather than as finished, so the calling program can tell the difference and, if it wants, resume the session with a higher limit.

When the context fills up. Everything accumulates in the model's context window: the conversation, every file read, every command's output. When it nears the limit, the harness compacts. It clears older tool outputs first, then summarises the conversation. The docs warn that "detailed instructions from early in the conversation may be lost", which is why persistent rules belong in CLAUDE.md, re-injected on every request, and not in the opening prompt. A team can tell the compactor what to keep, and a PreCompact hook can save the full transcript before it is summarised. If one file or command output is so large that the context refills right after each summary, the harness stops compacting after a few attempts and reports an error instead of looping.

Memory across sessions. A new session starts with an empty context. What survives is on disk: the project's CLAUDE.md, an auto-memory file where the model saves things it learned (the first 200 lines load at each start), and the full transcript of every session, which is what allows a session to be resumed or forked. Subagents are the other memory tool. Each one starts with a fresh context, does its work, and returns only a summary, so a large exploration costs the main conversation a paragraph rather than the full transcript.

None of these mechanisms makes the model smarter. They make a model's mistakes recoverable, its runs bounded and its working memory manageable, which is most of what separates a demo from something you can leave running.

Consider two other ways to build it. OpenAI's Codex CLI places the gate elsewhere. The operating system enforces its three sandbox modes: read-only, workspace-write and full access. It uses Seatbelt on macOS, Bubblewrap or Landlock on Linux, and a native sandbox on Windows. Network access is disabled by default. Commands that need permissions beyond the sandbox enter the approval flow.

At the small end is mini-SWE-agent, from the SWE-bench team at Princeton and Stanford. It's about 100 lines of Python, has bash as its only tool and keeps a completely linear history. Even so, it scores >74% on the SWE-bench verified benchmark. It implements steps 1, 3, 5 and 7, without a permission layer or compaction. Reduce a harness to its minimum and you have one tool and a loop.

What changes when the model stays the same

Harness engineering became a discipline largely because harness changes alone can shift results by double digits on the two benchmarks used to rate coding agents, with the model fixed throughout. SWE-bench Verified gives the agent a real GitHub repository and an issue and checks whether its patch passes the project's tests. Terminal-Bench gives it tasks to complete in a Linux terminal, such as building software or fixing a broken environment, and checks the result automatically.

  • In its January 2025 SWE-bench post, Anthropic reported moving the same Claude 3.5 Sonnet from 33% to 49% on SWE-bench Verified through scaffold changes alone. The best scaffold was intentionally small, consisting of a prompt, a bash tool and a file editing tool. Anthropic's point was that scaffolding can substantially change an agent's SWE-bench performance even when the underlying AI model is unchanged.
  • The February 2025 Claude 3.7 Sonnet announcement gave two results: 62.3% using the standard scaffold and 70.3% using a custom one. The custom scaffold tried solutions in parallel, rejected patches that failed visible tests and used a scoring model to rank the remaining patches.
  • LangChain kept gpt-5.2-codex fixed in a February 2026 Terminal-Bench 2.0 experiment. Harness middleware alone raised its score from 52.8% to 66.5%. The changes added a verification checklist before completion, injected environment context, detected loops of repeated edits and prompted the model about testing and time budgets. Reasoning effort produced a revealing result: the highest setting scored 53.9% because it timed out, while the setting one level below reached 63.6%.
  • Anthropic's November 2025 Opus 4.5 announcement also reported gains from the hosting environment. Keeping both the models and the Terminus-2 harness unchanged, a better environment took Gemini 3 to 56.7% and GPT-5.1 to 48.6%. Both results exceeded the figures reported by their developers. The machine running the harness is part of what affects performance.

These results change what benchmark scores mean. Frontier vendors in 2026 publish results only with their own harnesses. Terminal-Bench 4.0 separates MODEL and AGENT into different columns, yet each frontier entry still matches a model to its maker's harness. Claude models run with Claude Code, GPT models with Codex, and Grok with Grok Build.

Scale AI takes the opposite approach on its public SWE-bench Pro board. Its leading entries run under mini-SWE-agent with a 50-turn limit, while other entries receive 250 turns. A reported "model" score in 2026 therefore measures the model-plus-harness product. There is no public board that lets you disentangle their contributions.

The two kinds of harness engineering

Harness engineering happens at two levels, and it helps to keep them apart.

The builder harness is the harness itself: the program that Anthropic, OpenAI, Google or an open-source project ships. Engineering at this level means deciding how the loop runs, which tools exist and how they are described, in what order permission checks happen, how the conversation is compacted when it gets long, and when the run stops. If you write your own agent, this is the code you write. If you use Claude Code or Codex CLI, someone else did this work, and you choose it when you choose the product.

The user harness is everything a team adds around a shipped harness without touching the model or the product's code. Instruction files such as AGENTS.md tell the agent how the project works. Hooks block or allow specific actions. Linters and tests give the agent fast, automatic feedback on its own work. Documentation and CI make the repository legible to it. Engineering at this level means shaping the environment so the agent gets things right more often.

OpenAI's harness engineering post is an example of the second kind. The team's work was a short instruction file, a documentation directory the agent could read, and custom linters whose error messages told the agent how to fix a violation. None of it changed the model.

The practical point is that the user harness is the part you can start on tomorrow, whichever vendor's agent you use.

The vendors offering harnesses

As of October 2026, every major vendor sells three layers. There's a local agent to run yourself, a programmable SDK and a hosted "managed harness". In the hosted version, the vendor operates the loop and sandbox. The entries below all come from the vendors' own pages.

Vendor Agent you run locally SDK or framework Hosted offering Source availability
Anthropic Claude Code, available through terminal, IDE, desktop, web, Slack, Chrome and mobile Claude Agent SDK, providing the tools, agent loop and context management used by Claude Code Claude Managed Agents, in beta, supplies a pre-built, configurable agent harness on managed infrastructure Closed. The SDK wraps the closed Claude Code binary
OpenAI Codex CLI; the Codex app is also inside ChatGPT desktop OpenAI Agents SDK Agents API, in beta, operates the Codex harness and handles the agent infrastructure underneath it Apache-2.0 for Codex CLI, SDK and skills; the IDE extension and cloud remain closed
Google Antigravity CLI and IDE. For individual tiers, Antigravity CLI replaced Gemini CLI on June 18th, 2026 ADK, available in Python, TypeScript, Go, Java and Kotlin Managed Agents in the Gemini API, in public preview since 19 May 2026, provides a configurable agent harness in Google-hosted sandboxes ADK and Gemini CLI are open; I found no licence for Antigravity CLI
Microsoft and GitHub GitHub Copilot CLI, with local sandboxing GA on 7 October and an autopilot mode; also the Copilot cloud agent Microsoft Agent Framework includes a packaged "Harness Agent" Foundry hosted agents provides per-session VM-isolated sandboxes and accepts your own framework MIT for Agent Framework; closed for Copilot and Foundry
Amazon Kiro Strands Agents AgentCore Harness provides a managed agent loop to define and invoke AI agents through a single API call. It accepts any model, including OpenAI and Gemini, with no separate harness charge Open for Strands; closed for Kiro and AgentCore
Cloudflare I found none Agents SDK, which includes Project Think, described by Cloudflare as an opinionated harness that handles the full chat lifecycle: agentic loop, message persistence, streaming, tool execution, stream resumption and extensions Runs harnesses on Durable Objects and Sandboxes rather than selling its own loop: Pi's harness as a durable "Lifecycle capability" since 2 October 2026, and Claude Managed Agents with the loop at Anthropic and the sandbox on Cloudflare since 19 May 2026 MIT for the Agents SDK; Sandboxes and Workers are a service
xAI Grok Build, in beta since May 2026 I found none Grok Bot, available since August 2026, offers durable AI teammates working on a persistent cloud computer No statement found
Mistral Mistral Vibe, an open-source CLI I verified none When I checked the hosted agents documentation, it returned 404 Apache-2.0 for Vibe
Meta I found none I found none I found none Muse Spark 1.1 is sold as a model with zero-shot generalisation to new native tools, MCP servers and custom skills. It has no harness of its own

I see three patterns here. Hosted harnesses have become the product category of 2026. Anthropic, OpenAI, Google, Amazon and Microsoft now sell operation of the loop and sandbox as a service. Anthropic and OpenAI also have a mode that leaves compute and data with you while they operate the loop.

Model owners tend to package the model and harness together. Cloud platforms, meanwhile, sell a choice of models. Amazon explicitly supports OpenAI and Gemini models in its harness. Microsoft's Harness Agent asks you to supply a chat client.

Vendors are also increasingly hosting one another's products. AgentCore names the OpenAI and Claude SDKs among its supported frameworks. Codex added Bedrock support in June. GitHub Copilot added Grok, Gemini, Claude and GPT models within days of their releases. Cloudflare has gone furthest in this direction. It ships a harness of its own, Project Think, but its stated strategy is to be the layer underneath any harness, and it already runs Pi's, Anthropic's and others on its platform. Cloudflare's own definition is as plain as any: a harness is the agentic loop that calls tools, reads results, manages context and keeps going until the task is done.

One vendor explicitly says it trained a model on its harness. In the Grok 4.7 announcement, xAI says that training taught the model to understand the Grok Bot harness natively. LangChain describes Claude Code as post-trained with models and harnesses in the loop. That's LangChain making a claim about Anthropic. I couldn't find confirmation on an Anthropic page.

Open-source agents and frameworks

There are two groups on the open-source side: complete agents ready to install, and frameworks for building agents yourself. I obtained the star counts below through the GitHub API on 8 October 2026.

Project What it provides Stars Licence
OpenClaw A complete personal agent and chat-app gateway. Models and agent harnesses, including Claude, Codex and local models, can be swapped as plugins 391.6k MIT
Hermes Agent (Nous Research) A complete personal agent with seven execution backends and a built-in learning loop that turns experience into skills 252.0k MIT
OpenCode A complete coding agent for CLI, desktop and web, with modes for building and planning 212.2k MIT
Pi (Mario Zechner) A minimal coding agent using four tools and a prompt under 1,000 tokens. Subagents and permission popups are deliberately absent 113.3k MIT
OpenHands A coding agent that also runs Claude Code, Codex and Gemini through ACP, assigning each conversation a Docker sandbox 90.2k MIT
Cline A coding agent offering an SDK plus 15 lifecycle hooks 70.0k Apache-2.0
CrewAI A framework to build teams of role-playing agents 59.4k MIT
Goose (Agentic AI Foundation, from Block) A complete agent offering 70+ MCP extensions 55.0k Apache-2.0
Aider A pair programmer in the terminal, with a repo map and auto commits but no sandbox. Its last push was May 2026 49.4k Apache-2.0
LangGraph A framework at a low level for building stateful agents 42.9k MIT
Deep Agents (LangChain) An agent harness described as batteries-included and inspired by Claude Code 30.0k MIT
smolagents (Hugging Face) A framework whose agents express their actions in code 29.7k Apache-2.0
mini-SWE-agent A research harness in 100 lines, with bash as its sole tool 8.3k MIT

The list shows two changes. Activity has shifted away from IDE extensions towards terminal harnesses with SDK modes and always-on personal agents accessed through chat apps. Roo Code was archived in May 2026, and Aider has slowed. Together, OpenClaw and Hermes have more stars than all the coding harnesses combined. Stars measure attention, not quality or actual use.

Open harnesses are also increasingly becoming hosts for vendor harnesses. OpenHands runs Claude Code, Codex and Gemini within its own agent. Goose accepts logins through existing Claude, ChatGPT or Gemini subscriptions.

Their feature sets have almost entirely converged. Every complete harness now offers compaction, skills, MCP and permission modes. They also provide subagents, except Pi, whose author gives a written reason for omitting them: the user can't see what the sub-agent is doing, creating a black box inside another black box.

The four standards they share

Three conventions for files or protocols, plus one editor protocol, now underpin interoperability.

  • MCP, or Model Context Protocol, provides the tool layer. Every vendor's agent documentation I read mentions it, as does every open harness in the table. Its specification requires hosts to get explicit user consent before invoking any tool.
  • Agent Skills, using the SKILL.md format, provides the procedures layer. Each skill is a folder containing a SKILL.md file with a name, description and instructions. Loading happens in stages, so the full instructions enter context only when required. Anthropic started the format, which is now an open standard. On 8 October, agentskills.io listed more than forty adopting clients. They included Codex, Gemini CLI, Copilot, Cursor, OpenCode, OpenHands, Goose, Pi, Hermes and OpenClaw. The hosted harnesses from Microsoft and Amazon also support it.
  • AGENTS.md supplies repository instructions. Described as a README for agents, it is now stewarded by the Agentic AI Foundation under the Linux Foundation. Claude Code can read it instead of CLAUDE.md. Antigravity reads it together with GEMINI.md, while Pi loads it through the directory hierarchy.
  • ACP, or Agent Client Protocol, links agents with editors and hosts. Support includes OpenHands, Goose, Grok Build and Mistral Vibe.

Google's A2A protocol, for communication between agents, only partly fits this pattern. It belongs under the Linux Foundation. Its steering committee includes AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow. Microsoft and Amazon support it on their hosted platforms, but it doesn't appear in Anthropic's or OpenAI's agent documentation. Most coding harnesses use subagents within the same process, so they have little need for a wire protocol between agents.

Hooks haven't converged either. Cline provides fifteen named events. Claude Code implements hooks as shell commands. There is no common format.

How much evidence supports the techniques?

Practitioners tend to recommend the same few harness techniques. The support for those recommendations varies widely. Some have controlled measurements behind them; others have only anecdotes.

Fewer, better tools has the strongest evidence at a small scale. Vercel replaced about 80% of its agent's tools with a single bash-and-filesystem tool. It reports that task time dropped from 274.8 to 77.4 seconds, token use fell 37%, and steps fell 42%. The sample contained five tasks, a limitation the post states openly. Pi's author illustrates the cost of going the other way. Tool descriptions for the Playwright MCP server alone take 13.7k tokens per session, consuming 7-9% of the context window before work begins.

Context management is supported by a strong finding about what goes wrong. Chroma tested 18 models in its "context rot" study. Reliability declined as inputs grew longer. Even one distractor reduced performance, and focused prompts outperformed full prompts for every model tested. This gives an empirical basis for the compaction found in every harness. There is still no published head-to-head comparison of compaction strategies.

Repository instruction files are the most widely used technique and have the weakest support. The evidence here actually runs against their use. Researchers at ETH Zurich and LogicStar.ai tested several agents and models in a study accepted at ICLR 2026. Context files such as AGENTS.md did not improve success rates, while costs rose by over 20%. Agents spent more effort exploring without completing more tasks. Files written by people performed somewhat better than generated files. OpenAI's experience with a 100-line AGENTS.md remains an uncontrolled case study, rather than a measurement.

Subagents and todo tools attract disagreement, with neither side offering a benchmark. Anthropic, LangChain and Microsoft consider subagents with separate context a core part of the system. Pi's author argues that to-do lists usually confuse models more than they assist them, and that plans meant to persist should be stored in a file.

Sandboxing and permissions address safety rather than performance. Anthropic says users approve 93% of permission prompts in standard mode. It gives that approval rate as the reason for introducing auto mode with a classifier.

Could harnesses disappear?

The strongest case for a temporary harness goes like this: reinforcement learning on agent loops teaches models to do work that previously needed elaborate scaffolding. Eventually, little or none of that scaffolding remains.

There is evidence for it. The top public SWE-bench Pro entries now use a 100-line agent with bash alone. Vercel suggested that the best agent architecture might be almost no architecture. Pi uses a sub-1,000-token prompt because its author reasons that frontier models already understand what coding agents do. Anthropic made a similar point about its own product in March 2026. It said each harness component represents an assumption about something the model can't do by itself, and those assumptions should be stress tested. One part of its harness had already become unnecessary with Opus 4.6.

Open harnesses show how this happens. Hermes advertises trajectory compression to train next-generation tool-calling models. Every popular harness now doubles as a data pipeline for the next model. xAI presents that process as a feature in its statement about Grok 4.7. Anthropic's dynamic workflows, introduced in June, allow the model to write part of its harness during a run.

The layers being absorbed matter, though. The loop and tool set are shrinking: agents using only bash can compete, and tool lists are getting shorter. The other three layers show no comparable decline.

Context management addresses a property of the models themselves. Every harness in 2026 compacts context. That includes Pi, despite its initial argument against compaction. The safety layer deals with where real failures occur: runaway loops, exposed always-on agents and malicious skills. The repository environment is also outside the model by definition. That's the user harness described in OpenAI's post.

The commercial direction is different too. Every vendor introduced or expanded a hosted harness in 2026. Those services sell sandboxes, persistent sessions, compaction and audit trails. Their product goes well beyond a clever prompt.

My interpretation is that models are absorbing the thinking parts of the harness. Gatekeeping and memory are becoming services. I expect "harness" to remain the name for those services.

What I take from this

Around a model sits a loop, a set of tools, permission gates, context management and a stopping rule. Together, those make the harness. Thinking happens in the model; the surrounding plumbing does the rest. Changes to that plumbing can move benchmark results by ten to sixteen points without altering the model. That explains both the spread of the term and why every vendor now offers one.

For people using agents, the user harness is the part they can act on today. Instructions, hooks, linters and tests affect the agent's behaviour while leaving the model alone. But measure what those additions do. The one controlled study of instruction files found higher costs without improved results.

I used Claude Code for this research, including reading documentation and repositories. That means the Anthropic sections were researched by an Anthropic model reading Anthropic material. I chose Anthropic's product for the example because its public trace is the most detailed. I've tried to apply the same standard to those sections as to everything else. I'd like to hear from people who build or run harnesses and have had a different experience.

Sources

Everything above comes from these pages, read on 8 October 2026 unless a date is given.

Definitions and the worked example

Evidence that the harness moves results

Vendors

Open-source harnesses and frameworks

Standards

Techniques and their evidence

First published on ringarc.ai on 8 October 2026. Comments and corrections: vikas@probearc.ai.