Agentic Harness Explained: How AI Agents Get Work Done
A scientific Python script can look convincing long before it has read a real file. The imports seem reasonable, the plot has a useful title, and the explanation sounds complete. Then the first run finds a missing dependency, a different column name, or a waveform with an unexpected sampling rate.
Working through those surprises requires a sequence: inspect the data, run something, read the result, adjust the approach, and check what changed. An AI agent needs software that supports that sequence. This surrounding software is commonly called an agentic harness, or simply an agent harness.
Think of the harness as the workbench around a capable but fallible analyst. The workbench supplies instruments, a notebook, operating procedures, and checks. The analyst still has to make useful decisions, while the workbench determines which decisions can be carried out and what evidence comes back.
In this post, I explain the main pieces of a harness, follow them through a seismic-data quality-control example, and look at how to tell whether the extra machinery actually helps.
The key idea
A useful agent connects decisions to observable consequences. The model proposes a next step; the harness supplies context, controls execution, records what happened, and brings the result back into the next decision. Reliability depends on that whole system.
What is an agentic harness?
I use harness here for the runtime and supporting mechanisms that let a model pursue a task across multiple steps. The boundary varies between projects: some include the workspace, instructions, and evaluation machinery; others use the word mainly for the execution loop. OpenAI’s description of the Codex agent loop provides a concrete example of the latter, showing how the surrounding application coordinates a user, a model, and tools.
Claude Code provides a parallel example. Anthropic’s How Claude Code works describes a cycle of gathering context, taking action, and verifying results. For a bug fix, Claude Code can inspect the relevant files, make an edit, run tests, and use the output to decide the next step. Its surrounding software supplies the tools and manages the model’s context. Both examples make the same principle concrete: the model proposes the work, and the harness connects those proposals to execution and feedback.
Four terms help separate the pieces:
| Term | What it does | Example |
|---|---|---|
| Model | Produces text or structured action requests from its current input | Proposes a command to inspect a waveform file |
| Agent loop | Repeats model calls, actions, and observations | Reads a test failure, changes the code, and runs the test again |
| Harness | Manages the loop and its context, state, tools, boundaries, and records | Restricts file access and saves progress between sessions |
| Framework or SDK | Supplies reusable components for building the system | A library with a runner, tool interfaces, or checkpoint support |
These categories can overlap in a product. Installing a framework still leaves decisions about your data, permissions, success criteria, and failure handling.
An agent also differs from a fully prescribed workflow in how the next step is chosen. In a workflow, code may always run preprocessing, inference, and plotting in that order. In an agent, the model can choose to inspect a suspicious input before proceeding. Anthropic’s Building Effective Agents makes this distinction and recommends starting with the simplest arrangement that solves the task.
The loop turns an answer into a process
A basic agent loop repeatedly performs five jobs:
- Assemble context. Combine the task, relevant instructions, selected evidence, and current state.
- Ask the model for the next step. This might be a tool call, a question, or a proposed final answer.
- Check the action. Validate its arguments, scope, permissions, and remaining budget before execution.
- Run and observe. Return the actual tool result, including errors, rather than assuming the action succeeded.
- Update state and decide whether to continue. Finish when the required checks pass, or pause when a limit, blocker, or approval requirement is reached.
The ReAct paper is an influential research example of interleaving reasoning and actions so that observations can influence subsequent decisions. A production harness adds operational responsibilities around that interaction, including persistence and access control.
The feedback matters. A command exiting successfully establishes that the command completed under its own rules. It does not establish that every input file was inspected or that the scientific interpretation is correct. The harness needs a task-specific definition of completion.
What the harness has to manage
Context is the working set
For evidence about the current task, the model depends on the information supplied for that call. A repository may contain thousands of files, but the immediate question might need one function, its tests, and a configuration value.
Context engineering is the work of selecting and arranging that useful information. Retrieval, targeted file reads, and summaries can keep a run focused. Summarization can also discard an important constraint, so preserve source references and make critical instructions easy to recover. Anthropic’s context-engineering guide discusses this trade-off between limited context, relevance, and continuity.
For a scientific task, a compact data dictionary can be more useful than pages of raw numbers: units, channel naming, time convention, missing-value rules, and the location of the original files.
State carries the work forward
Context is what the model sees now. Persistent state is what the system saves so the task can continue later: completed steps, unresolved problems, file locations, decisions, and verification results.
Anthropic’s long-running-agent experiments used a feature list, progress records, and version history to help successive sessions resume work. Their reported failure modes included attempting too much at once and declaring completion too early.
A saved note is still a claim about the world. If it says a test passed yesterday, the current files may have changed. Good resumption includes checking the relevant state rather than blindly trusting the summary. For code, record the revision and test command alongside the result; for data, record the input manifest and processing settings.
Tools need clear contracts
A tool should make its purpose, inputs, outputs, and failure modes understandable. Compare a vague tool named process_data with one that inspects a specified archive directory and returns file counts, channel identifiers, sampling rates, and parsing errors.
The SWE-agent paper investigated this interface-design problem for software engineering. Its agent-computer interface was designed around repository navigation, editing, and testing, demonstrating that how a model interacts with its environment can affect task performance.
For our archive example, structured output should distinguish an empty result from a failed inspection. Include the file identifier and units with each measurement. Keep large arrays in files; return enough summary information and a path for further inspection.
Permissions need enforcement outside the prompt
An instruction to preserve raw data is helpful, but a read-only mount gives that instruction an independently enforced boundary. Apply the same principle to network destinations, credentials, and writable directories.
External documents and tool responses can contain misleading instructions. A harness should treat them as evidence with a source, rather than allowing them to grant new authority. Anthropic’s containment analysis describes complementary defenses around the model, environment, and external content, and explains why probabilistic model safeguards cannot stand alone.
Use human approval for consequential steps with a clear description of the action and target. An approval screen complements isolation; it does not replace it. Nor does a sandbox prove that a computation is correct or prevent every mistake inside its permitted workspace.
Recovery needs more than another attempt
A temporary service error may justify a bounded retry. A malformed input may need a different parser or an explicit failure record. A permission denial should become a blocker, not a search for a bypass.
Be especially careful when a tool times out after a write. The operation may already have happened. Idempotency means arranging an operation so repeating it has the same intended effect as performing it once. LangGraph’s Functional API documentation explains why resumable execution needs careful handling of side effects and unfinished tasks.
For a report, use a stable output identifier and inspect existing artifacts before rerunning a write. Set limits on elapsed time, tool calls, and spending, and return a useful partial result when the task cannot finish.
Traces make failures inspectable
An execution trace records the steps needed to understand a run: model and tool calls, returned results, timing, errors, approval events, and produced artifacts. The OpenAI Agents SDK tracing documentation gives one implementation with traces and nested spans.
The useful question is whether the record lets you locate the failure. Did the tool miss a file, did context omit a warning, or did the final summary misread a correct table? Apply redaction and retention rules because traces can also expose sensitive inputs. Observability should preserve necessary evidence without indiscriminately logging secrets or entire private datasets.
A concrete example: checking a seismic archive
Suppose the task is to prepare a quality-control report for a local waveform archive: identify unreadable files, compare expected and observed channels, flag sampling-rate inconsistencies, and summarize gaps and overlaps.
This is an illustrative design, not a reported experiment. No accuracy or speed improvement is implied. The model could help inspect an unfamiliar archive and explain exceptions, while tested Python functions perform the measurements.
I would start by making the task contract explicit:
task: Prepare a waveform archive QC report
inputs:
archive: read-only waveform directory
inventory: expected channels and operating intervals
outputs:
- file inventory with parsing status
- channel coverage and sampling-rate summary
- gap and overlap table
- reviewable report with limitations
constraints:
- preserve raw files
- write only to the report workspace
- no external uploads
- no automatic interpolation or resampling
This contract exposes an important scientific detail: a missing channel can be identified only relative to what was expected during a particular interval. A file listing alone cannot establish that a station should have recorded three components all day.
The harness would then support a sequence like this:
Inspect first. Read the directory structure and available metadata. Build a manifest with one record per input file, including failures. If the expected station inventory is unavailable, mark channel completeness as unresolved rather than inventing the missing reference.
Measure with explicit assumptions. Run a reviewed inspection script. ObsPy’s Stream.get_gaps documentation describes gap and overlap information derived from trace timing. Its result still needs interpretation: an entirely absent channel requires comparison with the inventory, and acquisition boundaries should not automatically be treated as instrument faults.
Investigate exceptions. Let the model choose a focused next step, such as opening the metadata for files with conflicting sample rates. Return parsing errors to the loop with their file paths. Keep the original measurements available so the final explanation can be checked against them.
Verify before summarizing. Test the inspection code on small fixtures with known gaps, overlaps, mixed sampling rates, and unreadable inputs. Reconcile successful and failed inspections against the manifest. Check that report counts match the generated tables and that output paths exist. Protect these acceptance checks from being weakened merely to make the run pass.
Deliver with provenance. Save the script, settings, input manifest, software versions, tables, figures, and unresolved questions together. The report should separate observations from interpretations. A gap is an observation; an explanation involving telemetry loss needs additional evidence.
If this becomes a stable daily procedure, much of it can become an ordinary scheduled pipeline. The agent is most useful where the path genuinely depends on what the inspection finds.
The agent says the archive passed QC, but one input file could not be read. What should the harness do?
How do we know the harness helps?
Evaluate the complete system on representative tasks. Keep the model and task set fixed when comparing harness changes, and repeat runs when variation matters. Measure accepted outcomes, missed requirements, unsafe-action attempts, recovery behavior, cost, and elapsed time. A polished final answer is only one part of the evidence.
Include difficult cases deliberately: stale progress notes, interrupted runs, malformed tool responses, missing inputs, and instructions embedded in untrusted files. Check whether the system stops honestly when it cannot proceed. A run that reports a blocker correctly can be more useful than one that manufactures completion.
Additional agents and review loops also have costs. In Anthropic’s March 2026 harness-design study, a planner, generator, and evaluator helped on demanding application-building tasks, but the design was expensive and required tuning. As the underlying model improved, the author removed some orchestration and reassessed where evaluation still added value. Those are task-specific engineering observations, not a universal ranking of harnesses over models.
OpenAI’s harness-engineering account makes a related point from repository design: useful documentation, executable checks, and accessible logs help agents work with evidence. The broader lesson I take from these examples is to fix the observed failure, then measure again.
A few common questions
Is a harness the same as a system prompt?
A system prompt contributes instructions. The harness also runs tools, carries state, enforces access boundaries, and handles execution outcomes. A paragraph of instructions cannot supply those capabilities by itself.
Do I need a multi-agent system?
Start with one agent and clear tools. Add separate workers when there is a demonstrated benefit from independent investigation, parallel work, or specialized review, and test whether that benefit justifies the coordination cost.
Can a harness make an agent completely reliable?
No. It can make failures easier to prevent, detect, contain, and recover from. Its tests can be incomplete, its environment can be misconfigured, and its model can still make a poor decision.
The practical starting point
Choose one bounded task. Define its inputs and acceptable outputs. Give the model a small, understandable tool set, preserve the evidence it needs, and enforce the permissions the task actually requires. Then inspect the failures before adding more machinery.
For scientific computing, the most valuable outcome is a result someone else can check: which data were used, which operations ran, which tests passed, and what remains uncertain. A well-designed agentic harness makes that chain visible, while leaving the scientific judgment open to scrutiny.
Further reading
- Unrolling the Codex agent loop, Michael Bolin, OpenAI, January 2026.
- How Claude Code works, Anthropic, official documentation.
- Building Effective Agents, Anthropic, December 2024.
- ReAct: Synergizing Reasoning and Acting in Language Models, Yao et al., ICLR 2023.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, Yang et al., 2024.
- Effective harnesses for long-running agents, Anthropic, November 2025.
- Harness design for long-running application development, Prithvi Rajasekaran, Anthropic, March 2026.
Sources and linked documentation checked in October 2026. Product implementations will continue to change; the examples here emphasize the underlying design choices.
Disclaimer of liability
The information provided by the Earth Inversion is made available for educational purposes only.
Whilst we endeavor to keep the information up-to-date and correct. Earth Inversion makes no representations or warranties of any kind, express or implied about the completeness, accuracy, reliability, suitability or availability with respect to the website or the information, products, services or related graphics content on the website for any purpose.
UNDER NO CIRCUMSTANCE SHALL WE HAVE ANY LIABILITY TO YOU FOR ANY LOSS OR DAMAGE OF ANY KIND INCURRED AS A RESULT OF THE USE OF THE SITE OR RELIANCE ON ANY INFORMATION PROVIDED ON THE SITE. ANY RELIANCE YOU PLACED ON SUCH MATERIAL IS THEREFORE STRICTLY AT YOUR OWN RISK.
Leave a comment