Agent harness: what it is, and why it matters more than the model

How context, tools, memory, permissions, verification, and observability turn a model into a system that can act inside a product.

person at a laptop beside a sheet of drawn layers, the harness between the model and the product

An agent harness is the program around an AI model that lets it work as an agent. The model understands the situation, thinks, and decides what to do. The harness gives it the context, the tools, a memory of what it already did, and the controls to act on real systems.

Put simply:

Agent = model + harness.

That is why the same model can behave very differently depending on the product it sits in. What changes is what it is told, which actions it can take, the permissions, what it remembers, and the rule that says the job is finished.

Choosing a good model is only one part. A large share of how the agent behaves depends on what we build around it. We cover that in the model is not the agent.

The loop that turns a model into an agent

A model on its own, given some text, returns an answer. An agent has to go further: take an action, see what happened, and decide the next step.

In its simplest form, the harness repeats this loop:

while (!state.finished) {
  const response = await model.generate({
    messages: state.messages,
    tools: availableTools
  });

  if (!response.toolCall) {
    return response.content;
  }

  const result = await executeTool(response.toolCall);

  state.messages.push(response, {
    role: "tool",
    content: JSON.stringify(result)
  });
}
while (!state.finished) {
  const response = await model.generate({
    messages: state.messages,
    tools: availableTools
  });

  if (!response.toolCall) {
    return response.content;
  }

  const result = await executeTool(response.toolCall);

  state.messages.push(response, {
    role: "tool",
    content: JSON.stringify(result)
  });
}
while (!state.finished) {
  const response = await model.generate({
    messages: state.messages,
    tools: availableTools
  });

  if (!response.toolCall) {
    return response.content;
  }

  const result = await executeTool(response.toolCall);

  state.messages.push(response, {
    role: "tool",
    content: JSON.stringify(result)
  });
}

The model chooses the next step. The code around it decides which actions exist, how they run, what result goes back to the model, and when to stop.

In a real product, that loop also has context, the state of this task, permissions, a check of the result, and a record of what happened. That is where the harness work starts.

LangChain describes it in The anatomy of an agent harness as the infrastructure around the model so it can repeat steps and use tools. Microsoft uses a similar idea in its Agent Framework: tools, memory, state, and control of the run.

The pieces of a harness

What information the model receives

The model can only think with what is in front of it on each question. The harness decides what goes in: the instructions, earlier messages, the result of actions, documents, files, or where the task stands.

This is sometimes called context engineering. It is not the same as the harness. The first chooses what information is needed at that moment. The second is the whole system the model works inside.

The difference matters when the task is long. We cannot keep every message and every result forever, so we have to decide what to keep, what to summarize, and what to look up again later.

Summarizing is not deleting whatever looks unimportant on its own. A small detail can be needed many steps later to understand why the agent decided something.

We made that point in the piece on Jev AI. Scoring each message or each action on its own, to decide what to keep, looks efficient. But it mixes two questions: whether that piece looks useful by itself, and what information the agent needs in order to keep thinking well. The second depends on everything that has happened in the task, not on one isolated piece.

How it goes from thinking to doing

Tools are the actions the model can ask for: look something up, read a file, send an email, run a program, or work in a closed space.

An action can be shown to the model with a contract, a text that says what it is called and what data it needs:

const getOrder = {
  name: "get_order",
  description: "Returns an order for this customer",
  parameters: {
    type: "object",
    properties: {
      order_id: { type: "string" }
    },
    required: ["order_id"]
  }
};
const getOrder = {
  name: "get_order",
  description: "Returns an order for this customer",
  parameters: {
    type: "object",
    properties: {
      order_id: { type: "string" }
    },
    required: ["order_id"]
  }
};
const getOrder = {
  name: "get_order",
  description: "Returns an order for this customer",
  parameters: {
    type: "object",
    properties: {
      order_id: { type: "string" }
    },
    required: ["order_id"]
  }
};

The model can decide to call get_order. That decision, on its own, changes nothing. The harness checks that the data is valid, adds who the user is, and runs the function on the right system.

The split matters: the model proposes an action, and the system decides how it runs.

When the agent can do many things, you also have to decide which ones it knows about at each moment. Loading hundreds of definitions from the start fills up what the model can read, and makes the choice harder.

In a real Devic run, the agent needed a capability that was not loaded at the start and used discover_tools to find it during the task itself.

Asking for actions is not only defining functions. It is also deciding how they are discovered, when they are loaded, where they run, and with which keys.

What to remember now, and what to remember next time

If the task has several steps, you need to know what has happened: which actions were used, what failed, whether someone has to approve something, or what should be retried. That is the state of this task.

Memory is something else. It keeps what can still be useful in later tasks: preferences, earlier decisions, or knowledge about the project.

Separating them answers two questions: what the agent needs in order to continue this task, and what it should remember when the next one starts.

It also decides what stays in front of the model, what is stored outside, and what can be dropped when the task ends.

The model proposes, the system authorizes

As soon as an agent can change files, run commands, or touch real systems, we cannot assume that every idea from the model runs on its own.

Some rules are fixed. Reading information can be allowed. Deleting data, making a payment, or touching passwords can require approval, or be forbidden.

Others depend on what the action means. If the agent has a tool that runs commands, git status and rm -rf ./data go through the same tool, and the risk is completely different.

One option is to put an evaluator in the middle:

const evaluation = await commandEvaluator.evaluate({
  command: toolCall.arguments.command
});

if (evaluation.risk === "low") {
  return execute(toolCall);
}

if (evaluation.risk === "medium") {
  return requestHumanApproval(toolCall);
}

throw new PolicyViolation(evaluation.reason);
const evaluation = await commandEvaluator.evaluate({
  command: toolCall.arguments.command
});

if (evaluation.risk === "low") {
  return execute(toolCall);
}

if (evaluation.risk === "medium") {
  return requestHumanApproval(toolCall);
}

throw new PolicyViolation(evaluation.reason);
const evaluation = await commandEvaluator.evaluate({
  command: toolCall.arguments.command
});

if (evaluation.risk === "low") {
  return execute(toolCall);
}

if (evaluation.risk === "medium") {
  return requestHumanApproval(toolCall);
}

throw new PolicyViolation(evaluation.reason);

This is one of the things we are working on at Devic: look at certain commands before they run, to catch the ones that can break something.

That evaluator does not replace the fixed rules. Operating-system permissions, isolation of the space where the code runs, API permissions, and access to keys still have to live in code and infrastructure. The evaluator adds a reading of the meaning, for when a fixed rule is not enough.

The same applies when a person has to approve. In real Devic runs, the agent can stop and ask for an explicit yes before it changes something.

In a SaaS it also matters who is who: who started the task, which customer it acts on, and which permissions it has. The model can propose the action. The final authorization still belongs to the system. We go into that in identity, permissions, and limits.

An action that does not fail is not a finished job

A function that ends without an error does not mean the agent achieved what it wanted. A good harness looks at the result and checks it before continuing.

await template.update(change);

const result = await template.capture();

const verification = await verifier.check({
  expected: change,
  actual: result
});

if (!verification.ok) {
  await agent.retry(verification.reason);
}
await template.update(change);

const result = await template.capture();

const verification = await verifier.check({
  expected: change,
  actual: result
});

if (!verification.ok) {
  await agent.retry(verification.reason);
}
await template.update(change);

const result = await template.capture();

const verification = await verifier.check({
  expected: change,
  actual: result
});

if (!verification.ok) {
  await agent.retry(verification.reason);
}

An agent that writes code can change a file and run the tests. One that updates a CRM can read the record back. One that works on a screen can look at a later screenshot.

This also happens in real Devic runs. After changing a template, the agent captures the page again to check the result. When an outside image did not load, it saw that the change had not come out as expected and did not mark the task as done.

An action that finishes well confirms that the function ran. The check answers something more important: whether the job was actually solved.

Being able to reconstruct what happened

When a task has several questions to the model and several actions, keeping only the final answer is not enough. You need to reconstruct what the model decided, which action it used, what came back, how long it took, and what it cost.

In Devic runs, each step can note the action, the time, the input tokens and the ones that came from cache, the output, and the cost.

That shows whether the problem was in the reasoning, in an action, in the check, or because the agent started filling up with too much context. The record does not make the agent more correct, but it lets you understand it, fix it, and operate it. We explain that in traceability.

How it looks together

Imagine a customer writing to an ecommerce SaaS: "The order arrived broken. I want my money back."

The harness loads that customer's context and where the conversation stands. The model decides to look up the order with an action and, after seeing the result, looks up the returns policy. It finds that the case allows a refund, but that anything over 50 euros needs approval.

The model proposes the refund. The harness checks the permissions and the business rule, pauses the task, and asks a person to confirm. Once approved, it runs the refund, reads the system again to see that it was processed, and writes the whole path into the record.

The model did the thinking. The context, the actions, the state, the permissions, the check, and the record were handled by the harness.

That is the jump from a demo that calls a system to an agent that can work inside a product. We also cover it in a demo is not an agent in production.

Where your effort should go

The case, the business rules, the actions that belong to your product, and the data are yours. That is usually where what sets you apart lives.

The loop, the state, the retries, the approvals, the record, the closed space where the code runs, separating customers, or being able to change models are problems that repeat across almost every product. For a team of a few dozen people, building and maintaining all of that inside can eat a lot of resources without an equivalent advantage.

The question is not only whether you can build your own harness. It is which part of that infrastructure should really be yours. If the advantage is knowing the business, the integrations, and the experience of the person using the product, the effort should go there.

Building that infrastructure or relying on a platform is covered in build vs buy.

The model matters. So does the harness

As agents get more capable, the name of the model explains less on its own. The result also depends on what it is told, which actions it can use, what it remembers, which permissions it acts under, and how it checks its work.

The model supplies the ability to think. The harness turns that ability into controlled actions on real systems.

So when you put agents in a product, the question is no longer only which model to use. You also have to decide what system it needs around it in order to do the job well.

That system is the agent harness.

Alberto Iglesias

CEO

Turn your SaaS AI-Native

Get a free trial just by signing up


Ready to turn your platform AI-Native?

Centralize agents, tools, and flows in one platform and start scaling with less friction.

Build and run agents without friction

Connect your current infrastructure

  • Devic AI

  • Devic AI

  • Devic AI

Ready to turn your platform AI-Native?

Centralize agents, tools, and flows in one platform and start scaling with less friction.

Build and run agents without friction

Connect your current infrastructure

  • Devic AI

  • Devic AI

  • Devic AI