Agent Harness with LangChain Middleware and Jev
Part 3: Retries, fallbacks, routing, guardrails, permissions, and the first human-in-the-loop interrupt.
Recap
Part 2 gave the support agent a context engine. Each call assembled its prompt from the tenant, environment, and role.

The middleware removed stale tool results from the model's view, moved large outputs to a store, and summarized long runs. It also kept memories inside tenant namespaces.
The next failures sat outside context management:
- The prompt told a viewer not to suggest write actions, but no code blocked an attempted call of a possible write tool.
- We described tool results as data, but an injected instruction still reached the model.
- One transient provider error could still crash the run.
In each failure above, the context was correct. The model had the intended instructions and data. The missing code sat around the model call, where the application decides what may happen regardless of what the model requests. LangChain calls that code the harness and this part builds it.
Prerequisites: Parts 1 and 2 of this series. You can read them below:


Let's begin!
What a harness is
Part 1 used a simple definition, an agent is a model calling tools in a loop, and the harness is the code around that loop. Now that the surrounding code is the subject, the definition needs more precision.
So what we need to understand is, the harness connects the model to its environment, data, memory, and tools. The model-and-tool loop is familiar. The surrounding controls vary by application.
Hence, harness engineering is engineering the execution environment and policies surrounding an agent loop.
create_agent already provides a harness. Even without custom middleware, it has a prompt, tool registry, fixed loop shape, and structured-output handling. This part replaces its implicit defaults with components chosen for specific failures.
Middleware (contd.)
Part 2 used some of the hooks for context management. In this part, we'll explore the complete execution model.
The six hooks
The agent loop has two steps, a model call and a tool call, plus a start and an end. Middleware exposes hooks at each.
| Hook | Runs | Kind | Used in this part for |
|---|---|---|---|
before_agent |
Once, at invocation | Node | Model-based request classifier |
before_model |
Before each model call | Node | PII redaction on input, model-call limit |
wrap_model_call |
Around each model call | Wrapper | Router, retry, fallback |
wrap_tool_call |
Around each tool call | Wrapper | Permission check, injection guard, audit, Jev risk gate |
after_model |
After each model reply, before tools run | Node | Approval interrupt, tool-call limit, output guard |
after_agent |
Once, at completion | Node | Memory write (from Part 2) |
Hooks of the same kind run in list order for before_* and reverse list order for after_*. Wrapper hooks nest with the first item on the list as the outermost wrapper. So a before_model hook always runs before every wrap_model_call hook, regardless of their positions in the middleware list.
Composition
Two ordering relationships matter most here.
Retry and fallback both wrap the model call. With [fallback, retry], fallback is outermost. We call the primary model up to three times. If every attempt fails, the exception reaches fallback, which calls the second model.

after_model, where it can pause after seeing the requested call but before the tools node runs. Permission uses wrap_tool_call, at the point of execution. Without another check, a human could approve a call that the permission guard later denies. A predicate in the approval configuration prevents that wasted review.Note on reference project:
The code and project setup are attached below as a zip file. You can extract it and run uv sync to get going.
Download the zip file below:

