Skip to main content
21 min read

Planning and Verification Loops with LangGraph and Jev

Part 4: Plans, independent verification, stuck detection, budgets, and loops that stop when the task is done.

👉

Recap

Part 3[link] built the harness around the agent.

Retry and fallback kept one provider error from ending a run. A model router picked the model, and hard limits capped model and tool calls. Guardrails redacted PII, fenced injected instructions, and screened the request with Jev.

The request classifier as a before_agent gate that runs once at invocation, before the main model exists: user request enters, a small classifier model replies ALLOW or BLOCK; ALLOW continues to the main model and tools, BLOCK jumps straight to end with a decline message and zero main-model calls; label the three safety properties, runs once per run, can only jump to end (no tool call, no edit, no approval), and runs before the main model exists.

A permission matrix enforced roles at the tool boundary, and an approval interrupt paused destructive calls for a human.

The permission matrix drawn as a grid. Rows are roles (viewer, engineer, oncall, admin). Columns are the three write tools named in the printed MATRIX , open_ticket, post_status, rollback_flag , each split into dev, staging, prod. Cells are filled or empty (e.g., viewer empty everywhere; engineer filled for open_ticket in all three environments but only dev/staging for post_status and rollback_flag; oncall and admin filled everywhere). A separate bar across the top labeled "read tools: always allowed" and a bar across the bottom labeled "unknown tools: always denied".

Part 1 opened with a failure that none of this was built to catch. A model asked for services/checkout/flag.py, but the real file is flags.py. The tool reported the error, and the scripted model ignored it and answered anyway. The loop saw no more tool calls and returned the answer as a success.

Everything we built so far let it through, and there was a good reason. No permission was violated, because reading a file is allowed. No limit was hit, because the run took one tool call. The injection guard found nothing to fence.

The Answer schema was satisfied, because evidence is a string and a file path is a string.

The harness governs what each call may do. It does not know what the task was, so it cannot tell whether the task is done. The stop rule from Part 1 still gives that decision to the model.

Here, in this part, the agent's work becomes a LangGraph loop with an explicit plan. The Part 3 agent works one checklist item at a time, with every Part 3 control still in force. A verifier that the model does not control decides whether each item is done.

The loop repairs, replans, or stops, and every way it can end is a named terminal state. By the end, a careless model ends in a stuck status and nothing is written to memory.

Prerequisites: Parts 1 to 3 of this series. You can read them below:

First Steps Towards a Production-Ready Agent with LangChain
Part 1: Tools, structured output, runtime context, and tracing, in one small agent.
Middleware-Driven Context Engineering in LangChain
Part 2: Dynamic prompts, memory, compression, and isolation with middleware
Agent Harness with LangChain Middleware and Jev
Part 3: Retries, fallbacks, routing, guardrails, permissions, and the first human-in-the-loop interrupt.

Let's begin!


Introduction

Every agent so far has used the same loop. Call the model. If it asked for tools, run them and call the model again. If it did not, stop. The stop rule is the last line, and it hands the decision to the model.

That rule has a known weakness. Huang et al. studied intrinsic self-correction. That is a model reviewing its own answer using only its own judgment.

Huang et al., 2024

On reasoning tasks, models struggled to correct themselves without external feedback. Sometimes the attempt made their performance worse. Asking a model whether it is finished relies on the same judgment and carries the same blind spot.

So what we need is a control cycle designed around the agent, so that "done" is established rather than claimed. Here we add things the default loop lacks:

  • A goal the loop can check against, stored as data rather than implied by a prompt.
  • A completion check the model does not control, grounded in tool output.
  • Progress detection, so a loop that repeats itself is noticed.
  • Hard stops, so every run ends in a known state with a report.

Every loop has the same basic parts, but control can sit in different places. A trigger starts it, a goal defines success, and the agent acts and observes. Then something decides whether to continue.

The loop stops because the goal is met or a budget has run out. In the default loop, the model owns that decision. Here, code does.

Two loops side by side. Left, "default loop": model → tools → model, with a single exit labeled "no more tool calls" owned by the model. Right, "verified loop": plan → act → verify → decide, with exits labeled done, stuck, budget exhausted, declined, all owned by code. A thin line from verify to "tool output" labeled "ground truth".
👉
The loop is not a replacement for the harness. The harness decides what each call may do. The loop decides whether the task is done.

Note: Verification costs time and model calls. A verified loop makes more calls than a single agent run, and planning adds a call before any work happens. Thus, add complexity only when it improves outcomes you can measure.

For example, the support agent's answers can trigger a rollback. Preventing a plausible but wrong answer is worth the extra calls.


Note on the reference project:

The code and project setup are attached below as a zip file. Extract it and run uv sync to get going.

Download the zip file below:

Published on Oct 2, 2026