Planning and Verification Loops with LangGraph and Jev
Part 4: Plans, independent verification, stuck detection, budgets, and loops that stop when the task is done.
97 posts published
Part 4: Plans, independent verification, stuck detection, budgets, and loops that stop when the task is done.
Part 3: Retries, fallbacks, routing, guardrails, permissions, and the first human-in-the-loop interrupt.
Part 2: Dynamic prompts, memory, compression, and isolation with middleware
Part 1: Tools, structured output, runtime context, and tracing, in one small agent.
Part 15: Block-Attention, Cartridges, and what running preloading in production actually involves, with implementations
Part 14: Shrinking a preloaded cache and the problems with those approaches, covered with implementations
Part 13: Reading the corpus once before any query arrives, and the two limitations that decide whether it works in production, with implementations
Part 12: Why prefill dominates RAG latency, and how to reuse KV caches that share no prefix, with implementations
RL Part 13: An exploration of real-world RL case studies.
RL Part 12: From a single judged group to a full multi-step training loop with ART and RULER.
RL Part 11: From verifiable rewards to LLM-as-a-judge.
RL Part 10: Dropping two models from the four-model pipeline, and building rewards you can trust.
Part 9: From human preferences to a trained reward signal, and the four-model PPO pipeline.
RL Part 8: Trust regions, the clipped surrogate, and the workhorse of modern RL.
RL Part 7: Learning the policy directly, from REINFORCE to actor-critic.
RL Part 6: From linear features to neural networks, and the engineering choices that makes deep value-based RL possible.
RL Part 5: From tables to parameterized value functions.
RL Part 4: Learning value functions and policies without a model. Monte Carlo methods, TD(0), SARSA, Q-learning, and the bias-variance bridge between them.
RL Part 3: Bellman expectation and optimality equations, policy iteration, value iteration, and why dynamic programming needs a model.
RL Part 2: Markov decision processes, returns, policies, and value functions.
RL Part 1: Agents, environments, rewards, and why RL is different from supervised learning.
Diffusion LLMs Part 2: How dLLMs scale to 100B parameters, the inference stack that makes them fast, hands-on code, and when to actually use them.
Diffusion LLMs Part 1: Understanding how diffusion language models work from first principles, the math behind masked diffusion, and why they represent a fundamentally different approach to text generation.
An exploration of real-world MLOps and LLMOps case studies, examining the importance of reliable ML and AI engineering and their significance for business outcomes.
LLMOps Part 14: An overview of the fundamentals of LLM serving, including API-based access, inference with vLLM, and practical decisions.
LLMOps Part 13: Exploring the mechanics of LLM inference, from prefill and decode phases to KV caching, batching, and optimization techniques that improve latency and throughput.
LLMOps Part 12: Understanding LLM fine-tuning, parameter-efficient methods like LoRA and QLoRA, and alignment techniques such as RLHF, DPO, and GRPO.
LLMOps Part 11: Understanding evaluation of conversational LLM systems, tool evaluations, tracing with Langfuse, and automated red teaming.
LLMOps Part 10: Understanding model benchmarks, LLM application evaluation, and tooling.
LLMOps Part 9: A foundational guide to the evaluation of LLM applications, covering challenges and a practical taxonomy of evaluation methods.