Agent Engineering

Research and field notes on Agent Harness, runtime design, reliability, orchestration, context, and evaluation.

Agent Engineering

An Engineering Guide to LLM Token and Cost Optimization

A practical cost-optimization guide for agents, skills, MCP, AI coding, and model APIs, covering task economics, context, caching, tools, model selection, and stop conditions.

LLM Engineering · Agent Efficiency · Context Engineering · LLM Cost · Model Routing · 13 min read ·

Read Essay

Agent Engineering

The Bumps and Bruises of Building Multi-Agent Systems

A field guide to seven recurring Multi-Agent failure modes: task decomposition, hidden dependencies, duplicate work, ownership, context handoffs, correlated errors, and debate.

Multi Agent · Agent Orchestration · Agent Reliability · Context Engineering · Agent Evaluation · 14 min read ·

Read Essay

Agent Engineering

A Personal Skill Map for the AI Era: Building a Complete Loop in 2026

An AI learning roadmap for 2026 is more than a tool list. Explore six durable skills--discovery, specification, modeling, orchestration, verification, and delivery--for turning ideas into real outcomes.

AI Learning · Agent Workflow · Education AI · Career Skills · Personal Productivity · 14 min read ·

Read Essay

Agent Engineering

Can AI Beat CAPTCHAs in 2026? From Recognition to Risk Scoring

How well can AI solve CAPTCHAs in 2026? A practical analysis of text, image grids, sliders, GUI agents, end-to-end pass rates, and modern bot risk systems.

Agent Engineering · GUI Agent · AI Security · Computer Vision · CAPTCHA · 23 min read ·

Read Essay

Agent Engineering

From Generation to Verification: Google ScientistOne's Chain-of-Evidence Framework

Explore Google ScientistOne's Chain-of-Evidence (CoE) framework for verifiable AI research, from claim classification and traceability to integrity audits.

Agent Engineering · Autonomous Research · Agent Evaluation · Verifiable AI · ScientistOne · 12 min read ·

Read Essay

Agent Engineering

The Real-World Challenges of Loop Engineering—and Why I’m Skeptical

A critical analysis of Loop Engineering through completion criteria, proxy metrics, error amplification, context corruption, engineering productivity, and agent security.

Agent Harness · Coding Agent · Agent Reliability · Agent Loop · Loop Engineering · 15 min read ·

Read Essay

Agent Engineering

What Is Loopcraft? From Prompt Engineering to Agent Loop System Design

A practical guide to Loopcraft: why AI agent engineering is moving from one-off prompt optimization toward event triggers, verification loops, retries, state, and continuously improving agent systems.

Agent Harness · Coding Agent · Agent Orchestration · Agent Loop · Prompt Engineering · Loopcraft · 15 min read ·

Read Essay

Agent Engineering

Field Notes: How Agentic RAG Handles the Real Mess of Enterprise Data

Starting from a real customer-support ticket, this piece breaks down the five stages of Agentic RAG -- orchestration, query rewriting, parallel retrieval, sufficiency checking, and synthesis -- and covers cross-system permission boundaries plus three engineering decisions on routing, iteration depth, and latency.

Agent Engineering · Agent Orchestration · Retrieval · Enterprise AI · Agentic RAG · 15 min read ·

Read Essay

Agent Engineering

Claude Code Incident Review: What Anthropic's Three Production Bugs Teach Agent Engineers

A practical review of three Anthropic Claude Code production incidents, covering reasoning history, system prompts, cache invalidation, default reasoning settings, and real-world AI agent testing.

Coding Agent · Agent Harness · Agent Reliability · Context Engineering · Claude Code · 10 min read ·

Read Essay

Agent Engineering

Lessons from LangChain: Designing a Reliable Runtime for Production-Grade Agents

A practical breakdown of LangChain's Runtime for production deep agents — durable execution, layered state, human-in-the-loop, permissions, middleware, streaming, observability, and a concrete design checklist for business-grade Agent Runtimes.

Agent Harness · Agent Runtime · Agent Reliability · Agent Observability · LangChain · 28 min read ·

Read Essay

Agent Engineering

Anthropic Managed Agents: 2026 Agent Harness Architecture for Production AI Agents

A practical breakdown of Anthropic Managed Agents, covering agent harness architecture, sessions, sandboxes, credentials, context builders, traces, evals, and production AI agent runtime patterns.

Agent Harness · Agent Runtime · Agent Reliability · Anthropic · Managed Agents · 18 min read ·

Read Essay

Agent Engineering

Agentic Harness Engineering (AHE): How Coding Agent Harnesses Evolve Automatically

A practical deep dive into Agentic Harness Engineering, coding agent harness design, agent observability, middleware, memory, evaluation, and rollback for AI engineering teams.

Agent Harness · Coding Agent · Agent Evaluation · Agent Observability · AHE · 23 min read ·

Read Essay

Agent Engineering

DSPy Tutorial: Why Signatures Are Easier to Optimize Than Raw Prompts

A practical DSPy tutorial for AI builders in North America. Learn how DSPy Signatures, Modules, and Optimizers improve prompt engineering, prompt optimization, and LLM pipeline design for production AI workflows.

LLM Engineering · Prompt Engineering · LLM Evaluation · DSPy · LLM Pipeline · 11 min read ·

Read Essay

Agent Engineering

PaperBanana Explained: How Multi-Agent AI Generates Academic Method Diagrams

A technical breakdown of PaperBanana's Retriever, Planner, Stylist, Visualizer, and Critic agents, showing how multi-agent AI generates academic method diagrams for research writing and scientific figure workflows.

Multi Agent · Agent Orchestration · Academic Figures · Agent Workflow · PaperBanana · 9 min read ·

Read Essay

Agent Engineering

How Use Multi-Agent Collaboration for Academic Figure Generation?

A focused breakdown of how PaperBanana coordinates Retriever, Planner, Stylist, Visualizer, and Critic agents to generate academic method diagrams.

Multi Agent · Agent Orchestration · Academic Figures · Role Design · PaperBanana · 6 min read ·

Read Essay