Why Mathematical Problems Are Hard for Agents in Practice
Cases from First Proof, Momus, LeanMarathon, Danus, and Aletheia reveal the central challenges math agents face in strategy selection, dependency management, long proofs, and preserving information from failed attempts.
AI for Math · Math Agent · Multi Agent · Formal Reasoning · Proof Verification · 17 min read ·
An Engineering Guide to LLM Token and Cost Optimization
A practical cost-optimization guide for agents, skills, MCP, AI coding, and model APIs, covering task economics, context, caching, tools, model selection, and stop conditions.
LLM Engineering · Agent Efficiency · Context Engineering · LLM Cost · Model Routing · 13 min read ·
A Personal Skill Map for the AI Era: Building a Complete Loop in 2026
An AI learning roadmap for 2026 is more than a tool list. Explore six durable skills--discovery, specification, modeling, orchestration, verification, and delivery--for turning ideas into real outcomes.
AI Learning · Agent Workflow · Education AI · Career Skills · Personal Productivity · 14 min read ·
Can AI Beat CAPTCHAs in 2026? From Recognition to Risk Scoring
How well can AI solve CAPTCHAs in 2026? A practical analysis of text, image grids, sliders, GUI agents, end-to-end pass rates, and modern bot risk systems.
Agent Engineering · GUI Agent · AI Security · Computer Vision · CAPTCHA · 23 min read ·
From Generation to Verification: Google ScientistOne's Chain-of-Evidence Framework
Explore Google ScientistOne's Chain-of-Evidence (CoE) framework for verifiable AI research, from claim classification and traceability to integrity audits.
Agent Engineering · Autonomous Research · Agent Evaluation · Verifiable AI · ScientistOne · 12 min read ·
The Real-World Challenges of Loop Engineering—and Why I’m Skeptical
A critical analysis of Loop Engineering through completion criteria, proxy metrics, error amplification, context corruption, engineering productivity, and agent security.
What Is Loopcraft? From Prompt Engineering to Agent Loop System Design
A practical guide to Loopcraft: why AI agent engineering is moving from one-off prompt optimization toward event triggers, verification loops, retries, state, and continuously improving agent systems.
Field Notes: How Agentic RAG Handles the Real Mess of Enterprise Data
Starting from a real customer-support ticket, this piece breaks down the five stages of Agentic RAG -- orchestration, query rewriting, parallel retrieval, sufficiency checking, and synthesis -- and covers cross-system permission boundaries plus three engineering decisions on routing, iteration depth, and latency.
Agent Engineering · Agent Orchestration · Retrieval · Enterprise AI · Agentic RAG · 15 min read ·
Claude Code Incident Review: What Anthropic's Three Production Bugs Teach Agent Engineers
A practical review of three Anthropic Claude Code production incidents, covering reasoning history, system prompts, cache invalidation, default reasoning settings, and real-world AI agent testing.
Coding Agent · Agent Harness · Agent Reliability · Context Engineering · Claude Code · 10 min read ·
DeepMind AlphaProof Nexus Explained: 4 System Paradigms for AI Math Research
A breakdown of how DeepMind AlphaProof Nexus combines LLMs, Lean, AlphaProof, evolutionary search, and multi-agent orchestration into a formal proof system for mathematical research.
AI for Math · Formal Reasoning · Multi Agent · AlphaProof · Deepmind · 17 min read ·
Lessons from LangChain: Designing a Reliable Runtime for Production-Grade Agents
A practical breakdown of LangChain's Runtime for production deep agents — durable execution, layered state, human-in-the-loop, permissions, middleware, streaming, observability, and a concrete design checklist for business-grade Agent Runtimes.
Anthropic Managed Agents: 2026 Agent Harness Architecture for Production AI Agents
A practical breakdown of Anthropic Managed Agents, covering agent harness architecture, sessions, sandboxes, credentials, context builders, traces, evals, and production AI agent runtime patterns.
Agentic Harness Engineering (AHE): How Coding Agent Harnesses Evolve Automatically
A practical deep dive into Agentic Harness Engineering, coding agent harness design, agent observability, middleware, memory, evaluation, and rollback for AI engineering teams.
DSPy Tutorial: Why Signatures Are Easier to Optimize Than Raw Prompts
A practical DSPy tutorial for AI builders in North America. Learn how DSPy Signatures, Modules, and Optimizers improve prompt engineering, prompt optimization, and LLM pipeline design for production AI workflows.
PaperBanana Explained: How Multi-Agent AI Generates Academic Method Diagrams
A technical breakdown of PaperBanana's Retriever, Planner, Stylist, Visualizer, and Critic agents, showing how multi-agent AI generates academic method diagrams for research writing and scientific figure workflows.
Multi Agent · Agent Orchestration · Academic Figures · Agent Workflow · PaperBanana · 9 min read ·
AlphaGeometry DSL Guide: Google Geometry DSL, defs.txt Actions, and Predicates
A technical guide to AlphaGeometry DSL covering problem format, defs.txt action definitions, geometric predicates, rules.txt reasoning, and construction workflow for geometry solvers, dataset generation, and AlphaGeometry reproduction.
AI for Math · Geometry Reasoning · Formal Reasoning · Geometry DSL · AlphaGeometry · 11 min read ·
AlphaGeometry2 Deep Dive: How Does Google AI Solve IMO Geometry Problems?
How AlphaGeometry2 combines geometry reasoning, auxiliary-line search, and symbolic deduction to solve IMO geometry problems, with engineering takeaways for math AI builders.
AI for Math · Geometry Reasoning · Formal Reasoning · AlphaGeometry · Deepmind · 11 min read ·