An AI Agent Walks into Production
A practical talk on AI agent failure modes in production: explicit state, idempotent side effects, recovery, evaluation, and the human stop button.
Direct answer
Direct answer: Agents in production
This talk shows engineering audiences why AI agents fail after the demonstration and what to build so they can be operated: explicit state, side effects that cannot silently repeat, replay, evaluation, and a stop path a person controls. Attendees leave with a production checklist they can apply to their own agent the following week. A security-focused variant, “The Agent That Couldn’t Wire Money”, covers least privilege, approvals, and replay across tools.
The question
“Our audience wants a practical talk about real agent failure modes. What will they be able to do differently afterwards?”
Who it is for
Conference programme committees, engineering meetups, platform and SRE events, and internal engineering days.
What a production agent must be able to do
Abstract
What the talk covers
Most AI agents that impress in a demonstration have never met a retry, a crash halfway through a payment, a model upgrade, or two workers changing the same record. In production those are ordinary events, and an agent built as a chat loop handles them badly: it repeats side effects, loses track of what it has done, and leaves operators unable to explain or recover the outcome. This talk treats an agent as a long-running distributed system. It walks through explicit state, durable checkpoints, idempotent and reconcilable effects, versioning of prompts, models, tools and policy, evaluation tied to real state transitions, and the stop conditions that hand control back to a person. Examples come from public open-source systems and published preprints, with the limits of each result stated. Attendees leave with a checklist they can apply to their own agent before its next release.
The audience leaves able to
- →Where agents duplicate or lose side effects, and how idempotency keys and reconciliation close that gap
- →What a useful checkpoint contains, and what a checkpoint alone cannot prove
- →Which versions and decisions to record so an operator can replay and explain an outcome
- →When an agent should stop for a person, and what a real recovery path looks like
Audiences
Formats
- Conference talk
- Keynote
- Technical deep dive
- Webinar
- Podcast
- Workshop variant
What you leave with
Talk abstract, evidence references, audience prerequisites, and a production checklist for attendees.
- 01 An abstract and title tailored to the event’s audience and track
- 02 Short and long speaker biographies from the speaker kit
- 03 A one-page production checklist attendees can keep
- 04 Links to the public systems and papers referenced in the talk
How it runs
Structure and format
- Duration
- 30–45 minutes for a conference talk; 60 minutes with questions; extendable to a hands-on lab
- Delivery
- In person in the UK, travel by agreement, or remote
- Participants
- Engineers, platform teams, SREs, architects, and technical leaders
The demo that worked
Why an agent that succeeds in a demonstration still fails under retries, crashes, concurrent writers, and changing models.
State and side effects
Explicit state, durable checkpoints, idempotency keys, receipts, and reconciliation for effects that may or may not have happened.
Evaluation and observability
Connecting decisions to state transitions, tool effects, versions, cost, and quality so an operator can explain an outcome.
The stop button
When an agent should stop for a person, and how to give operators a real resume, compensate, abandon, or escalate path.
Audience assumptions
- →Familiarity with how LLM-based applications call tools or APIs
- →No specific framework is assumed; examples are portable across orchestration libraries
Evidence
What this draws on
fast-langgraph
Durable agent state and checkpointing, with project-published benchmarks of up to 700× faster checkpoint operations. These are operation-level results, not total application speed.
Inspect the source ↗Before the Pull Request
In a controlled experiment, duplicate or conflicting multi-agent rework fell from 78% to 0% and useful throughput more than tripled. A controlled setup, not a universal productivity forecast.
Inspect the source ↗Azure agent platform delivery
Recent work deploying agentic AI services on Azure AI Foundry and AKS/Kubernetes, with custom sandboxing and MCP in .NET. Customer identity, architecture, and scale remain confidential.
CloseGate
A Python and MCP policy layer with action tiers, segregation of duties, materiality routing, mandatory approval for irreversible actions, and hash-chained replayable audit. No published independent certification or customer deployment.
Inspect the source ↗Scope and limits
What this is not
- —The talk teaches portable architecture and failure patterns; it is not a product pitch or a vendor endorsement.
- —Benchmark figures cited from fast-langgraph are operation-level results, not end-to-end application speed-ups.
- —Customer architectures from delivery work stay confidential; examples use public systems and papers.
If the need is different
Common questions
Answers before you commission
What will attendees be able to do afterwards?+
Identify where their agent can duplicate or lose a side effect, add explicit state and idempotency to that boundary, decide what to log for replay, and define when the agent must stop for a person.
Is the talk tied to a particular framework?+
No. The patterns apply to LangGraph, custom loops, and managed agent platforms. A framework-specific version is possible when the audience uses one stack.
Can it be a workshop instead?+
Yes. The Production Agent Reliability Lab turns the same material into a two-hour, half-day, or full-day hands-on session with a working artefact.
Is there a security-focused version?+
Yes. “The Agent That Couldn’t Wire Money” covers action tiers, approval boundaries, an evidence schema, and adversarial tests for tool-using agents.
Related