Talk

An AI Agent Walks into Production

A practical talk on AI agent failure modes in production: explicit state, idempotent side effects, recovery, evaluation, and the human stop button.

Direct answer

Direct answer: Agents in production

This talk shows engineering audiences why AI agents fail after the demonstration and what to build so they can be operated: explicit state, side effects that cannot silently repeat, replay, evaluation, and a stop path a person controls. Attendees leave with a production checklist they can apply to their own agent the following week. A security-focused variant, “The Agent That Couldn’t Wire Money”, covers least privilege, approvals, and replay across tools.

The question

“Our audience wants a practical talk about real agent failure modes. What will they be able to do differently afterwards?”

Who it is for

Conference programme committees, engineering meetups, platform and SRE events, and internal engineering days.

What a production agent must be able to do

What a production agent must be able to do Explicit state: Goals, inputs, decisions; Safe side effects: Idempotent or reconciled; Durable checkpoint: Resume after a crash; Stop for a person: When authority runs out 01 Explicit state Goals, inputs, decisions 02 Safe side effects Idempotent or reconciled 03 Durable checkpoint Resume after a crash 04 Stop for a person When authority runs out

Abstract

What the talk covers

Most AI agents that impress in a demonstration have never met a retry, a crash halfway through a payment, a model upgrade, or two workers changing the same record. In production those are ordinary events, and an agent built as a chat loop handles them badly: it repeats side effects, loses track of what it has done, and leaves operators unable to explain or recover the outcome. This talk treats an agent as a long-running distributed system. It walks through explicit state, durable checkpoints, idempotent and reconcilable effects, versioning of prompts, models, tools and policy, evaluation tied to real state transitions, and the stop conditions that hand control back to a person. Examples come from public open-source systems and published preprints, with the limits of each result stated. Attendees leave with a checklist they can apply to their own agent before its next release.

The audience leaves able to

  • →Where agents duplicate or lose side effects, and how idempotency keys and reconciliation close that gap
  • →What a useful checkpoint contains, and what a checkpoint alone cannot prove
  • →Which versions and decisions to record so an operator can replay and explain an outcome
  • →When an agent should stop for a person, and what a real recovery path looks like

Audiences

AI engineering Platform SRE Engineering leadership

Formats

  • Conference talk
  • Keynote
  • Technical deep dive
  • Webinar
  • Podcast
  • Workshop variant

What you leave with

Talk abstract, evidence references, audience prerequisites, and a production checklist for attendees.

  • 01 An abstract and title tailored to the event’s audience and track
  • 02 Short and long speaker biographies from the speaker kit
  • 03 A one-page production checklist attendees can keep
  • 04 Links to the public systems and papers referenced in the talk

How it runs

Structure and format

Duration
30–45 minutes for a conference talk; 60 minutes with questions; extendable to a hands-on lab
Delivery
In person in the UK, travel by agreement, or remote
Participants
Engineers, platform teams, SREs, architects, and technical leaders
01

The demo that worked

Why an agent that succeeds in a demonstration still fails under retries, crashes, concurrent writers, and changing models.

02

State and side effects

Explicit state, durable checkpoints, idempotency keys, receipts, and reconciliation for effects that may or may not have happened.

03

Evaluation and observability

Connecting decisions to state transitions, tool effects, versions, cost, and quality so an operator can explain an outcome.

04

The stop button

When an agent should stop for a person, and how to give operators a real resume, compensate, abandon, or escalate path.

Audience assumptions

  • →Familiarity with how LLM-based applications call tools or APIs
  • →No specific framework is assumed; examples are portable across orchestration libraries

Evidence

What this draws on

Public open-source system

fast-langgraph

Durable agent state and checkpointing, with project-published benchmarks of up to 700× faster checkpoint operations. These are operation-level results, not total application speed.

Inspect the source ↗
arXiv preprint

Before the Pull Request

In a controlled experiment, duplicate or conflicting multi-agent rework fell from 78% to 0% and useful throughput more than tripled. A controlled setup, not a universal productivity forecast.

Inspect the source ↗
Delivery experience

Azure agent platform delivery

Recent work deploying agentic AI services on Azure AI Foundry and AKS/Kubernetes, with custom sandboxing and MCP in .NET. Customer identity, architecture, and scale remain confidential.

Public reference system

CloseGate

A Python and MCP policy layer with action tiers, segregation of duties, materiality routing, mandatory approval for irreversible actions, and hash-chained replayable audit. No published independent certification or customer deployment.

Inspect the source ↗

Scope and limits

What this is not

  • —The talk teaches portable architecture and failure patterns; it is not a product pitch or a vendor endorsement.
  • —Benchmark figures cited from fast-langgraph are operation-level results, not end-to-end application speed-ups.
  • —Customer architectures from delivery work stay confidential; examples use public systems and papers.

If the need is different

Common questions

Answers before you commission

What will attendees be able to do afterwards?+

Identify where their agent can duplicate or lose a side effect, add explicit state and idempotency to that boundary, decide what to log for replay, and define when the agent must stop for a person.

Is the talk tied to a particular framework?+

No. The patterns apply to LangGraph, custom loops, and managed agent platforms. A framework-specific version is possible when the audience uses one stack.

Can it be a workshop instead?+

Yes. The Production Agent Reliability Lab turns the same material into a two-hour, half-day, or full-day hands-on session with a working artefact.

Is there a security-focused version?+

Yes. “The Agent That Couldn’t Wire Money” covers action tiers, approval boundaries, an evidence schema, and adversarial tests for tool-using agents.

Related

Next step

Invite Dipankar

A short written brief is enough to establish fit. I reply personally, and say plainly when the work belongs elsewhere or is not worth commissioning.

Talk, panel, or event workshop: what to include

  • →Event name, organiser, and website
  • →Audience: size, roles, and technical depth
  • →Preferred topic or the problem the audience faces
  • →Format and length (keynote, talk, panel, podcast, workshop)
  • →Date, time zone, and location or remote
  • →Fee, travel, and accommodation arrangements
  • →Recording, publication, and reuse terms
  • →AV and lab environment, if a workshop