An AI Agent Walks into Production
State, side effects, recovery, and the stop button
A practical architecture for explicit state, idempotent effects, replay, evaluation, and human stop conditions.
I speak about the engineering between an impressive demo and a dependable system: authority, state, evaluation, recovery, local execution, and the boundary between a model and the real world.
Formats
Every title can be tailored to the audience and format without weakening its evidence or turning it into a product pitch.
State, side effects, recovery, and the stop button
A practical architecture for explicit state, idempotent effects, replay, evaluation, and human stop conditions.
Enforcing least privilege, approvals, and replay across tools
An action-tier model, approval-boundary pattern, evidence schema, and adversarial test plan.
Testing generated CUDA and Triton with operator-aware oracles
A layered correctness harness and a reproducible way to distinguish plausible output from correct output.
What coordination logs reveal about multi-agent coding
A failure taxonomy and minimal protocol for measuring duplicate and conflicting work before code review.
Making physical AI operable from cloud plan to edge execution
A systems map for capability grounding, review, permissioned execution, telemetry, recovery, and replay.
A polyglot path from cloud API to embedded local inference
A decision framework for local versus hosted inference, resource limits, compatibility, and offline operation.
From occasional prompts to connected, reviewable work
A map for choosing repeatable workflows, connecting tools safely, keeping people in review, and measuring time returned.
What changes when a model can plan, call tools, and retain state
A plain-language architecture connecting goals, models, tools, MCP, orchestration, memory, security, evaluation, and production ownership.
Participants build, break, observe, and improve a bounded system using local or explicitly authorised environments.
Participants add action tiers, human approval, replayable audit, and adversarial policy tests to a tool-using agent.
Participants turn a fragile agent loop into explicit state, durable checkpoints, idempotent effects, replay, and stop conditions.
Participants expose duplicate work, add task claims and append-only events, reconcile divergent replicas, and mine the log.
Participants choose one recurring task, map its inputs and review points, build an AI-assisted workflow, and define a useful time-and-quality measure.
Dipankar Sarkar is the founder and principal consultant at Neul Labs, a fractional AI CTO, and a hands-on applied AI engineer with 18+ years of experience taking ambitious systems from research and architecture into production.
His recent work focuses on governed tool-using agents, durable orchestration, evaluation, local inference, and AI-generated GPU-kernel correctness. Earlier, he architected regulated banking AI at Aveni and was part of its team in the first FCA Supercharged Sandbox cohort, built RobotGPT at Orangewood Labs, and worked on NLP, recommendation, computer vision, and trust-and-safety systems serving more than 100M users at Hike. He is the author of AI for Everyday Automation and Nginx 1 Web Server Implementation Cookbook, publishes the open GenAI and Agentic AI Playbooks, has published applied ML research, and holds an M.S. in Computer Science (Cybersecurity) from Arizona State University and a B.Tech. from IIT Delhi.
Previous recorded appearance
Democratizing Quick Commerce with KiranaPro — Network Capital ↗A founder conversation on first-principles problem solving, Indian digital public infrastructure, and building KiranaPro.
Dipankar speaks about production AI agents, runtime governance, agent security, durable execution and recovery, AI-generated code correctness, multi-agent software engineering, local inference, physical AI and robotics, cloud operations, and blockchain infrastructure.
Remote podcasts, webinars, panels, fireside conversations, conference talks, technical deep dives, and hands-on workshops. In-person events are considered from a base in St Andrews, Scotland.
Yes. Public projects and delivery experience provide evidence, while the talk teaches portable architecture, failure modes, and operating patterns. A platform-specific version is possible when an audience needs one.
Ordinary event or podcast recording and retention are welcome, subject to reviewing the organiser’s specific presenter and content terms.
Yes. Workshops can run for 90–120 minutes or expand to a half or full day after the lab, environment, participant prerequisites, and material rights are agreed.
Bring the audience problem
Include the audience, format, date, location or remote setup, and the outcome you want attendees to leave with.
Invite Dipankar