Mayden.AI
← PerspectivesEngineering

Agents that survivecontact with production

A demo agent and a production agent are different animals. The difference is almost entirely in the engineering you don't see.

Arjun NairHead of the Engineering Studio12 May 20269 min read

It takes an afternoon to build an agent that demos well. It takes real engineering to build one that holds up when thousands of people use it on data that changes, against systems that fail, under rules that matter. The gap between those two things is where most agentic AI projects die.

A demo agent and a production agent are different animals, and the difference is almost entirely in the engineering you do not see.

A production agent needs an evaluation harness before it needs more capability. If you cannot measure whether a change made the agent better or worse, you are not engineering — you are guessing. We build evals first, then iterate against them.

It needs to handle failure as a normal condition, not an exception. In production, tools time out, APIs return nonsense, retrieved documents are stale, and the model occasionally produces confident rubbish. A demo assumes the happy path; a real agent assumes the unhappy one, with retries, fallbacks, timeouts, and sane behaviour when a dependency is simply down.

It needs guardrails that are part of the architecture, not bolted on. Tool access, data scope, and action permissions are designed in, reviewed, and audited. The agent can only do what it is allowed to do, and every action leaves a trace.

It takes real engineering to build one that holds up when thousands of people use it on data that changes, against systems that fail, under rules that matter.

It needs a clear boundary between reading and acting. An agent that can draft a reply is a very different risk from one that can send it, move money, or change a record. We separate the two deliberately, gate the consequential actions behind explicit checks, and keep a human in the loop wherever the cost of being wrong is high. Autonomy is earned, capability by capability, as the evidence accumulates.

And it needs to be observable. When something goes wrong in production — and it will — you need to see the full chain of reasoning, retrieval, and action that led there. Without that, you cannot fix it, and you cannot earn the trust required to expand its remit.

Cost and latency are engineering constraints, not afterthoughts. An agent that makes a dozen model calls to answer one question may be delightful in a demo and ruinous at scale. Production work means budgeting tokens, caching what can be cached, choosing the smallest model that clears the bar, and measuring the cost per successful outcome — because that is the number the business actually pays.

Most of this work is invisible to the user and obvious to the engineer. It is the difference between a system that impresses a steering committee and one that quietly handles ten thousand real interactions a day without anyone noticing. The first is a project; the second is an asset.

None of this is glamorous. All of it is the difference between an agent that ships and one that stays a demo.

Written by

Arjun NairHead of the Engineering Studio

Start a conversation
Start a conversation

Let's put your AIinto production.

Tell us where you're stuck. We'll bring senior people and a working plan — not a pitch.

DXBDubaiDubai International Financial Centre
RUHRiyadhRiyadh