Before an AI agent touches customer data: six production gates
An agent can complete a clean demo and still be unready for customer data. Before release, I want answers to six questions: what exact workflow does it run, where can its data go, how was reliability measured, what may it change, how do we recover from a bad action, and who owns it after launch?
A successful demonstration proves that an agent can complete a task. It does not answer what happens when data is missing, who may change what, how failures are detected, or whether the organization can investigate what the agent did.
The Measuring Agents in Production study surveyed 86 practitioners responsible for deployed systems across 26 domains and examined 20 case studies. In that sample, 68% of agents ran at most 10 steps before human intervention, 74% relied primarily on human evaluation, and reliability was the top development challenge. The study reports these observations about its sample. They are useful context for the gates below, not release thresholds.
What follows is a go/no-go framework for preparing an agent to handle real data. Each gate governs one decision, names the evidence to collect, and ends with a pass condition.

Gate 1: Is the workflow bounded?
Before discussing models, define the exact job. A production agent is not a general assistant. It is a system that executes one workflow, with named inputs, named outputs, and a named stopping condition.
Map the workflow before you touch the agent:
- Trigger: what starts it? A user message, a scheduled job, an event in an upstream system.
- Inputs: what data does it receive, in what format, with what permissions attached?
- Tools: what can it call, what does each tool do, and what does each tool cost?
- Decisions: where does the agent decide between options, and where does it follow a fixed rule?
- Outputs: what does it produce, where does it go, and who reads it?
- Stopping condition: when is the job done? How many steps are allowed before it must stop or escalate?
- Reversibility: which actions can be undone, and which cannot?
The MAP study found that 68% of the deployed agents in its survey ran for 10 steps or fewer before a human intervened. That describes current practice. It does not prove that ten steps is right for every workflow.
Pass condition: the workflow can be drawn as a bounded sequence with named exceptions, a maximum step count, and one accountable owner.
Gate 2: Is the data boundary understood?
The first check is the data flow. Map where data enters, where it moves, and where it persists.
Start with a data-flow diagram. Draw every component: the user, the agent, the model, the retrieval system, every tool, every log sink, and every external service. Then trace the data through it.
For each data path, answer four questions:
- Classification: what is the sensitivity of this data? Public, internal, confidential, restricted?
- Access rights: who is allowed to see this data? Does the agent’s identity carry the right scope, and only the right scope?
- Retention: where is this data stored, for how long, and who can delete it? Prompts, responses, tool-call arguments, traces, and evaluation datasets all count.
- Exposure surface: can this data appear in logs, in vector stores used by other agents, in evaluation sets reviewed by vendors, or in model training data?
Indirect prompt injection is one reason this map matters. A retrieved document, email, or API response can contain instructions that try to change the agent’s behavior. The OWASP LLM Top 10 and the AI Agent Security Cheat Sheet describe this attack class and the need to treat retrieved content as untrusted input.
The boundary also informs the deployment choice. A public API, a private cloud service, and a model running on company-owned machines have different data paths, costs, latency, and model options. Choose among them after the data classification and access requirements are known.
Pass condition: every data transfer, storage location, and access decision is documented in a diagram someone outside the build team can read.
Gate 3: Can reliability be measured before release?
One accuracy number cannot describe production reliability.
A production evaluation set should cover the full distribution of situations the agent will meet in production, not only the easy cases:
- Normal tasks: the common path.
- Ambiguous requests: inputs that could reasonably mean two things.
- Missing and conflicting information: data that is incomplete or contradictory.
- Tool failures: timeouts, rate limits, malformed responses.
- Permission violations: cases where the agent should refuse.
- Prompt injection and malicious documents: adversarial inputs designed to change the agent’s behavior.
- Multi-step compounding errors: cases where a small early mistake cascades.
- Cases requiring refusal or escalation: the agent must recognize what it cannot do.
The numbers below come from Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems. The paper evaluates six agents on 300 enterprise tasks. It reports that optimizing accuracy alone produced agents 4.4 to 10.8 times more expensive than cost-aware alternatives with comparable performance. Its reliability check fell from 60% on a single run to 25% consistency across eight runs of the same task. Those are study results, not promises for a new system.
The paper proposes CLEAR, which covers Cost, Latency, Efficacy, Assurance, and Reliability. Its expert study used 15 evaluators and found a correlation of 0.83 between CLEAR scores and production success, compared with 0.41 for accuracy-only evaluation. Measure cost, latency, policy behavior, and run-to-run consistency alongside accuracy, then define thresholds for the workflow you actually operate.
The MAP survey found that 74% of the deployed systems it studied still relied primarily on human evaluation. Human review catches cases that automated checks miss, but it needs a versioned rubric and a path to turn findings into regression tests.
Pass condition: the team has a versioned evaluation set covering all eight categories above, defined release thresholds for each metric, and regression testing that runs on every prompt, model, or tool change.
Gate 4: Does the agent have limited authority?
An agent with broad permissions can hide the real authority boundary during a demo. The boundary needs to be explicit before release.
Least privilege starts with identity. Treat the agent as a principal with a purpose, an owner, and an identity managed through its lifecycle. The Microsoft guidance on least privilege for AI agents recommends a dedicated principal, task-based roles, approved tools, and end-to-end records of what happened under which authority.
Keep permissions tied to the task and resource. Separate evidence gathering from remediation. Let the agent read before it can write. Use time-limited role activation, short-lived tokens, or per-action approvals when a workflow needs more access than its baseline. Downstream services should check identity and scope again rather than trusting the orchestrator.
Expose a curated tool list. High-impact operations such as delete, export, send, transfer, or deploy should require approval for the exact action. The OWASP AI Agent Security Cheat Sheet describes tool abuse and excessive agency as risks that need controls outside the prompt.
A practical authority ladder for any agent:

- Answer questions from data it can read.
- Recommend an action a human reviews.
- Prepare an action for human approval.
- Execute reversible actions autonomously.
- Execute consequential actions only with step-up approval.
Use the levels as an evidence sequence. Move upward only when the level below has passed its evaluation threshold.
Pass condition: the agent has a dedicated identity, task-scoped access, an allowlisted tool set, and an authority level that matches its evaluation evidence.
Gate 5: Can failures be detected and reversed?
Agent failures can be semantic. A tool may return HTTP 200 while containing the wrong record. Retrieval may return plausible context that does not support the answer. A step may succeed while the overall task moves in the wrong direction.
The operational layer needs four things.
Tracing. Record enough data to reconstruct a run: model and prompt versions, retrieval and tool versions, inputs, outputs, decisions, identity, effective scope, and the resulting state change. Anthropic describes full production tracing as the way its team diagnosed agents that were not finding obvious information in its multi-agent research system.
Outcome monitoring. Count success, partial completion, failure, and escalation as well as latency and token use. Alert on repeated retries, unusual resources, unexpected tool sequences, and cost spikes.
Kill switches and rollback. Test a way to disable the agent identity and invalidate its credentials. Test rollback or compensating actions for the state changes the workflow can make.
Incident response. Keep a procedure with an on-call owner, investigation steps, communication responsibilities, and a decision about rollback or redeployment.
Anthropic reports that its agents use about four times as many tokens as chat interactions and its multi-agent systems about 15 times as many. That is a measurement from its research system, not a universal multiplier. It does show why an unbounded loop needs a budget. The AI Agent Security Cheat Sheet also describes Denial of Wallet as a risk.
Pass condition: a failed action can be detected, explained, contained, and recovered from within a defined time window, using evidence the system already collects.
Gate 6: Is there an owner and operating model?
Name the people and routines that keep the system useful after launch.
Prompts can change in effectiveness as models change. Evaluation sets go stale, tools get deprecated, and new failure modes appear under load. An operating model gives those changes an owner.
Define, before launch:
- Business owner: the person accountable for the outcome the agent produces.
- Technical owner: the person accountable for the agent’s behavior, performance, and maintenance.
- Evaluation maintenance: who updates the evaluation set, how often, and against what new cases.
- Change approval: how prompt changes, model upgrades, and tool additions are reviewed and approved.
- Runbook: the documented procedures for common incidents, including the kill switch from Gate 5.
- Cost tracking: cost per successful task completed, not cost per token. Token cost hides behind task complexity.
- Support expectations: what users can expect, what they cannot, and where they escalate.
- Success criteria for scaling, redesigning, or stopping: the conditions under which the agent gets more scope, gets rebuilt, or gets retired.
Documentation needs its own owner and update path. A runbook that predates the last prompt or tool change is incomplete.
Pass condition: a named owner exists for performance, risk, cost, and maintenance, and that person can answer “what does this agent do, what does it cost, and what happens if it breaks” without asking the original builder.
The production-readiness scorecard
Use this worksheet before launch and after material changes. The status cells are intentionally blank for the team to fill in with its own evidence.
| Gate | Evidence required | Status |
|---|---|---|
| 1. Bounded workflow | Workflow map, step limit, exceptions, owner | |
| 2. Data boundary | Data-flow diagram, access map, retention policy | |
| 3. Evaluation | Versioned evaluation set, metrics, thresholds, regression run | |
| 4. Authority | Dedicated identity, task-scoped access, tool allowlist, authority level | |
| 5. Recovery | Tracing, outcome monitoring, tested kill switch, incident procedure | |
| 6. Ownership | Named owners, runbook, cost-per-task tracking, change approval |
Passing these gates does not prove that an agent is safe in every situation. It gives the launch decision an evidence trail and makes the remaining uncertainty visible. That is the standard I want before an agent can change data that a person may have to repair.
Work with Nazmi
Build your AI system with Nazmi.
Tell us what you are building, what exists today, and where your team needs help.
Start a conversation or book a 20-minute call →