
What a useful talk about production AI should cover: architecture, bounded autonomy, data readiness, evaluation, and the human decisions that hold it all together.
Most enterprise AI diagrams put a foundation model in the center and arrange everything else around it.
That is backwards.
The center of an enterprise AI system is the control plane: the people, policies, tests, data contracts, and deterministic software that decide when a model may act. The model is one component inside that system.
This distinction matters because a compelling demo and a dependable product optimize for different things. A demo asks, “Can the model do this once?” A production system asks, “Can we make the outcome useful, observable, affordable, and safe thousands of times?”
Start With the Least Autonomous Design That Works
The golden-hammer problem arrived quickly in AI. Once teams see an agent use tools and solve a multistep task, every process starts to look like an agent problem.
Most are not.
If the steps are known in advance, use a workflow. If a value can be calculated, calculate it. If a policy can be expressed as code, enforce it in code. Use a model where interpretation, synthesis, or adaptation is actually required.
A useful progression is:
- Deterministic software for rules, calculations, validation, and state transitions.
- A single model call for classification, extraction, rewriting, or summarization.
- A model inside a workflow when language reasoning is one step in a known process.
- A bounded agent when the correct sequence of steps cannot be specified beforehand.
- Multiple agents only when delegation creates a measurable advantage that simpler designs cannot provide.
Every step down that list adds nondeterminism, latency, cost, and a larger failure surface. Autonomy should therefore be earned by the problem, not selected because it makes the architecture diagram more exciting.
“Everything should be made as simple as possible, but not simpler.” — often attributed to Albert Einstein; this wording is a paraphrase of his 1933 lecture
For AI systems, “not simpler” is the important half. Some tasks genuinely require planning and tool selection. But the burden of proof belongs to the more autonomous design.

Put Determinism Around Nondeterminism
Language models are probabilistic components. Enterprise systems still need predictable behavior.
The answer is not to pretend the model is deterministic. It is to surround it with deterministic controls:
- Validate inputs before they reach the model.
- Request structured output and validate it against a schema.
- Keep authorization outside the prompt.
- Give tools narrow, typed interfaces.
- Make consequential writes idempotent where possible.
- Require approval for irreversible or high-impact actions.
- Set limits for time, tokens, tool calls, and retries.
- Record the prompt, model version, retrieved context, tool activity, and result.
- Define a fallback when confidence or validation fails.
This is defense in depth applied to AI. A prompt may tell an agent not to issue a refund above a threshold; the refund API must still enforce that threshold. Retrieved text may contain instructions; the tool layer must still reject unauthorized actions. The model can propose. The system decides what is permitted.
The human belongs in this design too, but “human in the loop” should not mean routing every output to a tired reviewer. Human attention is scarce. Reserve it for ambiguity, exceptions, high-impact decisions, and the creation of better policies from observed failures.
The person is not merely in the middle of a flowchart. People define the boundaries of acceptable behavior and decide when those boundaries need to change.

Treat Tools as Capabilities, Not Convenience Functions
Tool use turns generated text into real-world effects. That is the point at which an AI feature becomes an operational and security concern.
Model Context Protocol can standardize how models discover and invoke tools, but a standard interface does not make a capability safe. The server still needs conventional security engineering:
- least-privilege credentials;
- explicit read and write separation;
- user-scoped authorization;
- input validation and output filtering;
- audit records tied to an identity;
- rate and spend limits;
- approval gates for destructive operations.
Do not expose execute_sql(query) when the business need is get_open_orders(customer_id). Do not expose a general file system when the task needs one approved document collection. A narrow tool is easier to authorize, test, observe, and explain.
This also limits the damage from prompt injection. If untrusted content convinces a model to misuse a tool, the tool should still be incapable of exceeding its assigned authority.
“Programs must be written for people to read, and only incidentally for machines to execute.” — Harold Abelson and Gerald Jay Sussman, Structure and Interpretation of Computer Programs
The same principle applies to AI capabilities. A tool contract should make its purpose, inputs, effects, and limits obvious to the engineers and auditors who must reason about it.

Data Readiness Is Product Work
“Prepare the data” sounds like a preprocessing task. In practice, it is a product and governance program.
Retrieval quality depends on more than embeddings. Teams need to answer ordinary but difficult questions:
- Which source is authoritative?
- Who owns it?
- How fresh must it be?
- Which users may see each document or field?
- How are deleted and superseded records handled?
- Can an answer cite the exact evidence used?
- What happens when the sources disagree?
Chunking and indexing matter, but provenance, permissions, versioning, and lifecycle management matter more. A beautifully tuned retrieval system can still return an obsolete policy to the wrong employee.
The minimum useful retrieval record should preserve source identity, access-control metadata, timestamps, document version, and enough location information to produce a citation. Retrieval should enforce permissions before context is sent to the model, not ask the model to hide information afterward.
When an answer affects money, rights, safety, or compliance, the user should be able to inspect its evidence. Fluent text is not evidence.

Token Efficiency Is User Experience
Token efficiency is often presented as cost control. It is also latency, readability, and trust.
A simple question that produces a two-page answer makes the system feel less intelligent, not more. Long prompts also crowd the context window with low-value material and make the model’s attention harder to direct.
Good systems use an information budget:
- retrieve only relevant passages;
- summarize old interaction history;
- store large tool results outside the active context;
- give the model references instead of repeatedly copying raw data;
- cap response length based on task type;
- route simple tasks to smaller models;
- stop execution when the answer is already sufficient.
The cheapest model that reliably meets the quality target is usually the right model. “Reliably” prevents this from becoming a race to the lowest price. A cheaper model that causes retries, escalations, or incorrect actions is not cheaper at the system level.
Measure cost per successful task, not cost per token. Include retrieval, tool calls, retries, human review, and failure handling. The unit that matters to the business is a resolved case, a completed analysis, or an approved change.

Evaluation Must Resemble the Real Job
AI ROI is hard to estimate before a team has operational evidence. That does not justify shipping on intuition. It means experimentation must be designed to create knowledge.
There are two kinds of knowledge in an AI project:
Expert knowledge comes from domain specialists, security engineers, legal teams, operators, and users. It defines what good looks like and which failures are unacceptable.
Experimental knowledge comes from trials: which model works on actual cases, where retrieval fails, how users phrase requests, and which guardrails create friction.
Neither is sufficient alone. Experts can specify the wrong workflow if they never observe real use. Experiments can optimize the wrong metric if experts do not define the stakes.
Begin with a small evaluation set drawn from representative work. Include ordinary cases, edge cases, adversarial inputs, missing data, conflicting sources, and actions the system must refuse. For each case, record the expected properties of a good result rather than demanding one exact sentence.
Track several dimensions separately:
- task success;
- factual grounding and citation quality;
- policy compliance;
- tool-call correctness;
- latency;
- cost per successful task;
- escalation and override rates;
- severity of failures.
An average score can hide a catastrophic edge case. Segment results by task type, user group, data source, and risk level.
“In God we trust; all others must bring data.” — W. Edwards Deming
The necessary addition is: bring representative data. An unrealistic benchmark produces confidence without knowledge.
Build a Learning Loop, Not a One-Time Launch
An enterprise AI stack is not finished when the first model reaches production. Models change, data changes, prompts change, tools change, and user behavior changes.
The architecture needs a feedback loop:
- Observe production outcomes with appropriate privacy controls.
- Capture failures, overrides, and user corrections.
- Turn those cases into regression evaluations.
- Change one component at a time where possible.
- Compare quality, risk, latency, and cost before rollout.
- Release gradually and retain a rollback path.
This creates an institutional asset: a growing body of evidence about what works for this organization’s users, data, and constraints. Vendor benchmarks cannot supply that knowledge.
Version the pieces that can alter behavior: system instructions, model and parameters, tool schemas, retrieval configuration, policies, and evaluation sets. Without versioning, a team cannot explain why an answer changed or reproduce an incident.

A Practical Reference Architecture
A credible enterprise stack has clear ownership boundaries:
Experience layer: the application surface, user identity, consent, citations, and feedback.
Orchestration layer: deterministic workflows, model routing, agent state, limits, retries, and approval checkpoints.
Capability layer: narrow tools and APIs with typed contracts, authorization, and idempotency controls.
Knowledge layer: governed source data, retrieval, provenance, permissions, freshness, and deletion.
Model layer: interchangeable models selected by measured task requirements rather than brand loyalty.
Control and evaluation layer: policy enforcement, traces, cost accounting, offline evaluations, online monitoring, incident response, and rollout controls.
The model layer is deliberately replaceable. Models will improve and prices will move. The durable investment is the surrounding system: high-quality data, explicit capability boundaries, representative evaluations, and operational feedback.
The Architecture Review Questions That Matter
Before approving an AI use case, ask:
- Could a deterministic workflow solve this more reliably?
- What decision is the model making, and why does it require a model?
- What is the worst plausible failure?
- Which actions can the system take, under whose authority?
- Where are access controls enforced outside the model?
- What evidence can the user inspect?
- Which cases require human judgment?
- How will we test quality before release and detect regressions afterward?
- What is the cost per successful outcome?
- Can we reproduce, explain, and roll back a behavior change?
If a team cannot answer these questions, it does not yet have an AI architecture. It has a prototype.
The strongest enterprise AI systems will not be the ones with the most agents or the largest models. They will be the ones that know where probabilistic reasoning creates value, where ordinary software is better, and how to tell the difference with evidence.
That is the stack worth building.
Sources and Further Reading
- Albert Einstein, “On the Method of Theoretical Physics,” Herbert Spencer Lecture, Oxford, 1933. The popular “as simple as possible” line is a later paraphrase.
- Harold Abelson and Gerald Jay Sussman with Julie Sussman, Structure and Interpretation of Computer Programs, MIT Press.
- Quote Investigator, “In God We Trust. All Others Must Bring Data,” documenting attribution to W. Edwards Deming.
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024.
- OWASP, Top 10 for Large Language Model Applications.