PReference architecture for enterprises
How I build AI platforms.
A reference architecture for enterprises that want engineers and agents to use AI, with control over cost, data and risk. It is written for platform leaders and for the people who sign off the spend. Each building block lists the real alternatives, what each one costs you, and the option I reach for first.
01The enterprise problem
What goes wrong without a platform.
AI arrives in an enterprise one team at a time. Each of these problems is small on its own. Together they are the reason a platform exists.
| Problem | What you see | What it costs | What answers it |
|---|---|---|---|
| Keys in every repository | Each team signs up with a provider and stores its own key. | A leaked key is an incident, and nobody can list who holds one. | AI gateway |
| Spend with no owner | The cost arrives on one invoice, weeks after the usage. | Finance cannot attribute it and engineering cannot cap it. | AI gateway |
| Shadow AI | People paste company data into personal accounts because the approved route is slower. | Data leaves the boundary with no record that it left. | Governance |
| Agents with broad access | Tools are installed per laptop and run with personal tokens. | Nobody can say which agent can reach which system. | MCP gateway |
| No evidence | An auditor asks what the model was told and what it answered. | The answer depends on logs that were never kept. | Tracing |
| Quality depends on the driver | Two engineers with the same tool get different results. | Good practice stays in individual heads and onboarding repeats. | Engineering harness |
| Knowledge in heads and threads | Decisions are made in meetings and chat, then repeated to every new joiner. | Every agent session starts from nothing, and old decisions are reopened. | Context hub |
| Locked to one supplier | Provider calls are written into every service. | A price change or a better model means a rewrite. | AI gateway |
02Reference architecture
One request, every layer.
Follow one request from the person who asks to the model that answers, and back. Every layer it passes leaves a record, and one set of rules checks it at each step. Switch the scale to see what a smaller organisation keeps.
Pick a layer
Point at a layer or tap it to see what it holds. The list below does the same from the keyboard.
Four ways in
These are moments, not audiences. The same engineer uses a coding tool to do the work and the portal to find out who owns a service.
| Surface | Who uses it | What it is for |
|---|---|---|
| Chat | Everyone, especially colleagues who never open a coding tool. | Ask, clarify, approve and receive. |
| Developer portal | Any engineer or service owner. | Discover what exists, see who owns it, onboard a service and get a key. |
| Coding tools | Engineers and engineering agents. | Do the work, with shared context and approved tools. |
| Git | Platform engineers. | Change the platform. Every change is a reviewed merge that a reconciler applies. |
03Principles
Eight opinions I build on.
- One governed gateway
- Every model call and every tool call goes through one internal gateway that knows who is calling, what it costs and what it is allowed to reach. No application holds a provider key.
- Evals define done
- A prompt, model or agent change is finished when it passes a versioned evaluation suite, the same way code is finished when its tests pass. A named person then approves the release.
- Trace everything
- Every request leaves a full trace: prompt, retrieved context, tool calls, tokens, latency and cost. Debugging, cost attribution, evaluation data and audit all read from that one record.
- Context is infrastructure
- What the model is told decides what it produces. Context has owners, versions and a review step. A meeting is evidence, an observation is provisional, and a reviewed decision is the team’s position.
- Cost is a design constraint
- Budgets per team and per key exist from the first day, and model choice is a routing decision made per task. The cheapest model that passes the evals is the right one.
- Governance that expects failure
- Models will be wrong, tools will be misused and prompts will be injected. Controls fail closed, irreversible actions wait for a person, and every decision can be reconstructed afterwards.
- Enforcement outside the prompt
- Identity, permissions and approvals are checked by code that the model cannot talk its way past. An instruction in a prompt is guidance. It is never the control.
- Evidence before claims
- The target architecture and the current state are kept apart, and every statement about the platform cites what supports it. When a decision is reversed, the record shows the reversal.
04Building blocks
Twelve decisions, with the options.
Each block says what it is, why it matters, which real options exist and what each one costs you. Then it names the one I would pick first.
04.1Building block
Model access and the AI gateway
What it is
One internal endpoint that every application and agent calls instead of a provider’s own API. It holds the provider credentials, issues its own keys per team, applies budgets, rate limits and guardrails, picks a model, and records the call.
Why it matters
Without it, provider keys spread across repositories, spend shows up only on the invoice, and changing a model means changing every caller. With it, a model change is a routing change.
| Option | What you get | What it costs you |
|---|---|---|
| LiteLLM (self-hosted proxy)My default | An OpenAI-compatible API over most providers, with virtual keys, per-team budgets, fallbacks and spend tracking. The same proxy can govern tools. Open source, with paid enterprise features. | You run and patch it, along with its database. It moves fast, so you own upgrade testing and dependency hygiene. |
| Portkey | Routing, guardrails and observability in one gateway. Available as an open-source gateway or a hosted service. | On the hosted service, prompts cross another vendor. Palo Alto Networks bought Portkey in May 2026, so check the roadmap for the open-source gateway and the features you depend on. |
| Kong AI Gateway | AI plugins on the Kong API gateway. Model traffic and ordinary API traffic share one policy engine. | It pays off mainly when Kong already fronts your APIs. Basic AI plugins are free. Load balancing, semantic caching, the MCP proxy and token rate limits need an AI Gateway Enterprise licence. |
| Cloud-native (Amazon Bedrock, Microsoft Foundry, Gemini Enterprise Agent Platform) | Identity, private networking, guardrails and billing from the cloud you already use. No extra service to run. | The catalogue stops at what that cloud offers. Per-team budgets and routing across providers are thin, so a gateway often ends up in front anyway. |
My default
LiteLLM in front of Bedrock and direct provider accounts
- The OpenAI-compatible surface means every harness and SDK works without changes.
- Each team has one record: its model allow-list, rate limits and spend cap.
- Callers use approved model aliases. The gateway records the real model, provider, region and data policy behind each one.
- A fallback route must meet the same data policy as the route it replaces.
- It is open source, so I can read what it does with a prompt before I trust it.
04.2Building block
The Model Context Protocol gateway
What it is
The Model Context Protocol (MCP) is an open standard that lets an agent call tools and read resources from a server. An MCP gateway is the place where those servers are registered, authenticated and discovered.
Why it matters
Tools installed per agent run on laptops with personal tokens, and nobody can list which agent reaches which system. A model call returns text, but a tool call can change a system. Tool policy therefore needs one thing that model policy does not: whether the tool reads or writes.
| Option | What you get | What it costs you |
|---|---|---|
| Per-agent local servers | The fastest start. No infrastructure, and it works offline. | Credentials sit on laptops, versions drift, there is no central audit, and every harness is configured separately. |
| Central servers, reached directly | One deployment per system, with the caller’s identity on each request. | There is no shared discovery and no policy per call. Every agent is configured with every server. |
| Central servers behind the model gatewayMy default | Tool permissions use the keys and teams that already govern models, and tool calls land in the same traces. LiteLLM includes an MCP gateway. | The gateway knows which tools a team may reach. It does not know which of them write, so you define that label yourself. |
| A separate agent-platform gateway | A managed gateway with policy and interceptors that sit outside the agent. Amazon Bedrock AgentCore Gateway is one example. | Used for tools, it puts tool policy on a second surface with a second identity mapping, and two places to investigate a denial. |
| Vendor-hosted connectors | The software vendor runs the server for its own product. Nothing to deploy. | Data leaves under the vendor’s terms, scopes are coarser, and you still need an allow-list. |
My default
Central servers, registered in the gateway that governs models
- One policy plane: a team’s model allow-list, tool allow-list, rate limits and spend cap sit in one record.
- Every tool is labelled as read or write. Write tools are limited to the agent profiles that need them.
- Discovery is filtered by the policy that enforces the call, so an agent never learns about a tool it cannot use.
- Access is the intersection of the caller, the agent profile, the task and the source’s own permissions.
- Who may invoke an agent is a different question from which tools it may call. A separate inbound control answers it.
04.3Building block
Context hubs and the company brain
What it is
The company brain is the knowledge that people and agents share: knowledge bases, project context hubs and reusable skills, each kept in a source that someone owns. A project context hub is a repository of its own that sits beside the code repositories. It holds what the project is for, what is understood now, what has been decided and how the team works.
Why it matters
A project is wider than one repository. It spans several, plus meetings and decisions that belong to none of them. Without a hub, that understanding lives in heads and chat threads, and every agent session starts from nothing.
| Option | What you get | What it costs you |
|---|---|---|
| Context files in each repository | Instruction and rule files next to the code. Versioned with it and read by every harness. | They stop at the repository boundary. Project-wide decisions are copied between repositories or lost. |
| A project context hubMy default | One repository per project for context, skills, conventions, dated observations and reviewed decisions. It is searchable as a corpus and installable as a set of assets. | It needs a project owner and a review habit. Decide the access scope of each hub before the first sensitive one exists. |
| Knowledge bases behind one retrieval serviceMy default | Domain evidence and approved guidance, returned with citations, permissions and review status. | Permissions must hold at query time for every source behind the endpoint. Indexes have to follow source changes and revocations. |
| Retrieval-augmented generation over a vector store | Search by meaning across a corpus too large to curate. | There are embedding pipelines to run and retrieval quality to evaluate. I start with keyword search and add vectors when they beat that baseline. |
My default
A hub per project, a knowledge base per domain, one retrieval path
- A meeting is evidence, an observation is provisional, and a reviewed decision is the team’s position. The hub keeps the three apart.
- People and agents may propose a change with evidence. An owner reviews it, and publication happens in the source that owns the record.
- Skills and standards come in two tiers, organisation and project, resolved together at pinned versions.
- A catalogue answers what exists and who owns it. Retrieval serves the content. An agent needs both.
- A summary never widens the audience of its sources.
04.4Building block
Memory
What it is
Memory is what an agent keeps between interactions. I separate four scopes, each with its own store, readers and retention: the task in hand, the person, the project and the organisation.
Why it matters
Memory without scope turns one person’s remark into company policy, or shows one client’s information to another. The scopes are not a ladder. A lesson can be proposed from a task, but a checkpoint never becomes shared guidance by itself.
| Option | What you get | What it costs you |
|---|---|---|
| Workflow checkpoints in your databaseMy default | Task state that survives a restart, in a database you already operate. | It holds task state only. Give it a short retention window, separate from the audit record. |
| A managed memory serviceMy default | Extraction and retrieval of a person’s preferences without building them, with isolation enforced by the cloud’s identity system. | Check how deletion works. Where it is per record, removing everything for one client becomes an enumeration. Purpose, notice and access policy stay with you. |
| Project memory in GitMy default | Dated, attributed observations beside reviewed decisions, with history and review built in. | Git history keeps what you delete. Revocation has to cover clones and indexes as well. |
| A vector memory store | Recall by similarity across long histories. | It is hard to audit and hard to revoke by source. A namespace string is not isolation. |
My default
Four scopes, four stores, no automatic promotion
- Task state lives in checkpoints, a person’s preferences in a managed memory service, project memory in the context hub and organisation knowledge in owned knowledge bases.
- Person memory is opt-in and private by default.
- An observation becomes guidance only through an owner’s review.
- A paused task checks its permissions again when it resumes.
04.5Building block
Engineering towers
Draft definition. I am still refining how I describe this one.
What it is
A tower is a domain-aligned engineering capability group, for example data engineering, backend or cloud. It owns the reusable AI assets for its domain: skills, prompts, context packs, evals and guardrails. The platform team provides the rails and the towers provide the domain content.
Why it matters
A central platform team cannot write good guidance for every domain, and shared assets without a domain owner decay.
| Option | What you get | What it costs you |
|---|---|---|
| Central platform team owns everything | Consistency, and a fast start. | The team becomes a bottleneck and the guidance lacks domain depth. |
| Each product team owns its own | Assets written closest to the work. | Duplication, uneven quality, and little that another team can reuse. |
| Towers own, the platform curatesMy default | Domain depth with shared standards. Assets are versioned and released. | It needs named owners with time set aside, and some coordination. |
| Guild or community of practice | Low ceremony and voluntary. | Nobody is accountable, so assets depend on enthusiasm. |
My default
Towers, with a published contract for every asset
- Each asset has an owner, a version, an eval and a retirement path.
- The platform team reviews for safety and consistency and leaves domain judgement to the tower.
04.6Building block
The general company agent
What it is
One agent that every colleague can reach from chat, including people who never open a coding tool. It answers with citations, prepares documents, proposes changes to shared context and hands engineering work to the engineering path. Its harness is the loop, the state and the tool handling around the model.
Why it matters
Coding tools serve engineers, and most of a company is not engineers. A company agent gives everyone the same governed models, tools and knowledge. It keeps durable state, so an approval can take a day without losing the task.
| Option | What you get | What it costs you |
|---|---|---|
| LangGraphMy default | Explicit workflow state, database checkpoints, and interrupt and resume for approvals. Generally available since version 1.0, with a large community. | More assembly than a packaged harness. A resumed step runs again from its start, so side effects must be safe to repeat. |
| Deep Agents on LangGraph | A ready-made harness with planning and delegation, built on LangGraph. | One more layer to pin and understand. The company adapters are still yours to build. |
| Strands Agents | An open-source SDK from AWS with human-in-the-loop interrupts before tool calls, and an MCP client. | Younger than LangGraph, with a smaller community. |
| Vendor agent SDKs | The OpenAI Agents SDK and the Claude Agent SDK integrate closely with their own platforms. | One vendor’s runtime sits at the core of the company agent. |
| Agent as configuration | A managed console defines the agent. It is the fastest route to something running. | The agent sits outside source control and review, unlike every other workload. |
My default
LangGraph, shipped as a reviewed container image
- The agent is code in a repository, built in continuous integration and deployed as an immutable version. Rollback points the endpoint at the previous version.
- A managed agent runtime can isolate each session. Kubernetes is the alternative when you already run it.
- The approval ledger is a business record in its own table. It is never session state.
- Every action is tracked as planned, authorised, submitted, confirmed, failed or uncertain. An uncertain write is checked downstream before anything is retried.
- The agent prepares engineering tasks. It does not inherit engineering credentials.
- Default profiles have no shell and no host file system.
04.7Building block
The engineering harness
What it is
The harness is everything around the model in a coding agent: instructions, skills, tools, hooks and permissions. A company harness is one shared, versioned set of those that works across Claude Code, Codex, Cursor and Kiro. Three parts do the work: a registry of approved assets, an installer that resolves them at pinned versions, and a renderer that writes each tool’s own files and checks them for drift.
Why it matters
Engineers choose different tools. If the guardrails live in one tool’s configuration, every other tool bypasses them.
| Option | What you get | What it costs you |
|---|---|---|
| Standardise on one tool | One configuration format and the deepest integration. | Lock-in. Engineers work around it, and one vendor’s outage or price change reaches everyone. |
| Per-team configuration | Freedom for each team. | Drift. Guardrails differ per repository and onboarding is repeated. |
| A shared harnessMy default | One source of skills, rules, hooks and MCP configuration, rendered into each tool’s format. Versioned and reviewable. | You maintain adapters as tool formats change. |
| Open conventions only | Portable files such as AGENTS.md, skills and MCP, with no build step. | Coverage differs per tool. Hooks and permissions are not standardised. |
| Install from a developer portal | One place to browse skills, owners and how assets relate. Backstage is the usual choice. | A catalogue holds pointers. Installing from search results cannot guarantee exact, pinned content. |
My default
A shared harness: registry, installer and renderer
- Assets resolve from the registry at a pinned reference, never from a discovery page.
- One source record produces two projections: the registry for installs and the catalogue for browsing.
- Generated files are recorded with hashes, so drift is detected and hand edits survive.
- Open conventions carry instructions, skills and tools. Thin adapters cover hooks and permissions.
- All model traffic goes through the gateway, so the guardrails hold whichever tool an engineer opens.
04.8Building block
Tracing
What it is
A trace records one request as a tree of spans: the prompt, the retrievals, the tool calls, the model calls, and the tokens, latency and cost of each.
Why it matters
An agent failure is rarely one bad call. You need the whole tree to debug it, to attribute cost, to build evaluation data from production, and to answer an auditor.
| Option | What you get | What it costs you |
|---|---|---|
| MLflow TracingMy default | Open source and OpenTelemetry-compatible. Traces, evaluation runs and the model registry live in one tool. Self-hosted, or managed on Databricks. | The self-hosted interface and access control are plainer than dedicated products. The richest experience is on Databricks. |
| Langfuse | Open source, self-hosted or cloud. Strong prompt management, sessions, annotation queues and evals. | Self-hosting means running Postgres, ClickHouse, Redis and object storage. It joined ClickHouse in January 2026 and says it stays open source. |
| OpenTelemetry-native | AI spans flow through the telemetry pipeline and backend you already run. No new vendor. | The generative-AI semantic conventions are still at Development status, and general backends have no prompt views or eval workflow. You build those. |
| Vendor tools (LangSmith, Datadog, Arize, Braintrust) | Polished and quick to adopt, with evaluation built in. | Prompts and outputs leave your boundary unless you buy a self-hosted or hybrid enterprise plan, and Datadog offers none. Cost grows with the traces, spans or data you send. |
My default
MLflow Tracing, fed over OpenTelemetry
- One system holds traces, evaluation runs and models, so an eval can cite the trace it scored.
- Agents export standard OpenTelemetry, so the backend stays replaceable and the agent framework can change.
- The trace store holds prompts, tool arguments and retrieved context. It sits behind authentication, on an internal route, with traces separated per agent.
- Approval and security records are kept in their own store. Sampled traces are not the audit trail.
- I would choose Langfuse when a team needs prompt management and annotation queues more than lakehouse integration.
04.9Building block
Evaluation
What it is
Evaluation is two systems with two jobs. Controlled evals test whether an agent completes representative tasks before a release. Monitoring watches how real interactions behave after it. The methods below feed one or the other.
Why it matters
Without evals, a prompt or model change is an opinion. With them, a change can be gated like code, and a person can approve a release on evidence.
| Option | What you get | What it costs you |
|---|---|---|
| Controlled task evalsMy default | A versioned suite of representative tasks, run on demand and in continuous integration. Inspect AI is an open-source framework for this. | Suites need owners and held-out cases. A suite built only from the failures that prompted a change proves little. |
| Deterministic checksMy default | Exact match, schema validity or passing tests. Cheap and repeatable. | They cover only the cases someone thought to write. |
| Large language model as judge | Scores open-ended output for faithfulness, relevance and tone, at scale. | The judge must be calibrated against human labels. It costs tokens and can drift when the judge model changes. |
| Human review | Ground truth for nuance, and the labels that calibrate a judge. | Slow, expensive, and inconsistent between reviewers without a rubric. |
| Monitoring of production traces | Quality signals, latency and cost from real use. | It reports after the user has seen the result. |
| Regression gatesMy default | A merge or rollout is blocked when scores fall below the threshold. | Noisy evals make flaky gates. Each threshold needs an owner. |
My default
Controlled evals gate the release, production traces watch it
- Neither system approves a release. A named person reviews the evidence and accepts the remaining risk.
- An agent cannot approve its own release.
- A release is bound to its image, model route, prompt and skill versions, retrieval configuration and eval result.
- Suites cover task quality, source fidelity, unsupported claims, privacy leakage, prompt injection and tool misuse.
- Corrections and incidents become new test cases after an owner reviews them.
04.10Building block
Open-source model serving
What it is
Running open-weight models, or your own fine-tuned models, on infrastructure you control. They sit behind the same gateway as hosted models, so callers cannot tell the difference.
Why it matters
Some workloads cannot leave the boundary. Some are narrow, high-volume tasks where a small fine-tuned model is cheaper and faster. Some need a model pinned for reproducibility.
| Option | What you get | What it costs you |
|---|---|---|
| vLLMMy default | High-throughput serving with continuous batching, an OpenAI-compatible server, adapter serving for fine-tunes, and wide model support. | You own the graphics processors, autoscaling, upgrades and capacity planning. |
| Text Generation Inference | Hugging Face’s serving stack. Mature and well documented. | In maintenance mode since December 2025, and its repository archived since March 2026. Hugging Face points new deployments to vLLM or SGLang. Keep what runs; do not start here. |
| OllamaMy default | One command runs a quantised model locally, behind a subset of the OpenAI API. Suits development, demos and offline work. | It is not built for multi-tenant production throughput. |
| Amazon SageMaker endpoints | Managed hosting with autoscaling, identity and private networking, for your own container. | Real-time endpoints bill per instance-hour whether they are busy or idle. Serverless inference scales to zero, with less control of the hardware. |
| Amazon Bedrock Custom Model Import | Bring fine-tuned weights and call them through the Bedrock API. No endpoints to manage. | Only supported architectures. A copy scales to zero after five idle minutes, so the next call waits for a cold start. Less control of the runtime. |
My default
Ollama for development, vLLM in production
- Both expose an OpenAI-compatible API, so moving from laptop to cluster is mostly a change of URL.
- Every self-served model is registered in the gateway and inherits its budgets, guardrails and tracing.
- Open weights do not always mean an open licence. The licence is checked before the route is approved.
- Managed import suits low or irregular volume on a supported architecture.
- Hosted models stay the default. I self-serve for a data boundary, a cost case or a latency case.
04.11Building block
Governance, retention and security
What it is
Governance decides what the platform may do and what evidence permits it: who can call which model and tool with which data, what is kept and for how long, and what happens when something goes wrong. It follows each use case through its life: define it, assess its impact, test it, approve its release, monitor it, then change or retire it.
Why it matters
In a regulated setting the model will fail at some point. What matters is whether the failure is contained, visible and explainable afterwards, and whether data can be withdrawn when permission ends.
| Option | What you get | What it costs you |
|---|---|---|
| Policy documents and training | Cheap and quick to publish. | Nothing is enforced and there is no evidence of compliance. |
| Block by default | Low exposure on paper. | Usage moves to personal accounts, where nothing is visible. |
| Controls enforced in the gateway, the harness and the sourcesMy default | Identity, allow-lists, guardrails, budgets and logging defined as code. Each control produces evidence. | Platform engineering effort, and guardrail false positives that need tuning. |
| Guardrail services | Ready-made detectors for personal data, harmful content and prompt injection. | Extra latency and cost per call. Detection is probabilistic, and you still decide what a hit does. |
| A separate governance platform | Inventories, assessments and approvals in one product. | One more system of record. I link the records the catalogue and the delivery pipeline already hold. |
My default
Enforced controls, designed around failure
- Identity, permissions and approvals are enforced by code outside the model. An instruction in a prompt is not a control.
- An approval is bound to the exact action, its parameters and an expiry. A changed parameter invalidates it.
- Revoking access and deleting copies are separate operations, with separate deadlines and evidence. Deletion follows lineage through indexes, embeddings, caches, checkpoints and backups.
- Every class of data has an owner, a purpose and a retention schedule before it is collected. A class with no schedule is not collected.
- An operator can disable a model route, a tool or a whole service, and stop pending work.
- Retrieved and tool-returned content is treated as untrusted input.
- Controls map to public frameworks such as the NIST AI Risk Management Framework and the 2026 OWASP Top 10 for LLM Applications.
04.12Building block
The AI Development Life Cycle
What it is
The delivery workflow defines how work moves from intent to running software when agents write most of the code: what gets specified, where a person approves, and what must pass before a merge. I treat it as a software factory with a contract at each handover.
Why it matters
Agents produce code faster than people can review it. The workflow decides where human attention is spent, and it keeps an agent from acquiring the right to merge or deploy.
| Option | What you get | What it costs you |
|---|---|---|
| AWS AI-DLC (aidlc-workflows) | Open-source AI-DLC workflows from AWS Labs. Five phases, from initialisation to operation, with a human approval gate after each stage and traceable artefacts. | Heavy for small changes. I am exploring it. It earns a place only where it adds something the existing gates lack. |
| Spec-driven tools | Requirements, design and tasks are written before code. | Effort up front, and specs that drift from the code. |
| Gated pipelineMy default | The agent works freely on a branch. A gate then runs review, tests, lint and documentation checks before the push, and stops on findings that need a person. | Problems are caught after the code is written. Each change pays pipeline time and needs a clear statement of intent. |
| Pull-request review alone | Familiar, with nothing new to run. | Human review becomes the bottleneck, and quality depends on reviewer stamina. |
My default
A gated pipeline in the style of no-mistakes
- Enforcement sits outside the agent, so the model never marks its own work.
- An agent’s task ends at a draft pull request. A deterministic publisher with its own scoped identity opens it.
- Model review is advisory. Executable checks and a human reviewer decide.
- The build produces one immutable image with its evidence. A separate, reviewed change selects that image for an environment.
- Merging and deploying stay human decisions.
05Data platform
The data platform underneath.
An AI platform is only as good as the data it can reach and the data it can be tested on. I build the lakehouse first and treat models as one more consumer of it.
My background is in data platforms for financial services: Kafka, Spark, Snowflake, Databricks and Airflow. I hold the Databricks Data Engineer Associate and Generative AI Engineer Associate certifications.
- Lakehouse
- One storage layer for tables, files and AI artefacts, in an open table format and behind one catalogue. Data is refined in stages: raw (bronze), cleaned (silver) and ready to use (gold).
- Data quality
- Expectations are written as code and run inside the pipeline. Failing rows are quarantined and freshness and volume are monitored. A model evaluated on bad data reports a false pass.
- Feeding models and evaluation
- Gold tables become curated context, retrieval indexes, fine-tuning sets and golden evaluation sets. Traces flow back in as ordinary tables, so production behaviour becomes new evaluation cases.
- Sources
- Bronze: raw
- Silver: cleaned
- Quality gate
- Gold: ready to use
- Curated context
- Retrieval indexes
- Fine-tuning sets
- Evaluation sets
Traces return to bronze as new data
| Option | What you get | What it costs you |
|---|---|---|
| DatabricksMy default | Unity Catalog governs tables, volumes, functions, models and AI Search (vector) indexes. Managed MLflow puts traces and evals next to the data. | Platform cost, and the workspace model is a commitment. |
| Snowflake | A strong warehouse with simple operations and mature sharing. | Machine learning and open-format workflows are newer there than the SQL core. |
| Open stack on object storage | Open table formats with your choice of engine. The most control and the least lock-in. | You integrate and operate the catalogue, engines and governance yourself. |
| Cloud-native warehouse | It fits the cloud account, identity and billing you already have. | It ties the data platform to one cloud, and AI tooling varies by provider. |
My default
Databricks for the lakehouse, open formats underneath
- Data, models and traces share one governance model.
- Open table formats keep the data readable by other engines.
06Operating model
Who owns what.
A platform fails when ownership is vague. These are the roles I set up, what each one decides, and the record each one keeps. A small organisation can combine roles, but the person who reviews a consequential action needs the authority to reject it.
| Role | Owns | Decides | Record it keeps |
|---|---|---|---|
| Business owner | The use case: its intended users, its benefit and its acceptable limits. | What the service is for, and what it must never be used for. | The use-case record and its success measures. |
| AI service owner | The application, its evals and its operating instructions. | What changes, and when a change needs re-evaluation. | The versioned service record and release evidence. |
| Knowledge and data owners | Sources, their classification and their review. | Which sources agents may read, and how long derived copies live. | The source register, access scope and retention schedule. |
| Platform team | The gateways, tracing, the shared harness and the paved road between them. | How access, isolation, quotas and recovery are enforced. | Configuration, verification results and runbooks. |
| Engineering towers | Domain assets: skills, context packs, evals and guardrails. | What good output looks like in their domain. | Versioned assets, each with an eval. |
| Security, privacy and responsible-AI reviewers | The assessment of risk for each use case. | Conditions and exceptions. | The impact assessment and the review decision. |
| Release approver | The decision to release. | Whether the remaining risk is acceptable for this version and scope. | An approval bound to a version, a scope and a review date. |
| Service operator | The running service. | When to stop, roll back or escalate. | Incident, corrective action and recovery records. |
| Finance | Budgets and chargeback. | The budget for each team. | Spend against budget, by team and by model. |
07Adoption and maturity
Where to start, and what comes next.
There are three honest places to start. Which one fits depends on where the pressure is today.
Gateway first
My default
Cost, keys or the data boundary is the pressure.
One endpoint, keys tied to single sign-on, budgets per team, and tracing switched on.
Harness first
Coding tools are already in use and results are uneven.
One shared harness across repositories, guardrails, and MCP servers for the systems engineers reach most.
Use case first
One business use case has a sponsor and a deadline.
A thin slice through every layer for that use case, with evals that define done.
I start with the gateway. The other two paths need it within weeks, and it is the step that makes spend and usage visible.
A maturity roadmap
Each stage ends with evidence you can show. Move on when you can show it.
| Stage | What you add | Evidence that you are there |
|---|---|---|
| Ad hoc | Individual keys and tools, chosen team by team. | You have a list of who is using what. |
| Governed access | The gateway, identity, budgets and allow-lists. | You can say who spent what, on which model, last month. |
| Observable | Tracing on every call, with a retention policy. | You can reconstruct any request an auditor asks about. |
| Evaluated | Owned evaluation suites and a named release approver. | A model swap is a routing change backed by eval results. |
| Shared context and tools | Context hubs, the knowledge path and tools registered in the gateway. | The company agent and a coding tool retrieve the same approved revision, within the caller’s access. |
| Agents that act | Durable tasks, an approval ledger, and engineering agents that end at a draft pull request. | Any agent action can be approved, stopped and reconstructed. |
| Optimised | Routing by cost and quality, self-served models where they pay, and evals fed from production. | Cost per accepted result falls while eval scores hold. |
08Evaluation criteria
How to judge any option.
These are the questions behind every table on this page. In a regulated setting I weigh the first three most.
- Data boundary
- Where do prompts and outputs go, and who can read them? Can it run inside our network?
- Identity and access
- Does it use our single sign-on and our groups? Are keys issued per person or per team?
- Auditability
- Can we reconstruct one request from end to end, long after it happened?
- Revocation and retention
- Can we stop use at once, delete the copies on a schedule, and prove both?
- Cost control
- Can we cap spend per team before the money is spent?
- Operability
- Who patches it, what does an upgrade involve, and what happens when it is down?
- Exit cost
- What do we rewrite if we leave? Are the interfaces open standards?
- Supplier risk
- Who owns it, how is it licensed, and what changed in the last year?
- Fit with the estate
- Does it work with the cloud, the data platform and the tools we already run?
09Track two: smaller organisations
The same principles, at a smaller scale.
A small company cannot run an enterprise platform, and it should not try. GuideX, the travel marketplace I founded, runs agents in production with a small team. It keeps the principles and leaves out most of the machinery.
The worked example is the GuideX site-reliability estate: agents that triage errors, check payments and answer the team’s questions. A person approves each case before a coding agent picks it up.
- Managed services first
- Hosted model APIs, serverless containers that scale to zero, and a managed error tracker. There are no graphics processors to buy and no model servers to patch.
- A thin gateway
- A shared policy library that every agent embeds, in place of a gateway service in every call path. It enforces tool tiers, filters personal data, writes the audit record and checks the budget. One small orchestrator, with no model inside it, routes the requests that come from people.
- One internal tool server
- A single MCP server brokers every upstream system. It fails closed and accepts service identities only. The chat assistant holds no production credentials.
- Lightweight tracing
- One trace identifier travels from the mobile app through the backend to every agent action. Audit records live in the database the product already runs. A dedicated tracing platform can wait.
- Lightweight evaluation
- Every model output is validated against a schema, golden cases run in continuous integration, and a person approves anything irreversible. A low-cost model is the default. A stronger one is switched on per agent when measured failure rates justify it.
- A pragmatic harness
- One repository holds the agent instructions, hooks, commands and quality gates as plain files. A script copies them into each product repository and stamps the version, and a check reports drift. The rules that matter are gates that run before a commit and again in continuous integration.
- Budgets and one switch
- Each agent has a daily budget and an alert before the cap. One switch halts every agent at once.
- A deterministic shell
- Validation, authentication, tool tiers and the audit trail are plain code. Models reason only inside that shell, so flexibility never loosens a control.
Pick a layer
Point at a layer or tap it to see what it holds. The list below does the same from the keyboard.
| Building block | Enterprise | Smaller organisation | What stays the same |
|---|---|---|---|
| Model access | Many providers behind a self-hosted gateway. | Hosted APIs, called through a shared policy library. | Every call has an owner and a budget. |
| Tools for agents | Central MCP servers behind a gateway, with single sign-on. | One internal MCP server that fails closed. | Tool access is a reviewed decision. |
| Context | A central hub, with towers that own domain assets. | Instruction files in each repository, distributed from one source. | Context is versioned and reviewed. |
| Memory | Four scopes in four stores, with reviewed promotion. | Task state, an audit log and project notes kept as files. | An observation is not guidance until someone reviews it. |
| Company agent | A company agent on a managed runtime, with an approval ledger. | A chat assistant that can read and draft only, behind the orchestrator. | The agent that talks to people holds no production credentials. |
| Ownership | Named owners for the platform, each service, each source and each release. | The engineers who build the product. | Every asset has a named owner. |
| Tracing | A tracing platform such as MLflow or Langfuse. | Trace identifiers and audit records in the existing database. | Any action can be reconstructed. |
| Evaluation | Golden datasets, calibrated judges and regression gates. | Schema checks, golden cases and human approval. | A change is done when its checks pass. |
| Model serving | Self-served models where the boundary or the cost requires them. | None. Hosted models only. | The cheapest model that passes is the right one. |
| Harness | A shared harness with adapters for four coding tools. | Plain files copied by a script, for one or two tools. | Guardrails are enforced by code. |
| Governance | Controls as code, mapped to public frameworks and tested by a risk function. | Tool tiers, a personal-data filter, budgets and one switch, in one library. | Controls fail closed and leave evidence. |
| Delivery | A gated pipeline, with written plans for boundary changes. | The same gates as scripts. An agent’s pull request passes the gates a person’s does. | The agent never marks its own work. |
| Data platform | A lakehouse. | The product database and the cloud’s own logging. A warehouse comes when analysis needs one. | Traces are data. |
When to add the heavier parts
- A second team starts building with models, and spend needs attributing.
- A customer or a regulator asks for evidence you cannot produce from the audit records.
- Engineers use more coding tools than one script can keep aligned.
- A workload cannot leave your boundary, or its volume makes a hosted model the expensive choice.
10Working with me
What working with me looks like.
I have built this kind of platform inside a financial-services engineering organisation: a governed gateway to hosted models on AWS, a self-service developer portal, MLflow tracing, team budgets, authenticated MCP servers, and a shared harness across Claude Code, Cursor, Kiro and GitHub Copilot.
- Start with a conversationSixty minutes on what you run today, what constrains you, and where the risk sits.
- Write the options downYou get the alternatives, the trade-offs and a recommendation in plain language, with diagrams.
- Build a thin sliceA gateway, tracing and one eval gate, working for one team, before anything is scaled.
- Hand it overRunbooks, decisions and ownership sit with your engineers.
Book a 60-minute technical deep dive
Bring an architecture, a constraint or a decision you are stuck on. We work through it together.