PReference architecture for enterprises

How I build AI platforms.

A reference architecture for enterprises that want engineers and agents to use AI, with control over cost, data and risk. It is written for platform leaders and for the people who sign off the spend. Each building block lists the real alternatives, what each one costs you, and the option I reach for first.

Running a smaller organisation? Go to that track

Written from public tools and my own practice. It describes no employer’s or client’s system.

01The enterprise problem

What goes wrong without a platform.

AI arrives in an enterprise one team at a time. Each of these problems is small on its own. Together they are the reason a platform exists.

ProblemsWhat ungoverned adoption looks like
ProblemWhat you seeWhat it costsWhat answers it
Keys in every repositoryEach team signs up with a provider and stores its own key.A leaked key is an incident, and nobody can list who holds one.AI gateway
Spend with no ownerThe cost arrives on one invoice, weeks after the usage.Finance cannot attribute it and engineering cannot cap it.AI gateway
Shadow AIPeople paste company data into personal accounts because the approved route is slower.Data leaves the boundary with no record that it left.Governance
Agents with broad accessTools are installed per laptop and run with personal tokens.Nobody can say which agent can reach which system.MCP gateway
No evidenceAn auditor asks what the model was told and what it answered.The answer depends on logs that were never kept.Tracing
Quality depends on the driverTwo engineers with the same tool get different results.Good practice stays in individual heads and onboarding repeats.Engineering harness
Knowledge in heads and threadsDecisions are made in meetings and chat, then repeated to every new joiner.Every agent session starts from nothing, and old decisions are reopened.Context hub
Locked to one supplierProvider calls are written into every service.A price change or a better model means a rewrite.AI gateway

02Reference architecture

One request, every layer.

Follow one request from the person who asks to the model that answers, and back. Every layer it passes leaves a record, and one set of rules checks it at each step. Switch the scale to see what a smaller organisation keeps.

Scale of the architecture
BronzeSilverGold8Data platform6Models5AI gateway4Context hub3MCP gateway2Harness1Surfaces7TracingGGovernance
A request descends through the stack, passing a checkpoint at every layer. Every layer reports to tracing.

Pick a layer

Point at a layer or tap it to see what it holds. The list below does the same from the keyboard.

Four ways in

These are moments, not audiences. The same engineer uses a coding tool to do the work and the portal to find out who owns a service.

SurfacesWho comes in, and where
SurfaceWho uses itWhat it is for
ChatEveryone, especially colleagues who never open a coding tool.Ask, clarify, approve and receive.
Developer portalAny engineer or service owner.Discover what exists, see who owns it, onboard a service and get a key.
Coding toolsEngineers and engineering agents.Do the work, with shared context and approved tools.
GitPlatform engineers.Change the platform. Every change is a reviewed merge that a reconciler applies.

03Principles

Eight opinions I build on.

One governed gateway
Every model call and every tool call goes through one internal gateway that knows who is calling, what it costs and what it is allowed to reach. No application holds a provider key.
Evals define done
A prompt, model or agent change is finished when it passes a versioned evaluation suite, the same way code is finished when its tests pass. A named person then approves the release.
Trace everything
Every request leaves a full trace: prompt, retrieved context, tool calls, tokens, latency and cost. Debugging, cost attribution, evaluation data and audit all read from that one record.
Context is infrastructure
What the model is told decides what it produces. Context has owners, versions and a review step. A meeting is evidence, an observation is provisional, and a reviewed decision is the team’s position.
Cost is a design constraint
Budgets per team and per key exist from the first day, and model choice is a routing decision made per task. The cheapest model that passes the evals is the right one.
Governance that expects failure
Models will be wrong, tools will be misused and prompts will be injected. Controls fail closed, irreversible actions wait for a person, and every decision can be reconstructed afterwards.
Enforcement outside the prompt
Identity, permissions and approvals are checked by code that the model cannot talk its way past. An instruction in a prompt is guidance. It is never the control.
Evidence before claims
The target architecture and the current state are kept apart, and every statement about the platform cites what supports it. When a decision is reversed, the record shows the reversal.

04Building blocks

Twelve decisions, with the options.

Each block says what it is, why it matters, which real options exist and what each one costs you. Then it names the one I would pick first.

04.1Building block

Model access and the AI gateway

What it is

One internal endpoint that every application and agent calls instead of a provider’s own API. It holds the provider credentials, issues its own keys per team, applies budgets, rate limits and guardrails, picks a model, and records the call.

Why it matters

Without it, provider keys spread across repositories, spend shows up only on the invoice, and changing a model means changing every caller. With it, a model change is a routing change.

AlternativesAI gateway options
OptionWhat you getWhat it costs you
LiteLLM (self-hosted proxy)My defaultAn OpenAI-compatible API over most providers, with virtual keys, per-team budgets, fallbacks and spend tracking. The same proxy can govern tools. Open source, with paid enterprise features.You run and patch it, along with its database. It moves fast, so you own upgrade testing and dependency hygiene.
PortkeyRouting, guardrails and observability in one gateway. Available as an open-source gateway or a hosted service.On the hosted service, prompts cross another vendor. Palo Alto Networks bought Portkey in May 2026, so check the roadmap for the open-source gateway and the features you depend on.
Kong AI GatewayAI plugins on the Kong API gateway. Model traffic and ordinary API traffic share one policy engine.It pays off mainly when Kong already fronts your APIs. Basic AI plugins are free. Load balancing, semantic caching, the MCP proxy and token rate limits need an AI Gateway Enterprise licence.
Cloud-native (Amazon Bedrock, Microsoft Foundry, Gemini Enterprise Agent Platform)Identity, private networking, guardrails and billing from the cloud you already use. No extra service to run.The catalogue stops at what that cloud offers. Per-team budgets and routing across providers are thin, so a gateway often ends up in front anyway.

My default

LiteLLM in front of Bedrock and direct provider accounts

  • The OpenAI-compatible surface means every harness and SDK works without changes.
  • Each team has one record: its model allow-list, rate limits and spend cap.
  • Callers use approved model aliases. The gateway records the real model, provider, region and data policy behind each one.
  • A fallback route must meet the same data policy as the route it replaces.
  • It is open source, so I can read what it does with a prompt before I trust it.

04.2Building block

The Model Context Protocol gateway

What it is

The Model Context Protocol (MCP) is an open standard that lets an agent call tools and read resources from a server. An MCP gateway is the place where those servers are registered, authenticated and discovered.

Why it matters

Tools installed per agent run on laptops with personal tokens, and nobody can list which agent reaches which system. A model call returns text, but a tool call can change a system. Tool policy therefore needs one thing that model policy does not: whether the tool reads or writes.

AlternativesWays to give agents tools
OptionWhat you getWhat it costs you
Per-agent local serversThe fastest start. No infrastructure, and it works offline.Credentials sit on laptops, versions drift, there is no central audit, and every harness is configured separately.
Central servers, reached directlyOne deployment per system, with the caller’s identity on each request.There is no shared discovery and no policy per call. Every agent is configured with every server.
Central servers behind the model gatewayMy defaultTool permissions use the keys and teams that already govern models, and tool calls land in the same traces. LiteLLM includes an MCP gateway.The gateway knows which tools a team may reach. It does not know which of them write, so you define that label yourself.
A separate agent-platform gatewayA managed gateway with policy and interceptors that sit outside the agent. Amazon Bedrock AgentCore Gateway is one example.Used for tools, it puts tool policy on a second surface with a second identity mapping, and two places to investigate a denial.
Vendor-hosted connectorsThe software vendor runs the server for its own product. Nothing to deploy.Data leaves under the vendor’s terms, scopes are coarser, and you still need an allow-list.

My default

Central servers, registered in the gateway that governs models

  • One policy plane: a team’s model allow-list, tool allow-list, rate limits and spend cap sit in one record.
  • Every tool is labelled as read or write. Write tools are limited to the agent profiles that need them.
  • Discovery is filtered by the policy that enforces the call, so an agent never learns about a tool it cannot use.
  • Access is the intersection of the caller, the agent profile, the task and the source’s own permissions.
  • Who may invoke an agent is a different question from which tools it may call. A separate inbound control answers it.

Sources, checked 28 September 2026

04.3Building block

Context hubs and the company brain

What it is

The company brain is the knowledge that people and agents share: knowledge bases, project context hubs and reusable skills, each kept in a source that someone owns. A project context hub is a repository of its own that sits beside the code repositories. It holds what the project is for, what is understood now, what has been decided and how the team works.

Why it matters

A project is wider than one repository. It spans several, plus meetings and decisions that belong to none of them. Without a hub, that understanding lives in heads and chat threads, and every agent session starts from nothing.

AlternativesWhere shared context lives and how it is served
OptionWhat you getWhat it costs you
Context files in each repositoryInstruction and rule files next to the code. Versioned with it and read by every harness.They stop at the repository boundary. Project-wide decisions are copied between repositories or lost.
A project context hubMy defaultOne repository per project for context, skills, conventions, dated observations and reviewed decisions. It is searchable as a corpus and installable as a set of assets.It needs a project owner and a review habit. Decide the access scope of each hub before the first sensitive one exists.
Knowledge bases behind one retrieval serviceMy defaultDomain evidence and approved guidance, returned with citations, permissions and review status.Permissions must hold at query time for every source behind the endpoint. Indexes have to follow source changes and revocations.
Retrieval-augmented generation over a vector storeSearch by meaning across a corpus too large to curate.There are embedding pipelines to run and retrieval quality to evaluate. I start with keyword search and add vectors when they beat that baseline.

My default

A hub per project, a knowledge base per domain, one retrieval path

  • A meeting is evidence, an observation is provisional, and a reviewed decision is the team’s position. The hub keeps the three apart.
  • People and agents may propose a change with evidence. An owner reviews it, and publication happens in the source that owns the record.
  • Skills and standards come in two tiers, organisation and project, resolved together at pinned versions.
  • A catalogue answers what exists and who owns it. Retrieval serves the content. An agent needs both.
  • A summary never widens the audience of its sources.

04.4Building block

Memory

What it is

Memory is what an agent keeps between interactions. I separate four scopes, each with its own store, readers and retention: the task in hand, the person, the project and the organisation.

Why it matters

Memory without scope turns one person’s remark into company policy, or shows one client’s information to another. The scopes are not a ladder. A lesson can be proposed from a task, but a checkpoint never becomes shared guidance by itself.

AlternativesWhere memory lives
OptionWhat you getWhat it costs you
Workflow checkpoints in your databaseMy defaultTask state that survives a restart, in a database you already operate.It holds task state only. Give it a short retention window, separate from the audit record.
A managed memory serviceMy defaultExtraction and retrieval of a person’s preferences without building them, with isolation enforced by the cloud’s identity system.Check how deletion works. Where it is per record, removing everything for one client becomes an enumeration. Purpose, notice and access policy stay with you.
Project memory in GitMy defaultDated, attributed observations beside reviewed decisions, with history and review built in.Git history keeps what you delete. Revocation has to cover clones and indexes as well.
A vector memory storeRecall by similarity across long histories.It is hard to audit and hard to revoke by source. A namespace string is not isolation.

My default

Four scopes, four stores, no automatic promotion

  • Task state lives in checkpoints, a person’s preferences in a managed memory service, project memory in the context hub and organisation knowledge in owned knowledge bases.
  • Person memory is opt-in and private by default.
  • An observation becomes guidance only through an owner’s review.
  • A paused task checks its permissions again when it resumes.

04.5Building block

Engineering towers

Draft definition. I am still refining how I describe this one.

What it is

A tower is a domain-aligned engineering capability group, for example data engineering, backend or cloud. It owns the reusable AI assets for its domain: skills, prompts, context packs, evals and guardrails. The platform team provides the rails and the towers provide the domain content.

Why it matters

A central platform team cannot write good guidance for every domain, and shared assets without a domain owner decay.

AlternativesWho owns reusable AI assets
OptionWhat you getWhat it costs you
Central platform team owns everythingConsistency, and a fast start.The team becomes a bottleneck and the guidance lacks domain depth.
Each product team owns its ownAssets written closest to the work.Duplication, uneven quality, and little that another team can reuse.
Towers own, the platform curatesMy defaultDomain depth with shared standards. Assets are versioned and released.It needs named owners with time set aside, and some coordination.
Guild or community of practiceLow ceremony and voluntary.Nobody is accountable, so assets depend on enthusiasm.

My default

Towers, with a published contract for every asset

  • Each asset has an owner, a version, an eval and a retirement path.
  • The platform team reviews for safety and consistency and leaves domain judgement to the tower.

04.6Building block

The general company agent

What it is

One agent that every colleague can reach from chat, including people who never open a coding tool. It answers with citations, prepares documents, proposes changes to shared context and hands engineering work to the engineering path. Its harness is the loop, the state and the tool handling around the model.

Why it matters

Coding tools serve engineers, and most of a company is not engineers. A company agent gives everyone the same governed models, tools and knowledge. It keeps durable state, so an approval can take a day without losing the task.

AlternativesWhat the company agent is built on
OptionWhat you getWhat it costs you
LangGraphMy defaultExplicit workflow state, database checkpoints, and interrupt and resume for approvals. Generally available since version 1.0, with a large community.More assembly than a packaged harness. A resumed step runs again from its start, so side effects must be safe to repeat.
Deep Agents on LangGraphA ready-made harness with planning and delegation, built on LangGraph.One more layer to pin and understand. The company adapters are still yours to build.
Strands AgentsAn open-source SDK from AWS with human-in-the-loop interrupts before tool calls, and an MCP client.Younger than LangGraph, with a smaller community.
Vendor agent SDKsThe OpenAI Agents SDK and the Claude Agent SDK integrate closely with their own platforms.One vendor’s runtime sits at the core of the company agent.
Agent as configurationA managed console defines the agent. It is the fastest route to something running.The agent sits outside source control and review, unlike every other workload.

My default

LangGraph, shipped as a reviewed container image

  • The agent is code in a repository, built in continuous integration and deployed as an immutable version. Rollback points the endpoint at the previous version.
  • A managed agent runtime can isolate each session. Kubernetes is the alternative when you already run it.
  • The approval ledger is a business record in its own table. It is never session state.
  • Every action is tracked as planned, authorised, submitted, confirmed, failed or uncertain. An uncertain write is checked downstream before anything is retried.
  • The agent prepares engineering tasks. It does not inherit engineering credentials.
  • Default profiles have no shell and no host file system.

04.7Building block

The engineering harness

What it is

The harness is everything around the model in a coding agent: instructions, skills, tools, hooks and permissions. A company harness is one shared, versioned set of those that works across Claude Code, Codex, Cursor and Kiro. Three parts do the work: a registry of approved assets, an installer that resolves them at pinned versions, and a renderer that writes each tool’s own files and checks them for drift.

Why it matters

Engineers choose different tools. If the guardrails live in one tool’s configuration, every other tool bypasses them.

AlternativesWays to standardise agent behaviour
OptionWhat you getWhat it costs you
Standardise on one toolOne configuration format and the deepest integration.Lock-in. Engineers work around it, and one vendor’s outage or price change reaches everyone.
Per-team configurationFreedom for each team.Drift. Guardrails differ per repository and onboarding is repeated.
A shared harnessMy defaultOne source of skills, rules, hooks and MCP configuration, rendered into each tool’s format. Versioned and reviewable.You maintain adapters as tool formats change.
Open conventions onlyPortable files such as AGENTS.md, skills and MCP, with no build step.Coverage differs per tool. Hooks and permissions are not standardised.
Install from a developer portalOne place to browse skills, owners and how assets relate. Backstage is the usual choice.A catalogue holds pointers. Installing from search results cannot guarantee exact, pinned content.

My default

A shared harness: registry, installer and renderer

  • Assets resolve from the registry at a pinned reference, never from a discovery page.
  • One source record produces two projections: the registry for installs and the catalogue for browsing.
  • Generated files are recorded with hashes, so drift is detected and hand edits survive.
  • Open conventions carry instructions, skills and tools. Thin adapters cover hooks and permissions.
  • All model traffic goes through the gateway, so the guardrails hold whichever tool an engineer opens.

Sources, checked 28 September 2026

04.8Building block

Tracing

What it is

A trace records one request as a tree of spans: the prompt, the retrievals, the tool calls, the model calls, and the tokens, latency and cost of each.

Why it matters

An agent failure is rarely one bad call. You need the whole tree to debug it, to attribute cost, to build evaluation data from production, and to answer an auditor.

AlternativesTracing options
OptionWhat you getWhat it costs you
MLflow TracingMy defaultOpen source and OpenTelemetry-compatible. Traces, evaluation runs and the model registry live in one tool. Self-hosted, or managed on Databricks.The self-hosted interface and access control are plainer than dedicated products. The richest experience is on Databricks.
LangfuseOpen source, self-hosted or cloud. Strong prompt management, sessions, annotation queues and evals.Self-hosting means running Postgres, ClickHouse, Redis and object storage. It joined ClickHouse in January 2026 and says it stays open source.
OpenTelemetry-nativeAI spans flow through the telemetry pipeline and backend you already run. No new vendor.The generative-AI semantic conventions are still at Development status, and general backends have no prompt views or eval workflow. You build those.
Vendor tools (LangSmith, Datadog, Arize, Braintrust)Polished and quick to adopt, with evaluation built in.Prompts and outputs leave your boundary unless you buy a self-hosted or hybrid enterprise plan, and Datadog offers none. Cost grows with the traces, spans or data you send.

My default

MLflow Tracing, fed over OpenTelemetry

  • One system holds traces, evaluation runs and models, so an eval can cite the trace it scored.
  • Agents export standard OpenTelemetry, so the backend stays replaceable and the agent framework can change.
  • The trace store holds prompts, tool arguments and retrieved context. It sits behind authentication, on an internal route, with traces separated per agent.
  • Approval and security records are kept in their own store. Sampled traces are not the audit trail.
  • I would choose Langfuse when a team needs prompt management and annotation queues more than lakehouse integration.

04.9Building block

Evaluation

What it is

Evaluation is two systems with two jobs. Controlled evals test whether an agent completes representative tasks before a release. Monitoring watches how real interactions behave after it. The methods below feed one or the other.

Why it matters

Without evals, a prompt or model change is an opinion. With them, a change can be gated like code, and a person can approve a release on evidence.

AlternativesEvaluation methods
OptionWhat you getWhat it costs you
Controlled task evalsMy defaultA versioned suite of representative tasks, run on demand and in continuous integration. Inspect AI is an open-source framework for this.Suites need owners and held-out cases. A suite built only from the failures that prompted a change proves little.
Deterministic checksMy defaultExact match, schema validity or passing tests. Cheap and repeatable.They cover only the cases someone thought to write.
Large language model as judgeScores open-ended output for faithfulness, relevance and tone, at scale.The judge must be calibrated against human labels. It costs tokens and can drift when the judge model changes.
Human reviewGround truth for nuance, and the labels that calibrate a judge.Slow, expensive, and inconsistent between reviewers without a rubric.
Monitoring of production tracesQuality signals, latency and cost from real use.It reports after the user has seen the result.
Regression gatesMy defaultA merge or rollout is blocked when scores fall below the threshold.Noisy evals make flaky gates. Each threshold needs an owner.

My default

Controlled evals gate the release, production traces watch it

  • Neither system approves a release. A named person reviews the evidence and accepts the remaining risk.
  • An agent cannot approve its own release.
  • A release is bound to its image, model route, prompt and skill versions, retrieval configuration and eval result.
  • Suites cover task quality, source fidelity, unsupported claims, privacy leakage, prompt injection and tool misuse.
  • Corrections and incidents become new test cases after an owner reviews them.

Sources, checked 28 September 2026

04.10Building block

Open-source model serving

What it is

Running open-weight models, or your own fine-tuned models, on infrastructure you control. They sit behind the same gateway as hosted models, so callers cannot tell the difference.

Why it matters

Some workloads cannot leave the boundary. Some are narrow, high-volume tasks where a small fine-tuned model is cheaper and faster. Some need a model pinned for reproducibility.

AlternativesModel serving options
OptionWhat you getWhat it costs you
vLLMMy defaultHigh-throughput serving with continuous batching, an OpenAI-compatible server, adapter serving for fine-tunes, and wide model support.You own the graphics processors, autoscaling, upgrades and capacity planning.
Text Generation InferenceHugging Face’s serving stack. Mature and well documented.In maintenance mode since December 2025, and its repository archived since March 2026. Hugging Face points new deployments to vLLM or SGLang. Keep what runs; do not start here.
OllamaMy defaultOne command runs a quantised model locally, behind a subset of the OpenAI API. Suits development, demos and offline work.It is not built for multi-tenant production throughput.
Amazon SageMaker endpointsManaged hosting with autoscaling, identity and private networking, for your own container.Real-time endpoints bill per instance-hour whether they are busy or idle. Serverless inference scales to zero, with less control of the hardware.
Amazon Bedrock Custom Model ImportBring fine-tuned weights and call them through the Bedrock API. No endpoints to manage.Only supported architectures. A copy scales to zero after five idle minutes, so the next call waits for a cold start. Less control of the runtime.

My default

Ollama for development, vLLM in production

  • Both expose an OpenAI-compatible API, so moving from laptop to cluster is mostly a change of URL.
  • Every self-served model is registered in the gateway and inherits its budgets, guardrails and tracing.
  • Open weights do not always mean an open licence. The licence is checked before the route is approved.
  • Managed import suits low or irregular volume on a supported architecture.
  • Hosted models stay the default. I self-serve for a data boundary, a cost case or a latency case.

04.11Building block

Governance, retention and security

What it is

Governance decides what the platform may do and what evidence permits it: who can call which model and tool with which data, what is kept and for how long, and what happens when something goes wrong. It follows each use case through its life: define it, assess its impact, test it, approve its release, monitor it, then change or retire it.

Why it matters

In a regulated setting the model will fail at some point. What matters is whether the failure is contained, visible and explainable afterwards, and whether data can be withdrawn when permission ends.

AlternativesGovernance postures
OptionWhat you getWhat it costs you
Policy documents and trainingCheap and quick to publish.Nothing is enforced and there is no evidence of compliance.
Block by defaultLow exposure on paper.Usage moves to personal accounts, where nothing is visible.
Controls enforced in the gateway, the harness and the sourcesMy defaultIdentity, allow-lists, guardrails, budgets and logging defined as code. Each control produces evidence.Platform engineering effort, and guardrail false positives that need tuning.
Guardrail servicesReady-made detectors for personal data, harmful content and prompt injection.Extra latency and cost per call. Detection is probabilistic, and you still decide what a hit does.
A separate governance platformInventories, assessments and approvals in one product.One more system of record. I link the records the catalogue and the delivery pipeline already hold.

My default

Enforced controls, designed around failure

  • Identity, permissions and approvals are enforced by code outside the model. An instruction in a prompt is not a control.
  • An approval is bound to the exact action, its parameters and an expiry. A changed parameter invalidates it.
  • Revoking access and deleting copies are separate operations, with separate deadlines and evidence. Deletion follows lineage through indexes, embeddings, caches, checkpoints and backups.
  • Every class of data has an owner, a purpose and a retention schedule before it is collected. A class with no schedule is not collected.
  • An operator can disable a model route, a tool or a whole service, and stop pending work.
  • Retrieved and tool-returned content is treated as untrusted input.
  • Controls map to public frameworks such as the NIST AI Risk Management Framework and the 2026 OWASP Top 10 for LLM Applications.

04.12Building block

The AI Development Life Cycle

What it is

The delivery workflow defines how work moves from intent to running software when agents write most of the code: what gets specified, where a person approves, and what must pass before a merge. I treat it as a software factory with a contract at each handover.

Why it matters

Agents produce code faster than people can review it. The workflow decides where human attention is spent, and it keeps an agent from acquiring the right to merge or deploy.

AlternativesDelivery workflows
OptionWhat you getWhat it costs you
AWS AI-DLC (aidlc-workflows)Open-source AI-DLC workflows from AWS Labs. Five phases, from initialisation to operation, with a human approval gate after each stage and traceable artefacts.Heavy for small changes. I am exploring it. It earns a place only where it adds something the existing gates lack.
Spec-driven toolsRequirements, design and tasks are written before code.Effort up front, and specs that drift from the code.
Gated pipelineMy defaultThe agent works freely on a branch. A gate then runs review, tests, lint and documentation checks before the push, and stops on findings that need a person.Problems are caught after the code is written. Each change pays pipeline time and needs a clear statement of intent.
Pull-request review aloneFamiliar, with nothing new to run.Human review becomes the bottleneck, and quality depends on reviewer stamina.

My default

A gated pipeline in the style of no-mistakes

  • Enforcement sits outside the agent, so the model never marks its own work.
  • An agent’s task ends at a draft pull request. A deterministic publisher with its own scoped identity opens it.
  • Model review is advisory. Executable checks and a human reviewer decide.
  • The build produces one immutable image with its evidence. A separate, reviewed change selects that image for an environment.
  • Merging and deploying stay human decisions.

Sources, checked 28 September 2026

05Data platform

The data platform underneath.

An AI platform is only as good as the data it can reach and the data it can be tested on. I build the lakehouse first and treat models as one more consumer of it.

My background is in data platforms for financial services: Kafka, Spark, Snowflake, Databricks and Airflow. I hold the Databricks Data Engineer Associate and Generative AI Engineer Associate certifications.

Lakehouse
One storage layer for tables, files and AI artefacts, in an open table format and behind one catalogue. Data is refined in stages: raw (bronze), cleaned (silver) and ready to use (gold).
Data quality
Expectations are written as code and run inside the pipeline. Failing rows are quarantined and freshness and volume are monitored. A model evaluated on bad data reports a false pass.
Feeding models and evaluation
Gold tables become curated context, retrieval indexes, fine-tuning sets and golden evaluation sets. Traces flow back in as ordinary tables, so production behaviour becomes new evaluation cases.
  1. Sources
  2. Bronze: raw
  3. Silver: cleaned
  4. Quality gate
  5. Gold: ready to use

Gold feeds

  • Curated context
  • Retrieval indexes
  • Fine-tuning sets
  • Evaluation sets

Traces return to bronze as new data

How data reaches models and evaluation, and how traces return.
AlternativesWhere the lakehouse runs
OptionWhat you getWhat it costs you
DatabricksMy defaultUnity Catalog governs tables, volumes, functions, models and AI Search (vector) indexes. Managed MLflow puts traces and evals next to the data.Platform cost, and the workspace model is a commitment.
SnowflakeA strong warehouse with simple operations and mature sharing.Machine learning and open-format workflows are newer there than the SQL core.
Open stack on object storageOpen table formats with your choice of engine. The most control and the least lock-in.You integrate and operate the catalogue, engines and governance yourself.
Cloud-native warehouseIt fits the cloud account, identity and billing you already have.It ties the data platform to one cloud, and AI tooling varies by provider.

My default

Databricks for the lakehouse, open formats underneath

  • Data, models and traces share one governance model.
  • Open table formats keep the data readable by other engines.

06Operating model

Who owns what.

A platform fails when ownership is vague. These are the roles I set up, what each one decides, and the record each one keeps. A small organisation can combine roles, but the person who reviews a consequential action needs the authority to reject it.

RolesOwnership, decisions and records
RoleOwnsDecidesRecord it keeps
Business ownerThe use case: its intended users, its benefit and its acceptable limits.What the service is for, and what it must never be used for.The use-case record and its success measures.
AI service ownerThe application, its evals and its operating instructions.What changes, and when a change needs re-evaluation.The versioned service record and release evidence.
Knowledge and data ownersSources, their classification and their review.Which sources agents may read, and how long derived copies live.The source register, access scope and retention schedule.
Platform teamThe gateways, tracing, the shared harness and the paved road between them.How access, isolation, quotas and recovery are enforced.Configuration, verification results and runbooks.
Engineering towersDomain assets: skills, context packs, evals and guardrails.What good output looks like in their domain.Versioned assets, each with an eval.
Security, privacy and responsible-AI reviewersThe assessment of risk for each use case.Conditions and exceptions.The impact assessment and the review decision.
Release approverThe decision to release.Whether the remaining risk is acceptable for this version and scope.An approval bound to a version, a scope and a review date.
Service operatorThe running service.When to stop, roll back or escalate.Incident, corrective action and recovery records.
FinanceBudgets and chargeback.The budget for each team.Spend against budget, by team and by model.

Engineering towers are a draft definition. See the building block for the alternatives. Go to engineering towers

07Adoption and maturity

Where to start, and what comes next.

There are three honest places to start. Which one fits depends on where the pressure is today.

I start with the gateway. The other two paths need it within weeks, and it is the step that makes spend and usage visible.

A maturity roadmap

Each stage ends with evidence you can show. Move on when you can show it.

RoadmapSeven stages, each with evidence
StageWhat you addEvidence that you are there
1Ad hocIndividual keys and tools, chosen team by team.You have a list of who is using what.
2Governed accessThe gateway, identity, budgets and allow-lists.You can say who spent what, on which model, last month.
3ObservableTracing on every call, with a retention policy.You can reconstruct any request an auditor asks about.
4EvaluatedOwned evaluation suites and a named release approver.A model swap is a routing change backed by eval results.
5Shared context and toolsContext hubs, the knowledge path and tools registered in the gateway.The company agent and a coding tool retrieve the same approved revision, within the caller’s access.
6Agents that actDurable tasks, an approval ledger, and engineering agents that end at a draft pull request.Any agent action can be approved, stopped and reconstructed.
7OptimisedRouting by cost and quality, self-served models where they pay, and evals fed from production.Cost per accepted result falls while eval scores hold.

08Evaluation criteria

How to judge any option.

These are the questions behind every table on this page. In a regulated setting I weigh the first three most.

Data boundary
Where do prompts and outputs go, and who can read them? Can it run inside our network?
Identity and access
Does it use our single sign-on and our groups? Are keys issued per person or per team?
Auditability
Can we reconstruct one request from end to end, long after it happened?
Revocation and retention
Can we stop use at once, delete the copies on a schedule, and prove both?
Cost control
Can we cap spend per team before the money is spent?
Operability
Who patches it, what does an upgrade involve, and what happens when it is down?
Exit cost
What do we rewrite if we leave? Are the interfaces open standards?
Supplier risk
Who owns it, how is it licensed, and what changed in the last year?
Fit with the estate
Does it work with the cloud, the data platform and the tools we already run?

09Track two: smaller organisations

The same principles, at a smaller scale.

A small company cannot run an enterprise platform, and it should not try. GuideX, the travel marketplace I founded, runs agents in production with a small team. It keeps the principles and leaves out most of the machinery.

The worked example is the GuideX site-reliability estate: agents that triage errors, check payments and answer the team’s questions. A person approves each case before a coding agent picks it up.

Read the GuideX case study

Managed services first
Hosted model APIs, serverless containers that scale to zero, and a managed error tracker. There are no graphics processors to buy and no model servers to patch.
A thin gateway
A shared policy library that every agent embeds, in place of a gateway service in every call path. It enforces tool tiers, filters personal data, writes the audit record and checks the budget. One small orchestrator, with no model inside it, routes the requests that come from people.
One internal tool server
A single MCP server brokers every upstream system. It fails closed and accepts service identities only. The chat assistant holds no production credentials.
Lightweight tracing
One trace identifier travels from the mobile app through the backend to every agent action. Audit records live in the database the product already runs. A dedicated tracing platform can wait.
Lightweight evaluation
Every model output is validated against a schema, golden cases run in continuous integration, and a person approves anything irreversible. A low-cost model is the default. A stronger one is switched on per agent when measured failure rates justify it.
A pragmatic harness
One repository holds the agent instructions, hooks, commands and quality gates as plain files. A script copies them into each product repository and stamps the version, and a check reports drift. The rules that matter are gates that run before a commit and again in continuous integration.
Budgets and one switch
Each agent has a daily budget and an alert before the cap. One switch halts every agent at once.
A deterministic shell
Validation, authentication, tool tiers and the audit trail are plain code. Models reason only inside that shell, so flexibility never loosens a control.
Scale of the architecture
BronzeSilverGold8Data platform6Models5AI gateway4Context hub3MCP gateway2Harness1Surfaces7TracingGGovernance
The same layers for a small company. Dashed outlines are the parts it leaves out until it needs them.

Pick a layer

Point at a layer or tap it to see what it holds. The list below does the same from the keyboard.

ComparisonEnterprise and smaller organisation, block by block
Building blockEnterpriseSmaller organisationWhat stays the same
Model accessMany providers behind a self-hosted gateway.Hosted APIs, called through a shared policy library.Every call has an owner and a budget.
Tools for agentsCentral MCP servers behind a gateway, with single sign-on.One internal MCP server that fails closed.Tool access is a reviewed decision.
ContextA central hub, with towers that own domain assets.Instruction files in each repository, distributed from one source.Context is versioned and reviewed.
MemoryFour scopes in four stores, with reviewed promotion.Task state, an audit log and project notes kept as files.An observation is not guidance until someone reviews it.
Company agentA company agent on a managed runtime, with an approval ledger.A chat assistant that can read and draft only, behind the orchestrator.The agent that talks to people holds no production credentials.
OwnershipNamed owners for the platform, each service, each source and each release.The engineers who build the product.Every asset has a named owner.
TracingA tracing platform such as MLflow or Langfuse.Trace identifiers and audit records in the existing database.Any action can be reconstructed.
EvaluationGolden datasets, calibrated judges and regression gates.Schema checks, golden cases and human approval.A change is done when its checks pass.
Model servingSelf-served models where the boundary or the cost requires them.None. Hosted models only.The cheapest model that passes is the right one.
HarnessA shared harness with adapters for four coding tools.Plain files copied by a script, for one or two tools.Guardrails are enforced by code.
GovernanceControls as code, mapped to public frameworks and tested by a risk function.Tool tiers, a personal-data filter, budgets and one switch, in one library.Controls fail closed and leave evidence.
DeliveryA gated pipeline, with written plans for boundary changes.The same gates as scripts. An agent’s pull request passes the gates a person’s does.The agent never marks its own work.
Data platformA lakehouse.The product database and the cloud’s own logging. A warehouse comes when analysis needs one.Traces are data.

When to add the heavier parts

  • A second team starts building with models, and spend needs attributing.
  • A customer or a regulator asks for evidence you cannot produce from the audit records.
  • Engineers use more coding tools than one script can keep aligned.
  • A workload cannot leave your boundary, or its volume makes a hosted model the expensive choice.

10Working with me

What working with me looks like.

I have built this kind of platform inside a financial-services engineering organisation: a governed gateway to hosted models on AWS, a self-service developer portal, MLflow tracing, team budgets, authenticated MCP servers, and a shared harness across Claude Code, Cursor, Kiro and GitHub Copilot.

Read the enterprise AI platform case study

  1. 01Start with a conversationSixty minutes on what you run today, what constrains you, and where the risk sits.
  2. 02Write the options downYou get the alternatives, the trade-offs and a recommendation in plain language, with diagrams.
  3. 03Build a thin sliceA gateway, tracing and one eval gate, working for one team, before anything is scaled.
  4. 04Hand it overRunbooks, decisions and ownership sit with your engineers.

Book a 60-minute technical deep dive

Bring an architecture, a constraint or a decision you are stuck on. We work through it together.

Book the deep diveEmail [email protected]