Over the past month I published a series of posts on what breaks in enterprise AI after the pilot works. A $440,000 consulting report delivered with fabricated citations no review process caught. A crafted email that talked Microsoft 365 Copilot into leaking private data — zero clicks involved. Shadow AI turning up in 43% of data breaches. Frontier labs admitting their own agents escaped their own sandboxes. The same routine task costing nine times more than it should because of configuration defaults nobody reviewed.
Writing that series, one thing became impossible to ignore: not one of these was a technology failure. In every case, the model did what models do. The failure was a decision that never got made — a default someone shipped, a permission nobody scoped, a check nobody owned.
That observation has a practical consequence. If the failures are decision failures, then the fix is not a better model, a bigger platform, or another policy document. The fix is a set of durable rules for making AI decisions when evidence is incomplete, pressure is high, and the demo looks great. Rules that hold up when someone pushes back.
There's a reason I frame these as first principles rather than best practices. Best practices are borrowed — they describe what worked somewhere else, in someone else's context, at a moment that has usually already passed. In a field where the technology reinvents itself every couple of quarters, tactics expire fast. First principles are what's left when you strip the tactics away: the small set of things that stay true across model generations, vendor cycles, and hype waves. A leadership team that has done the work of writing its own down gets three things that a policy binder never delivers — decisions that are consistent instead of personality-driven, the ability to delegate judgment without delegating chaos, and a way to say no (or yes) that survives being challenged. A team that hasn't done that work re-argues every AI decision from scratch, under pressure, every time.
These are the twelve first principles I developed to guide enterprise AI strategy. I use them as a decision filter: when a proposal arrives, I should be able to connect it to an outcome, an accountable owner, evidence, an authority grant, a risk boundary, and a review cadence. When I can't, one of these principles tells me what's missing. Each comes with why it matters, what failure looks like, and the one question that tests it.
01Outcomes before activity
AI demos are unusually cheap to build and unusually convincing to watch. That combination means visible activity can outrun actual performance for a long time — a growing portfolio of impressive pilots with no accountable value anywhere in it. The discipline is to translate every proposal into an owned business outcome with a baseline and a measure, before work starts. Not "we deployed a copilot" — "we cut claim-handling time 18%, and here's the owner who signed up for that number."
02Accountability must match authority
The fastest way to burn an AI leader — or a program — is responsibility without decision rights. If you are accountable for outcomes but can't approve, pause, or stop the work, you don't run the program; you narrate it. Authority needs to be written down: decision rights, budget visibility, escalation routes, named decision owners. And it needs to be tested on a real, survivable hard decision before it's tested on an existential one.
03Evidence before conclusions
AI evidence decays. A model update, a data change, or a shift in how people actually use the system can invalidate a result that was true last quarter. Most organizations treat an evaluation as a one-time gate: passed it in March, cite it forever. The discipline is to capture the sources, assumptions, baselines, and economics behind every claim — and to re-earn confidence after material change, not just re-assert it. Deck claims that outlive their evidence are how programs end up confidently wrong.
04Value before scale
Scaling an unproven use case doesn't multiply value — it multiplies cost, complexity, and eventual disappointment. The pressure to scale usually arrives before the proof does, because platform investment is easier to defend when it's serving "the whole enterprise." Set a measurable value threshold and fund the next stage only when the current one clears it. If the pilots can't show material benefit, the platform serving them is overhead, not infrastructure.
05Bound risk before production autonomy
Autonomy changes the blast radius from an incorrect answer to an incorrect action. The labs themselves cannot fully keep their agents in the box — OpenAI, Anthropic, and Moonshot have all documented models escaping test environments. So the question inside your company was never "which model is safest." It's which permissions, human involvement, failure modes, monitoring, and stop conditions get defined before an agent touches production. If nobody has agreed what the agent may do and when it must stop, the deployment decision hasn't actually been made — it's just been defaulted.
06Put controls where decisions occur
A policy document cannot compensate for a control that's absent from the path of action. Shadow AI is the proof: usage policies everywhere, and 43% of breaches involving tools no one sanctioned. Controls work when they live in intake, identity, tooling, deployment, logging, and monitoring — the places where a risky action can actually be blocked and evidence actually accumulates. If you can't point to where in the workflow a risky action gets stopped, you don't have a control. You have a hope.
07Share platforms, distribute ownership
Enterprise AI creates a real tension between centralized control and local context, and neither extreme is safe by default. Full centralization produces a center of excellence that becomes a queue — owning neither adoption nor results. Full federation produces incompatible controls and duplicated spend. The working answer is a thin shared platform for the capabilities that benefit from consistency — gateways, evaluation, standards — with value and risk owned by the business units closest to the work.
08Adoption is part of the use case
Value appears only when a safe capability changes how work is actually performed. Licenses and launch announcements measure procurement, not adoption — and the quiet return to the old workflow is the most common failure mode nobody reports. Workflow redesign, manager reinforcement, enablement, and trust-building are not change-management garnish added after deployment; they are part of the use case itself, and they belong in its plan and its budget from the start.
09Distinguish reversible from irreversible bets
Not every AI decision deserves the same speed. A bounded experiment you can unwind in a week should move fast — the learning is worth more than the risk. A one-way commitment — a platform standardization, a vendor lock-in, a workflow you can't restore — needs stronger proof and explicit exit criteria before you walk through the door. Organizations fail in both directions: freezing safe learning behind enterprise process, or rushing irreversible commitments because the demo was compelling.
10Transparency is safer than optimism
AI fails silently — quality drift, bias, cost creep — and the person who championed the program has a built-in incentive to soften bad news about it. That combination is dangerous. The discipline is to surface limitations, dissent, incidents, and unresolved assumptions early, especially when they threaten the program narrative. Deloitte's fabricated-citations episode is instructive: the failure was expensive, but the disclosure and refund preserved something. Bad news that reaches executives after the customer or the board already has it costs far more than the underlying failure.
11Data conditions govern what is possible
No model choice makes unavailable, untrusted, or unauthorized data a safe foundation. When an AI initiative fails on inaccessible, stale, biased, or improperly permissioned data, the model takes the blame — but the data condition was knowable in advance. Use data readiness, lineage, permissioning, and exposure evidence to rule a use case ready, constrained, or not suitable before committing to it. "Constrained" and "not yet" are legitimate answers that save quarters of wasted effort.
12Define how you will know it works
AI quality is probabilistic and context-dependent, and a fluent output is not evidence of fitness. A system is ready for production when there is a representative evaluation set, a rubric, a passing threshold, human review where it matters, and a trigger that catches regression afterward. Declared successful "because it demos well" is how fabricated citations end up in delivered reports. If no one can state the passing standard, the system hasn't passed anything.
When principles conflict
Principles that never tension against each other aren't doing any work. Three conflicts come up constantly:
Value before scale vs. reversible bets. A small pilot may not produce scale-level proof. Fund the smallest representative test that could falsify the value case, then scale only when the remaining uncertainty is material enough to justify the next bet.
Speed vs. bounded risk. Move quickly on reversible, low-blast-radius work. Slow down only the permissions, exposure, and autonomy that can create irreversible harm. Discipline applied indiscriminately is just slowness.
Central standards vs. local ownership. Centralize the controls and capabilities that benefit from consistency. Keep outcome ownership and workflow context with the business. When a dispute arises, ask which failure is worse: an inconsistent control, or an unowned outcome.
The doctrine in practice
Here's what this looks like under pressure. An executive wants an autonomous customer-service agent deployed quickly, because a competitor just announced one. The undisciplined responses are the two easy ones: "no" (institutional hesitation dressed as prudence) or "yes" (a high-blast-radius commitment made on competitive anxiety).
The doctrine produces a third answer. Translate the request into a customer outcome with an accountable business owner (01). Classify it: autonomous action in front of customers is high blast radius (05), but a bounded pilot is a two-way door (09). Check authority, data permissions, and budget (02, 11). Then move fast on a genuinely bounded version: a representative but low-risk queue, an evaluation threshold defined before launch (12), human escalation and stop conditions (05), cost instrumentation (03), and a dated decision on scaling (04). That is disciplined speed — faster than the "no," and vastly safer than the "yes."
When pressure rises, the thing to return to is the operating hierarchy — each layer doing a job the others can't:
Principles shape the judgment. The mandate enables the decision. The process makes it repeatable. The record makes accountability visible.
None of this requires believing anything unusual about AI. It requires believing something unfashionable about organizations: that most AI failures are ordinary governance failures wearing a new costume, and that the organizations that get durable value from AI will be the ones that made deliberate decisions where everyone else shipped defaults.
