The Standardization Illusion: Why AI Inference Efficiency Can’t Be Standardized Across Labs In Regulation
During my time at the Securities and Exchange Commission following the 2008 U.S. financial crisis, I watched regulators — under intense pressure to prevent another systemic breakdown — endeavor to develop new policy in response.
Every major economic or technological shift provokes an understandable reaction: an urge to control intricate, rapidly evolving subjects through rigid administrative formulas.
We are seeing this pattern repeat today in AI policy.
In his policy essay “We Must Pace the Frontier” over the weekend, Anthropic CEO Dario Amodei made a sensible case for pragmatic oversight: specifically focusing on physical hardware supply chains, safety testing, and securing physical data centers. Controlling the tangible, physical "chokepoints" of AI is clear, enforceable, and grounded in reality.
Where policy risks going off the rails is when regulators try to audit data centers based on "inference efficiency.” Essentially, this means trying to measure how much useful computing work or intelligence an AI lab extracts from every unit of electricity.
Attempting to regulate AI on inference efficiency creates the same trap as Scope 3 carbon accounting. Scope 3 forces organizations to quantify indirect, value-chain emissions using fluid boundary assumptions and proxy estimates. This resulted in outputs that were endlessly complex, vulnerable to being “gamed” and extremely difficult to accurately audit.
Attempting to force a uniform efficiency standard on AI inference is a similar mistake.
Trying to enforce a single efficiency metric across different research labs fails for three practical reasons:
Hardware and Software Customization: Two different labs can run the exact same AI model on the exact same computer chip, but one lab’s custom software tweaks will make it draw drastically less power. To audit efficiency, regulators would have to inspect every lab’s secret internal code.
Query Complexity Varies Wildly: Answering a simple question like "What’s the capital of France?" takes a tiny fraction of energy compared to asking an AI to solve a multi-step math problem or analyze a complex contract. Since user requests change second-by-second, efficiency numbers fluctuate wildly.
The Target Never Stops Moving: AI engineers optimize their software daily. Efficiency techniques evolve far faster than any government reporting cycle.
Why Policy Must Stay Narrow
When regulations target fluid software performance, they invite gaming. Companies end up engineering their AI to pass artificial government tests rather than making it genuinely better or safer.
Just as financial regulators eventually learned to focus on basic, physical balance-sheet capital and reserves, AI regulators must stick to physical, clear baselines:
Metered Power (MW): Measure physical electricity drawn straight from the power grid.
Physical Chip Counts: Track and audit actual physical computer chips sitting in a server rack.
Cooling and Facility Size: Monitor physical energy and infrastructure limits.
Practical regulation works best when anchored to tangible realities. Regulators should focus on the physical buildings and hardware that power AI, leaving labs free to innovate on the software without getting trapped in a web of unauditable compliance metrics.