Anthropic has started publishing measurements of how models get built — compute, inputs, process — not just what they can do. It's a new kind of AI disclosure.
Every conversation about AI progress runs on outputs. A model scores X on a benchmark. It solves Y percent of coding tasks. It matched experts on some hard problem. Outputs are what labs announce, what the press covers, and what enterprise buyers put on slides.
Anthropic has just published something different: a set of measurements about how models are made. Not what Claude can do, but what goes into producing it — the production process itself, with compute as an input and capability as an output. The company frames these metrics as complementary to capability evaluations, which it publishes separately through its Responsible Scaling Policy risk reports. The stated goal is to get better at correlating inputs with outputs, so that the pace of AI development becomes something you can observe rather than something you infer from press releases.
That sounds dry. I think it's one of the more consequential disclosure moves of the year, and for a reason that has nothing to do with Anthropic.
Why input metrics are different
Capability benchmarks answer a question about the present tense: is this model good at this task, right now? That is useful for a purchasing decision and almost useless for a planning decision.
Input metrics answer a different question: how fast is the underlying process moving, and what is driving the movement? If you can see that a given jump in capability took a certain amount of compute, a certain number of training runs, a certain iteration cadence, you start to have something that looks like a trend line instead of a series of surprises.
For anyone running an AI roadmap inside a real company, that distinction is the whole ballgame. The hardest thing about planning AI work in 2026 isn't picking a model. It's that the ground keeps moving underneath the plan. You scope a six-month project around what models can do in March, and by August the assumption that justified half the architecture is obsolete — or, just as disruptive, a capability you were counting on arriving hasn't arrived.
Output benchmarks don't help with that. They tell you where things stand. They don't tell you how fast the floor is rising.

The self-regulation problem this runs into
There's an obvious objection, and Nature published a version of it in the same news cycle: a piece arguing straightforwardly that AI companies can't be trusted to self-regulate.
The critique is fair on its own terms. Voluntary disclosure has structural limits. A lab chooses which measurements to publish, when, and with what framing. It can stop publishing the moment the numbers get inconvenient. Nothing about Anthropic's initiative changes the fact that the entity producing the measurement is the entity being measured.
But I'd argue the two things are less opposed than they look. Every regulatory regime that actually functions — financial reporting, drug trials, emissions standards — started with a fight about what to measure. The metric came first, then the argument about who verifies it, then the mandate. You cannot regulate a pace you have no vocabulary for describing.
Right now, AI oversight has extremely rich vocabulary for model behavior and almost none for model production. Regulators can ask whether a system is high-risk, whether outputs are labeled, whether an incident was disclosed. They have no standard way to ask how fast the provider's own development loop is accelerating. Publishing process measurements — even self-reported, even selectively — creates the categories that a mandate could later attach to.
That's my read, not a prediction. But if you were wondering why a lab would volunteer metrics about its own internal machinery, "shaping the measurement framework before someone else writes it" seems like a reasonable guess.
What to do with this as a buyer
You don't need to care about frontier-lab compute budgets to take something practical from this. The transferable idea is that process metrics predict better than outcome metrics — and that applies inside your organization as directly as it does inside a lab.
Most enterprise AI reporting I see is pure outcome measurement. Number of licenses. Monthly spend. Usage counts. A satisfaction score from a survey nobody wanted to fill out. All of it describes where you are and none of it describes how fast you're moving or why.
The process-side questions are the ones that actually forecast:
- How long does it take you to go from "we want to try this" to a governed pilot? If the answer is four months, your AI strategy is capped by your intake process, not by model capability.
- How long does a model swap take? When a better or cheaper model ships, can you move a production workload onto it in a week, or is it a quarter-long project? This is now a recurring cost, not a one-time one.
- How many of your AI systems can you actually name? Pace is meaningless without an inventory. If you can't enumerate what's running, you can't measure change in it.
- What's the gap between your fastest-moving team and your slowest? Aggregate maturity scores hide the variance that matters most.
And for vendors, this raises a question worth putting in your next diligence conversation: what do you publish about your own development process, and how often? A year ago that would have been a strange thing to ask. One frontier lab has now made it a reasonable one — and the labs that decline to answer are telling you something too.
The uncomfortable implication
If input metrics do turn out to correlate cleanly with capability jumps, the honest consequence is that the pace of change becomes more legible — and harder to be surprised by, but also harder to pretend away.
Plenty of organizations are quietly betting that AI capability plateaus long enough for their current architecture, vendor contracts, and governance model to hold. Measurable pace makes that bet explicit instead of implicit. You'd be able to look at the trend and say: our two-year plan assumes the floor rises this much and no more.
Maybe that's a bet you'd still take. But you'd be taking it with your eyes open, which is more than most AI roadmaps can currently claim.
