Anthropic gave outside researchers independent access to real Claude usage data — a first for a frontier lab, and a new benchmark for what AI vendors should tell you.
Every enterprise buying AI today is flying on vibes about one of the most important questions in the deal: how is this thing actually being used?
Vendors will tell you about capabilities. They'll tell you about benchmarks, context windows, uptime, and SOC 2 reports. What almost none of them will tell you — because they mostly don't share it — is what real people do with the product once it's deployed. Which tasks. Which failure modes. Which patterns of over-reliance. Which populations get worse results.
Anthropic just broke that pattern. The company handed a set of external academic partners access to real Claude usage data, with a deliberately narrow set of contractual review rights: user privacy, information that could help people violate usage policies, Anthropic's own confidential information, and research accuracy. Beyond that, the company says it had no say in the findings — and researchers are free to publish results "even if they are inconvenient for Anthropic."
That last clause is the whole story.
Why this is unusual
The standard model for studying how AI gets used is either a survey (people telling you what they think they do) or a vendor-authored report (the vendor telling you what it wants you to know). Both are cheap. Both are unreliable in the same direction: they flatter the tool.
Independent access to actual interaction data is a different category of evidence. It's the difference between asking employees whether they follow your AI policy and looking at the logs.
Anthropic has published early results from the program, including work from Carnegie Mellon's Social and Language Technologies Lab. But the specific findings matter less than the precedent. A frontier lab has now demonstrated that you can give outsiders real usage data, with privacy protections in place, and survive the results. Which means the excuse other vendors have been leaning on — that it's impossible, or too risky, or a privacy non-starter — just got a lot weaker.

The gap this exposes in your own house
Here's the uncomfortable mirror. If a lab with hundreds of millions of users can characterize how its product is used well enough to hand the data to skeptical academics, what could you say about AI usage inside your own organization?
For most companies, the honest answer is close to nothing. They can name their approved tools. They might know seat counts. Some know monthly spend. Very few can answer questions like:
- Which business processes now have an AI system in the critical path?
- What percentage of licensed seats are actually active — and which teams pay for tools nobody opens?
- Where are people using AI for decisions we said required a human?
- Which of our AI-touched workflows have a documented failure mode and an owner?
Seat counts are not usage data. Spend is not adoption. A tool inventory is not a map of where AI has become load-bearing in how work actually gets done. The Anthropic program is interesting partly because it makes that distinction impossible to ignore: the lab is measuring behavior, and most of its customers are measuring invoices.
Turn this into a procurement question
The practical takeaway isn't "read the CMU paper." It's that vendor usage transparency is now a thing you can ask for, with a live example to point at.
Three questions worth adding to your next AI vendor review:
- What usage telemetry do we get, and at what granularity? Not aggregate dashboards — can you see task types, abandonment, override rates, and where your users hit refusals or hallucinations? If the vendor's answer is a monthly PDF, you are buying a black box.
- Has any independent party examined how this system behaves in production? Internal red-teaming is table stakes. External, publishable scrutiny is the differentiator, and a small but growing number of vendors can now claim it.
- What happens to inconvenient findings? Ask specifically whether the vendor has ever published or permitted publication of a result that made the product look worse. The answer tells you a lot about what their safety claims are worth.
None of these require you to be a compliance specialist. They're the same questions a good buyer asks about any critical system — they've just been oddly absent from AI purchasing.
The maturity signal underneath
There's a broader pattern here that's worth naming. Over the past two years, AI governance has largely been an exercise in documentation: policies, registers, risk classifications, attestations. Useful, but paper-thin if nobody knows what's happening in production.
The frontier is shifting toward evidence. Testing frameworks that produce inspectable artifacts. Evaluations designed so scores can be verified rather than asserted. And now usage data that outsiders can actually examine. The common thread is a move from claims to observability — from telling people your AI is fine to being able to demonstrate it.
Organizations that treat AI maturity as a matter of observability will find this easy to adapt to. Organizations that treat it as a matter of policy documents will keep discovering problems from customers, auditors, and press releases.
What to do this quarter
Start small and concrete. Pick your three highest-consequence AI deployments — the ones touching customers, money, or people decisions. For each, write down: who owns it, what it's actually used for (not what it was bought for), how you'd know if it degraded, and what usage data the vendor gives you today.
If you can't fill in the last two columns, that's not a governance failure. It's an inventory failure, and it's fixable. But you can't prioritize what you can't see — and right now the labs have better visibility into how your teams use AI than you do.
Anthropic just proved that visibility can be shared. Ask your vendors why theirs isn't.
