CMS's top technology official told FedScoop the agency is moving away from counting AI usage and toward measuring outcomes. The FCC, separately, is building a scorecard that will grade phone companies on how well they block spam calls, not on whether they installed the required blocking technology. Two agencies, two different problems, the same underlying correction: activity metrics were standing in for results, and someone finally asked for the results instead.
From Adoption to Impact
Government technology programs have long been graded on activity. Did the agency deploy the tool. Did staff log in. Did the vendor deliver the contracted feature. CMS's move away from AI-usage metrics is a quiet admission that those numbers were never the point, they were a stand-in for the point, and stand-ins get gamed.
An agency can report high AI adoption while the underlying work gets no faster. A phone carrier can report that it installed call-blocking software while spam calls keep landing anyway. The FCC's scorecard is built to close exactly that gap: it grades carriers on blocking performance, on calls that get through or don't, rather than on whether a compliance box got checked. That distinction matters far beyond telecom and health tech. It's the same distinction between a media plan that reports impressions delivered and one that reports what those impressions actually moved.
For years, the safest number in a report was the one that was easiest to produce. Logins, seats provisioned, dashboards configured, campaigns launched on schedule. Those numbers are real, and they're not nothing, but they were never proof of value. They were proxies, and everyone quietly agreed to treat the proxy as the thing itself because measuring the actual outcome was harder, slower, or more embarrassing when the answer came back weak.
The Scorecard as a Public Ranking
What makes the FCC's approach worth watching isn't the robocall problem itself, it's the mechanism. A scorecard puts every carrier into the same public ranking on the same outcome measure. There's no separate story a laggard can tell about its roadmap or its good-faith effort. The number is the number, and it's visible to competitors and customers alike.
That's a different kind of pressure than a checklist creates. A checklist rewards documentation: show the auditor the policy, the log, the sign-off, and you pass. A scorecard rewards results, and does it in public, where a customer can compare one carrier's blocking rate against another's without needing to understand the underlying technology at all. Once one agency proves that model works for something as measurable as call blocking, it becomes harder for any buyer, public or private, to keep accepting activity reports as proof of value.
The same instinct shows up elsewhere in the current federal record, even where it isn't framed as a scorecard. The fight over the VA's electronic health records contract, which prompted a House committee to vote to subpoena Oracle after a difficult hearing, is fundamentally a dispute about whether a multibillion-dollar system delivers the outcome it was built for, not whether it was installed on schedule. The FTC's extended comment period on personalized pricing is a version of the same question aimed at algorithms: not whether a pricing tool exists, but what it actually does to the price a person pays. Installation was never the deliverable. It just used to be the easiest thing to check.
Verification Is Becoming the Product
There's a second thread running through the same batch of federal activity, and it's worth naming separately: agencies are getting stricter about verifying claims before they trust them, not just after. The FTC's warning to consumers about scanning QR codes parked in public places is a small, almost throwaway example, but it's the same logic in miniature. Don't trust the thing because it looks official. Verify it does what it claims before you act on it. Colorado's new mobile driver's license for TSA checkpoints works because it's built to be verified at the point of use, not because it was issued.
Marketing and agency work runs on the opposite habit more often than anyone wants to admit. A tool gets adopted because a vendor says it works, a metric gets reported because it's the one the platform surfaces first, a campaign gets called a success because the deliverables shipped on time. The scorecard model asks a blunter question at every one of those steps: verified against what, and by whom. Telecom industry groups are currently in the Sixth Circuit arguing against the FCC's data breach rules, which is itself a fight over how much verification a regulator is entitled to demand before it accepts a company's account of its own security. That argument is about telecom rules specifically, but the shape of it, how much proof is owed before a claim is trusted, is going to keep showing up wherever agencies decide activity reports aren't enough anymore.
What to Change Before the Next Pitch
None of this means usage metrics disappear. Adoption is still worth knowing, the same way installation is still worth knowing. But it stops being the headline slide, and it stops being sufficient on its own. The practical shift for anyone pitching or renewing work is to move the outcome question earlier in the conversation, before the client has to ask it.
- Replace activity slides, logins, seats, impressions, with the outcome those activities were supposed to cause, and be honest if the connection is thin.
- Ask what the client's own scorecard would measure if they built one today, then report against that instead of against your own dashboard.
- Treat any renewal conversation as a chance to show a before-and-after result, not a usage trend line.
- Where a claim can be independently verified, whether that's a call blocked, a lead converted, a cost avoided, build the reporting so the client can check it without taking your word for it.
The checklist asks whether you did the thing. The scorecard asks whether the thing worked.
There's a structural reason this pressure is building now rather than five years ago. A recent spending deal included provisions blocking political control of grants, which is a different fight but rests on the same premise as the CMS and FCC moves: that money and effort need to be judged against a standard set in advance, not against whoever is telling the story afterward. Debate over the Universal Service Fund's future without reforms runs on the same fuel, prominent Republicans questioning USF haven't been arguing that the fund does nothing, they've been arguing that nobody can currently prove what it does relative to what it costs. Outcome measurement is what ends that kind of argument, one way or the other.
Agencies that adjust their reporting now, before a client or regulator forces the switch, get to define what "outcome" means for their own work, and get to build the systems that make proving it easy. Agencies that wait will have someone else's scorecard handed to them, built around whatever the buyer decided to measure once they stopped trusting the checklist.