In July 2026 the Census Bureau asked a sample drawn from 1.2 million American businesses whether they had used AI in any business function in the previous two weeks. 21.8% said yes.
In the same economy, the Federal Reserve's Small Business Credit Survey put the figure at 46% of small employer firms. The Atlanta Fed's Survey of Business Uncertainty put it at 69.4% equal-weighted. McKinsey put it at 88% of organizations.
None of these is a lie. They are all measuring something real. They just are not measuring the same thing, and once you see what separates them you stop looking for the correct number and start working out how to get one for your own shop. That is the whole argument of this post: no published efficiency figure transfers to your business, the reason is visible in the sampling frames, and running your own measurement costs about eighty dollars a month at published prices.
The spread, with the sampling frame printed next to it
Here is the same question, roughly the same window, five answers.
| Source | Figure | Who was actually in the sample |
|---|---|---|
| Census BTOS, cycle covering 13 to 26 July 2026 | 21.8% of US businesses used AI in any business function in the prior two weeks | Probability sample from a 1.2 million business frame. About 24,500 responses. 12% unit response rate, non-response adjusted. Weighted by firm. |
| Fed Small Business Credit Survey, fielded Sept to Nov 2025 | 46% of small employer firms currently use AI | 6,525 responses from firms with 1 to 499 employees. The survey's own methodology says it "is not a random sample" and warns about convenience-sample bias. |
| Atlanta Fed Survey of Business Uncertainty, November 2025, equal-weighted | 69.4% for AI, 46.4% for large language models | Phone-recruited stratified firm panel. Equal weight per firm in this cut. |
| Same survey, employment-weighted | 78% of the labor force works at a firm that has adopted AI | Weighted by headcount, so one large adopter outweighs a thousand small non-adopters. |
| McKinsey State of AI, fielded 25 June to 29 July 2025 | 88% of organizations use AI in at least one business function | 1,993 online respondents across 105 nations, weighted by each nation's share of global GDP. 38% came from organizations above $1 billion in revenue. |
Three things fall out of that table immediately.
The weighting unit does most of the work. Census weights by firm, so a two-person landscaping company counts the same as a bank. The Atlanta Fed also publishes an employment-weighted cut, and the same survey moves from 69.4% to 78% just by switching the unit. Census researchers found the same effect in their own data: 18% of firms used AI in a business function, rising to 32% on an employment-weighted basis, and reaching 50% to 60% for very large firms in Information, Professional Services and Finance.
Probability sampling is not a formality. The Small Business Credit Survey produces a number more than twice the Census figure for a similar population, and it says in its own methodology that it is not a random sample. That is an honest disclosure and it is also the entire explanation. Firms that answer an optional survey about AI are firms with something to say about AI.
Nobody is standing behind the biggest number. Stanford's AI Index republishes McKinsey's 88% and attaches its own caveat: the results "are self-reported and should be viewed as directional rather than comprehensive." When the index that carries a statistic tells you not to treat it as a measurement, believe it.
This is not new. A Federal Reserve staff review cites an earlier survey of 16 studies that put work-related AI adoption anywhere between about 5% and 40% as of mid-2024. An eight-fold spread, same phenomenon, same year.
The pair that runs in opposite directions in the same weeks
The starkest example is not about adoption. It is about return.
MIT's Project NANDA published a report in July 2025 stating that "95% of organizations are getting zero return" from generative AI, with just 5% of integrated pilots extracting millions in value. That figure went everywhere.
Wharton's Human-AI Research group, with GBK Collective, fielded its enterprise study from 26 June to 11 July 2025 and reported that "nearly three-quarters already see positive ROI."
Same country. Overlapping weeks. Opposite conclusions. The resolution is in the screening criteria, and you should read them side by side.
Wharton's respondents had to be senior decision makers at a "U.S.-based enterprise commercial organization (1000+ employees and >$50 million revenue)." No small business was eligible to answer. Not one. And inside that sample, the belief tracks distance from the work: 81% of VP-and-above respondents believed ROI was positive, against 69% of mid-managers.
MIT NANDA's base was 300-plus publicly disclosed initiatives reviewed, 52 structured interviews, and 153 survey responses collected at four industry conferences. The report carries its own limitations note saying the figures are "directionally accurate based on individual interviews rather than official company reporting" and that "success definitions may differ across organizations." MIT no longer hosts the report at its original address; when I went looking on 21 August 2026, that address served the Media Lab group's overview page instead of the PDF, so I read the wording from a third-party mirror.
So the most quoted "AI does not pay off" statistic in the world rests on 153 conference attendees and a definitional shrug, and it has been quietly removed from its publisher's server. There is a much better version of the same finding.
The Atlanta Fed, working with the Bank of England and academic co-authors, surveyed almost 6,000 CFOs, CEOs and executives from stratified firm samples across the US, UK, Germany and Australia. More than 80% of firms reported no impact from AI on either employment or productivity over the past three years. In the US specifically, 89% reported no employment impact and 91% reported no labor productivity impact. The estimated average realized productivity gain across all firms was about 0.29%.
That is the same headline with roughly forty times the sample and a real sampling frame. If you only carry one statistic about AI not showing up in the numbers yet, carry that one.
The same executives forecast that AI will raise productivity 1.4% and cut employment 0.7% over the next three years. Employees, surveyed separately, predicted a 0.5% employment increase. Both groups are guessing about the future while sitting on three years of measured near-zero.
The size gradient is not the ramp you think it is
The story everyone tells is that AI adoption rises smoothly with company size and small business is being left behind. The Census employment-size file for July 2026 says something more interesting.
| Employees | Used AI in the last two weeks |
|---|---|
| 1 to 4 | 21.9% |
| 5 to 9 | 20.0% |
| 10 to 19 | 20.2% |
| 20 to 49 | 22.4% |
| 50 to 99 | 27.8% |
| 100 to 249 | 30.3% |
| 250 or more | 41.5% |
The gradient does not exist below 50 employees. Solo operators and micro firms use AI at essentially the same rate as firms with 20 to 49 people, and at a higher rate than firms with 5 to 19. The trough is the 5-to-19 band, not the bottom. The ramp only starts above 50.
Two honest caveats. First, the same agency published a story in May 2026 saying that "less than 20% of firms with four or fewer employees reported using AI" and that between December 2025 and May 2026 use "didn't change significantly among firms with fewer than 20 employees." That was an earlier cycle. The August 2026 file has the 1-to-4 band at 21.9%. Both are Census, they are three months apart, and the smallest firms moved. Print both rather than picking the flattering one.
Second, the expectations gap is wider than the usage gap: 50.6% of firms with 250 or more employees expect to be using AI in the next six months, against 25.8% of firms with 1 to 4.
And there is a series break sitting underneath all of this. On 17 November 2025 Census changed the core question from AI use "in producing goods or services" to AI use "in any business function." Under the old wording the firm-weighted rate had crawled from 3.5% to about 10% over two years. Post-revision it reads around 18% and climbing. Census's own working paper attributes the jump to three things at once: the wording change, continued adoption during the federal funding lapse when no data was collected, and the introduction of the AI supplement itself, which likely raised respondents' propensity to report use. Anyone drawing a hockey stick through that point without a break marker is manufacturing it.
What "AI adoption" actually means for the median adopter
Two numbers from the Census AI supplement explain more about this market than any adoption chart.
64.3% of AI-using businesses made no changes at all to adopt it. No staff training, no hardware or software purchase, no new workflow, no vendor. For comparison, 15.4% developed new workflows, 15.0% trained current staff, 8.5% purchased cloud services, 3.8% used a vendor or consultant, and 1.3% hired anyone with AI training.
95.7% reported no change in total employment from AI use over the prior six months. 2.3% reported an increase and 2.0% a decrease.
Read those together and "adoption" resolves to something very ordinary: somebody opened a chatbot. That is also why the surveys disagree so violently. When the threshold for a yes is approximately zero effort, small differences in question wording and respondent enthusiasm move the answer from 21.8% to 46% between two surveys of a similar population.
The supplement reinforces it. Only 10.1% of AI-using firms used AI to perform a task previously done by an employee. 43.7% used it to supplement or enhance a task an employee already performed. 51.5% picked none of the above.
And the reason non-adopters give is not cost. Asked why they have no plans to use AI in the next six months, 61.6% said "AI is not applicable to this business." Only 6.9% said it was too expensive. Only 3.2% said a previous attempt had not met expectations. "Not applicable" falls steadily with size, from 63.3% at 1 to 4 employees to 42.2% at 250 or more, which reads less like a capability gap than a perception gap.
If you are a small-business owner who has decided AI is not for you, you are in the largest single category of non-adopters, and you got there for reasons of fit rather than price. That is a defensible position. It is also worth testing once, cheaply, which is the back half of this post.
The gap between what firms say and what their staff do
Two Census probability samples, same statistical system, months apart in the same year, asking about the same thing from two ends.
The business survey says 21.8% of firms use AI. The March 2026 Household Trends and Outlook Pulse Survey asked workers directly, and 55% of US workers said they had used AI on the job for at least one of eleven tasks. That is roughly two and a half times the share of firms that say they use it.
The Census working paper addresses this head on and notes that "worker task use sometimes occurs without formal firm-level adoption." That has a name in most companies. It is shadow AI, and here it is measured by the government rather than guessed at by a vendor with a governance product to sell.
The operational consequence is concrete. If you run a business and you think your answer to the adoption question is no, there is a good chance your staff's answer is yes, and that gap is where your customer data is going. Before you run any experiment, that is the thing to ask about, and asking is free.
While workers were being asked whether they used AI, they were also asked how much time it saved them. 31% said one to two hours. 25% said less than an hour. 15% said three to four hours and another 15% said more than four. 10% said it saved no time at all, and 3% said it required additional time.
The productivity paradox in two official numbers
The Bureau of Labor Statistics does not publish an AI productivity figure. It says so plainly: it "implicitly captures AI use through its capital measure of software used in production." If you see a BLS AI productivity statistic quoted anywhere, it is not a thing BLS produces.
What BLS does publish is the money going in. Software investment grew at an 11.1% compound annual rate from 2019 to 2024, up from 7.9% over 2007 to 2019. Total factor productivity over that same 2019 to 2024 stretch rose 1.1%.
That is the paradox in two official series. Capital spending accelerating, measured output not yet moving.
A Richmond and Atlanta Fed study with Duke, covering 748 corporate executives, names it directly and documents "a productivity paradox, in which perceived productivity gains are larger than measured productivity gains, likely reflecting a delay in revenue realizations." CFOs reported a mean AI-attributed labor productivity increase of 1.8% for 2025. Measured against their own reported revenue and headcount changes instead of their perception, the implied figure was about 0.8% in high-skill services and finance and roughly 0.4% in low-skill services, manufacturing and construction.
Perception runs at roughly two to four and a half times measurement. That ratio is the single most useful thing in this post for designing your own test, because it tells you exactly which instrument not to use.
One more from that research: more than 40% of surveyed companies did not invest in AI at all in 2025, and the top stated reason was that the technology is not mature enough (42%), ahead of an untrained workforce (36%) and privacy concerns (36%). Meanwhile spending among those who do invest is wildly concentrated. More than half of firms expect to spend no more than $200 per employee on AI in 2026, while the top 10% plan at least $2,800 per employee. Any aggregate AI-spend number you read is describing that top decile.
The software engineering measurement, and why it cannot be re-run
The most cited productivity experiment in this field deserves careful handling, because the version most people quote is out of date and its own authors say so.
In 2025, METR ran a randomised controlled trial with 16 experienced open-source developers working 246 real issues on repositories averaging 22,000-plus stars and over a million lines of code. The result: "when developers use AI tools, they take 19% longer than without." The perception gap in the same study was larger than the effect. Developers forecast a 24% speedup, and after being measurably slowed, still believed AI had sped them up by 20%.
That finding is now labelled by METR itself: "These results are out of date. We have released results that are current as of early 2026... We believe these historical results no longer reflect the current impact of AI models on open-source developer productivity." Do not quote the 19% in the present tense. It is a 2025 artifact and its publisher retired it.
The follow-up flipped sign and then declined to claim a win. For the ten returning developers METR estimated an 18% speedup, with a confidence interval running from a 38% speedup to a 9% slowdown. For 47 newly recruited developers the estimate was a 4% speedup, with an interval from a 15% speedup to a 9% slowdown. Both cross zero. The pay rate also dropped from $150 an hour to $50, which METR flags as its own confound.
Then comes the part that actually matters, and I have not seen it made well anywhere else. METR is changing the experiment design because it cannot recruit a clean control group any more. "When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI."
The control arm has become unrecruitable. People will not do the work the old way for a study, which means the honest experiment is getting harder to run every quarter, not easier. If a well-funded research lab paying professional rates cannot hold a control group together, assume your own two-week baseline is fragile too, and design around it.
METR's separate 2026 survey of 349 technical workers is the coda: median self-reported change in the value of their work was 1.4x to 2x, while median self-reported speed change was 3x. The same organisation notes its earlier study found people overestimated AI's time effect by 40 percentage points.
Adoption up, trust down, measured twice
Two large independent surveys in the same year, both showing the lines diverging.
Stack Overflow's 2025 developer survey, with more than 49,000 respondents, found 84% saying they use or plan to use AI tools, up from 76% in 2024, while 46% said they do not trust the accuracy of AI output, up sharply from 31% the year before. 45% named debugging AI-generated code as time-consuming. 77% said vibe coding is not part of their professional work.
Google's DORA 2025 report, with roughly 5,000 respondents, found 90% adoption and more than 80% reporting productivity gains, alongside 30% reporting little or no trust in AI-generated code. DORA calls it the trust paradox. In the same data, developers spend a median of two hours a day working with AI and 59% report a positive influence on code quality.
They disagree on level. Stack Overflow's 46% distrust and DORA's 30% little-or-no-trust are not the same measurement, because one asks about accuracy directly and the other uses a five-point trust scale. Treat them as corroborating a direction, not a level. The direction is unambiguous: usage and confidence are moving apart.
The only failure evidence in this post that is not self-reported
Everything above is somebody telling a researcher how it went. There is one running dataset where a third party with subpoena power checked the work.
Damien Charlotin maintains a database of court decisions worldwide involving AI-hallucinated legal content. As of 21 August 2026 it stood at 1,936 cases, 1,327 of them in the United States, with 211 in Canada and 98 in Australia. NPR reported the same database at 206 cases in July 2025.
Roughly nine-fold growth in thirteen months, in documented instances of trained professionals filing fabricated material in court. And the count is a floor, not a total: the database describes itself as seeking to be exhaustive while remaining a work in progress that "will expand as new examples emerge."
For scale on the individual consequence, two attorneys in the Mike Lindell matter were each fined $3,000 over a filing containing more than two dozen errors including AI-fabricated case citations.
Note what this measures. Not "AI is unreliable" in the abstract. It measures the failure of a verification step in a profession where verification is the entire job. Whatever your equivalent of a citation is, a quoted price, a part number, a compliance date, a customer's balance, that is where your version of this lands.
Vendor numbers, with the qualifiers the vendors printed
Case-study figures are usable if you carry the fine print with them. Here is the fine print.
BlackLine reports that early adopters of its multi-agent reconciliation system achieved "up to a 92% reduction in manual reconciliation preparation time." Two qualifiers in one sentence: "up to" is a ceiling, not a typical result, and "early adopters" is a self-selected group of the customers who leaned in hardest.
Coupa announced that customers "used Coupa to manage over $425B of business spend, and realized almost $15B in savings" in a single quarter, in a release that also promotes its Navi AI agents. The $15 billion is a platform-wide spend-management figure. It is not savings attributed to the AI. The release places the two next to each other without claiming causation, and neither should you.
IBM says its AskHR agent achieved a 94% containment rate on common questions and contributed to a 40% reduction in HR operational costs over four years, with 75% fewer support tickets since 2016. IBM separately says AI and automation have it "on track to reach USD 4.5B in savings by the end of 2025," measured from January 2023. If you have seen $3.5 billion attributed to IBM, that is not the number on IBM's own page. And the AskHR page is loose about units: it cites over 2.1 million employee conversations annually and, elsewhere, more than 11.5 million employee interactions in 2024 alone, without defining either term.
Intercom publishes a customer stat for prediction market Kalshi of more than 80,000 monthly Fin resolutions at an 80% automation rate, as its customers page read on 21 August 2026. A resolution rate is not a satisfaction rate and it is not a cost saving.
Klarna is the most instructive because both halves are on the record. In February 2024 Klarna announced its assistant had handled 2.3 million conversations in a month, two-thirds of its customer service chats, doing "the equivalent work of 700 full-time agents," with resolution times down from 11 minutes to under 2 and an estimated $40 million profit improvement for 2024. In May 2025 its CEO conceded that AI customer service, while cheaper, produced "lower quality" output and that "investing in the quality of human support is the way of the future for us." Within the same week, CNBC reported the headcount side from Klarna's IPO prospectus: 5,527 full-time employees in December 2022 down to 3,422 in December 2024. Those two stories are five days apart. This was not a reversal so much as a split decision: keep the headcount reduction, restore the human escalation path. The prospectus figure is the strongest number in the file because it sits in a securities filing rather than a press release.
Salesforce against itself is the cleanest exhibit of all. Marc Benioff said he reduced customer support "from 9,000 heads to about 5,000 because I need less heads," with a company spokesperson clarifying the mechanism was declining case volume and not backfilling rather than layoffs, and hundreds of employees redeployed. Salesforce also sells the agent platform: Agentforce ARR passed half a billion dollars in Q3 FY26, up 330% year over year, across more than 9,500 paid deals. In May 2025,, Salesforce's own AI researchers published a benchmark of leading agents on realistic CRM work and found "only around 58% single-turn success... dropping significantly to approximately 35% in multi-turn settings," with agents showing "near-zero inherent confidentiality awareness." Same company, same year, both primary sources, opposite messages.
The pattern to internalise: in almost every case above, the party publishing the statistic sells the remedy the statistic implies. That does not make the numbers false. It makes them a category of evidence with a known lean, which is why the Census and Federal Reserve material carries most of the weight in this post.
Two independent results are worth keeping as reference points, because they are field measurements rather than marketing. The largest rigorous field study of AI in customer support, covering 5,179 agents, found a 14% average productivity gain, rising to 34% for novice and low-skilled workers. And an academic benchmark that simulated an entire company found "the most competitive agent can complete 30% of tasks autonomously." Stanford's AI Index summarises where measured gains cluster: 14% to 15% in customer support, 26% in software development, 50% in marketing output, with smaller gains in tasks requiring deeper reasoning.
One absence is worth reporting as a finding. Across all of this material I could not find a single named business with fewer than 100 employees with a measured AI result from a credible primary source. Every case study involves a company large enough to have a communications team. The evidence base for small business is not thin. It is missing.
What a real thirty-day experiment costs
Here is the whole budget for a two-person shop, at prices published on the vendors' own pages and read on 21 August 2026.
| Item | Published price | What it buys |
|---|---|---|
| Google Workspace Business Standard, 2 seats | $14 per user per month, annual commitment | Email and docs, with Gemini bundled in rather than sold as an add-on |
| Claude Pro | $17 per month on annual billing ($200 up front), $20 monthly | Includes Claude Code, Cowork, Design and Science |
| ChatGPT Plus | $20 per month | Includes ChatGPT Work and expanded Codex usage |
| GitHub Copilot Pro | $10 per user per month | Includes $15 of monthly AI credits and access to third-party agents |
| Netlify Free | $0 | 300 credits per month, functions and AI models |
| Cloudflare Workers Paid | $5 per month account minimum | 10 million requests and 30 million CPU-milliseconds included, no egress charge |
| Total | $80 per month | Two office suites, two frontier assistants, a coding agent, hosting |
The heavier build, with ChatGPT Business at $20 per seat annual, Claude Team standard seats at $20 per seat annual, Workspace Standard at $14 per seat for two people, plus GitHub Copilot Business at $19 per user and the same infrastructure, lands near $170 a month.
Compare that to the two comparables I could actually open. WebFX, one of the few agencies that publishes a real rate card, starts SEO services at $3,000 a month and says most businesses spend $2,500 a month, with small businesses typically paying $1,500 to $3,500. VC3 publishes managed IT at $150 to $400 per user per month. I looked hard for a primary survey of typical small-business agency retainers and could not open one. Every "industry average retainer" figure I found traces back to an agency's own blog citing an unnamed report, so I am not printing one.
If you would rather pay per token than per seat, these are the published API rates on the same date. Model names churn faster than prices, so this table is stamped 21 August 2026.
| Model | Input per million tokens | Output per million tokens |
|---|---|---|
| gpt-5.6-luna | $0.20 | $1.20 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
| Gemini 3.7 Flash | $0.75 (promotional) | $3.75 (promotional) |
| Haiku 4.5 | $1 | $5 |
| Sonnet 5 | $2 | $10 |
| gpt-5.6-terra | $2 | $12 |
| gpt-5.6-sol | $4 | $20 (short context) |
| Opus 5 | $5 | $25 |
| Fable 5 | $10 | $50 |
Four things about that table. Gemini 3.7 Flash's rate is explicitly promotional and the pricing page says it doubles to $1.50 and $7.50 on 1 January 2027, so any cost model built on it has a printed expiry; the OpenAI page similarly flags GPT-5.6 Sol's promotional pricing as running at least through 21 November 2026. All three vendors converge on exactly 50% off for batch processing. Caching is where they diverge, and Anthropic prices cache reads at 10% of each model's input rate, which puts Haiku 4.5 cache reads at $0.10 per million tokens. And there is a free floor under all of it: ChatGPT Free is $0, GitHub Copilot Free is $0 with 2,000 completions a month, Cloudflare Workers AI gives 10,000 Neurons a day free and charges $0.011 per 1,000 thereafter, and Cloudflare's AI Gateway core features (dashboard analytics, caching, rate limiting) are free on every plan.
One trap in the seat pricing, because it is pure money. On Microsoft's own site, Business Standard is $14 per user per month paid yearly and the Copilot Business add-on lists at $21 and is currently promoted at $18, so the list-price pair sums to $35 and the promoted pair to $32. The bundled SKU, Microsoft 365 Business Standard with Copilot, is $23.50. That is $11.50 per user per month cheaper at list, and still $8.50 cheaper at the promoted add-on price, for the same two things. Same pattern at the top: Business Premium at $22 plus Copilot at $21 is $43, against $32 for the bundle. If you already have a subscription and you add Copilot to it, you are on the expensive path. Switch SKU instead. And on the Google side, Business Plus at $22 lists no AI features beyond Standard at $14; its AI bullet reads "Everything included from Starter and Standard." Plus buys storage, eDiscovery and security, not more intelligence.
A measurement design you can actually run
The design has to survive the two problems this whole post has been documenting: people overestimate their own speedup by a wide margin, and control groups do not hold. So the design is deliberately dumb, and it counts things rather than asking about them.
1. Pick one task with a countable unit. Not "marketing." One task, done repeatedly, where you can write down a number at the end of each instance. Quotes prepared. Accounts reconciled. Support emails answered. Listings written. Job descriptions drafted. Inspection reports typed up. The measured gains in the literature cluster in exactly this kind of work: structured, repetitive and already monitored. Vague creative work is where measurement dies.
2. Write down the unit of output and the quality bar before you start. The unit is the thing you count: one quote, one reconciled account, one answered ticket. The quality bar is the thing that disqualifies a unit from counting: a quote with a wrong price, a reply that needed a second reply, a draft the customer pushed back on. Klarna's whole lesson is that cheaper and lower quality is a real, common outcome, and if you have not defined quality in advance you will book it as a win.
3. Baseline for two full weeks before anybody touches a tool. Count three things per unit: how many you completed, roughly how long each took, and how many needed rework. Two weeks, not two days, because otherwise a single busy Tuesday becomes your baseline.
4. Ask, before you start, who is already using AI for this. 55% of workers say they use it on the job while only 21.8% of firms say they do. If half your baseline period is quietly AI-assisted already, your experiment measures nothing. Ask without consequences attached, because a punitive version of that question gets a false answer.
5. Then run thirty days with the tool, counting the exact same three things. Same people, same task, same definitions. Do not change the workflow and the tool in the same month, or you will not know which one moved the number.
6. Count the rework, not just the drafts. 45% of developers named debugging AI-generated output as time consuming. The unit is not done when the draft appears. It is done when it is correct enough to send. Whatever your version of a fabricated citation is, that is the number that eats the gain.
7. Do not ask anyone whether it helped. This is the rule the research supports hardest. Developers in a controlled trial predicted a 24% speedup, were measured 19% slower, and still reported a 20% speedup afterwards. CFOs perceived 1.8% productivity gains against the 0.4% to 0.8% implied by their own revenue and headcount. A survey of technical workers self-reported 3x speed but only 1.4x to 2x value. Self-report runs two to four and a half times measurement in the two datasets here that produce comparable numbers, and in the METR trial it did not just exaggerate the effect, it got the sign wrong. The count on your sheet is the finding. The feeling in the room is not.
8. Write the kill criterion on the same sheet, before day one. Decide now what result makes you cancel: a specific minimum improvement in units per hour, or a maximum acceptable rework rate, in writing, dated. Deciding afterwards is how a $20 seat becomes a permanent line item nobody can defend.
Two weeks of baseline and thirty days of trial cost you one or two seats from the table above and a spreadsheet. The measurement is cheaper than the first hour of any consultant who would run it for you.
What to do when the answer is no
It will often be no, and no is a completely ordinary result. 61.6% of businesses with no plans to adopt say AI is simply not applicable to what they do, and only 6.9% say cost is the reason. Ninety-one percent of US firms report no measured labor productivity impact over three years. You are not the exception if it does not work.
When the sheet says no:
Change the task before you change the tool. The gains that show up in real measurements are concentrated in structured, high-volume, monitored work, and the largest single effect in the field literature went to novices doing customer support, at 34% against a 14% average. If your test task was your most judgment-heavy work, you tested the hardest case. Try the most repetitive one instead.
Cancel the seat the same week. This is the whole reason to run the experiment at $80 a month rather than at a $3,000 retainer: cancelling costs nothing but a calendar reminder. Auto-recharge, where a vendor offers it, should stay off. Netlify's, for one, is disabled by default.
Consider the cheapest form of yes. 64.3% of AI-using firms made no changes at all to adopt: no training, no purchase, no new workflow, no vendor. If a free tier and no process change gets you a small, real, countable improvement on one task, that is a complete and respectable outcome. It is what the median adopter is actually doing.
Do not let a vendor tell you what your number should have been. Every headline efficiency figure in the first half of this post came from an organisation with a product, a consultancy practice or a reputational stake in the answer. The two bodies of evidence that carry real weight, the Census business surveys and the Federal Reserve firm surveys, are both boring, both free, and both say the effect is smaller and later than the marketing does.
And keep the shadow-AI answer even if you keep nothing else. If the experiment taught you that three of your five staff were already pasting customer information into a free chatbot, the thirty days paid for itself regardless of what the units-per-hour column says.
If you are reading this because you are already paying an agency retainer for the marketing and search side of this work, the whole map of doing it yourself is in The $20 Dollar Agency, which is $9.99 on Kindle. The short version is on this page: measure one task, count the units, and decide with the sheet in front of you.
Fact-check notes and sources
- The 21.8% adoption figure, the size gradient and the expectations gap come from the Census Bureau's Business Trends and Outlook Survey, cycle 202616, reference period 13 to 26 July 2026, published 13 August 2026, in the National and Employment Size Class workbooks. Standard error on the national figure is 0.33%. The unit response rate is 12%, about 24,500 responses from a 1.2 million business sample, non-response adjusted, per the survey's own response-rate workbook. Census publishes new cycles fortnightly and the file addresses are stable, so re-pull rather than quoting this post's numbers later.
- The November 2025 question change is documented in Census's own story, Large Firms With at Least 20 Employees Biggest AI Users (26 May 2026), and the three-cause explanation for the jump is in the Census working paper The Microstructure of AI Diffusion, CES-WP-26-25 (April 2026), which also carries the 18% firm-weighted versus 32% employment-weighted split and the 117,000-firm supplement sample. The apparent conflict between "less than 20% of firms with four or fewer employees" in the May story and 21.9% in the August file is a genuine difference between cycles, not an error in either.
- The AI supplement figures (64.3% made no changes, 95.7% no employment change, 61.6% "not applicable", 10.1% replacing a human task) are from the BTOS AI Supplement Table 2026, pooled over six biweekly panels from 17 November 2025 to 8 February 2026.
- The survey spread is laid out in Monitoring AI Adoption in the U.S. Economy, FEDS Notes, Board of Governors, 3 April 2026, which is also the source for the Survey of Business Uncertainty's 78% employment-weighted and 69.4% equal-weighted figures, and for the 5% to 40% range across 16 earlier surveys reviewed by Crane, Green and Soto.
- The Small Business Credit Survey's 46% is from the 2026 Report on Employer Firms, 6,525 responses, fielded 3 September to 14 November 2025. The same report states in its methodology that it "is not a random sample" and warns about convenience-sample bias. Its self-reported outcomes (71% increased productivity, 39% improved quality, 31% higher sales) are firms describing themselves rather than anything measured, and only 7% of its AI users said they had fully integrated AI into business processes.
- The "no impact" finding is Firm Data on AI, Atlanta Fed Working Paper 2026-3 (March 2026), almost 6,000 executives across four countries from stratified firm samples. The perception-versus-measurement gap is from Artificial Intelligence, Productivity, and the Workforce, Working Paper 2026-4, 748 executives, and the spending concentration from the Atlanta Fed macroblog on AI spending (6 May 2026). That macroblog notes its per-employee average is winsorized at the 99th percentile, which is the authors' own point about how few firms drive the aggregate.
- The MIT NANDA 95% figure comes from The GenAI Divide: State of AI in Business 2025 (July 2025). MIT no longer serves the report at its original address; when I checked on 21 August 2026, nanda.media.mit.edu/ai_report_2025.pdf returned the Media Lab group's overview page rather than the PDF, so the wording, the 153-response methodology and the report's own limitations note were read from a third-party PDF mirror. Treat the figure accordingly. It is quoted here as the number everybody cites, not as a measurement I would rely on.
- The Wharton counterpoint is Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise, Wharton Human-AI Research with GBK Collective, October 2025, 801 respondents, fielded 26 June to 11 July 2025. The screening criterion quoted in the body ("1000+ employees and >$50 million revenue") is from its own methodology section.
- McKinsey's 88%, its sample size and its weighting were verified through the Stanford HAI 2026 AI Index, which reproduces them with a full source citation and attaches its own caveat that the results are self-reported and directional. McKinsey's own State of AI page would not open for me across repeated attempts, so any profit-impact figures from that survey are deliberately absent here rather than repeated second hand.
- BLS material: the statement that BLS publishes no AI productivity figure is on its own page, Productivity and Artificial Intelligence. The 11.1% software investment growth rate and the 1.1% total factor productivity change are from AI and the rise of software investment, Monthly Labor Review, May 2026.
- The worker-side figures (55% of workers, and the time-saved distribution) are from the Census Bureau's March 2026 Household Trends and Outlook Pulse Survey, written up at About a Third of Workers Who Used AI in the Last Week Said They Completed Tasks One to Two Hours Faster, 11 August 2026.
- METR: the 19% slowdown, the sample of 16 developers and 246 issues, and the out-of-date banner are all on the same page, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. The reversed estimates, the confidence intervals, the pay-rate change and the 30% to 50% task-withholding finding are in We are Changing our Developer Productivity Experiment Design (24 February 2026). The self-report survey of 349 technical workers is Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity.
- Trust data: Stack Overflow's 2025 Developer Survey release (29 July 2025, more than 49,000 respondents) and Google's DORA 2025 report summary (roughly 5,000 respondents). The two use different question wordings and are not comparable on level, only on direction.
- The hallucination database is Damien Charlotin's AI Hallucination Cases, read on 21 August 2026 at 1,936 cases. The July 2025 count of 206 comes from NPR's report (10 July 2025), which also quotes the sanctions order in the Lindell matter directly. The underlying court order is public. I relied on NPR's direct quotation of it rather than reading the order myself, which is a limitation worth knowing.
- Vendor claims are each linked to the vendor's own page and should be read as marketing rather than measurement: BlackLine (27 July 2026, "up to" a 92% reduction, for self-selected early adopters), Coupa via PRNewswire (15 December 2025, a platform-wide spend figure and not an AI-attributed one), IBM AskHR and IBM on enterprise productivity, Intercom's customers page as it read on 21 August 2026, Klarna's February 2024 release, and Salesforce's Q3 FY26 results. The Klarna reversal is reported by Entrepreneur (9 May 2025) and the prospectus headcount figures by CNBC (14 May 2025). The Salesforce support headcount and the attrition mechanism are reported by The Register (2 September 2025). A frequently quoted Agentforce autonomous-resolution percentage does not appear in that earnings release and I could not locate it on any Salesforce-owned page, so it is not in this post.
- Independent benchmarks: Salesforce's own CRMArena-Pro (arXiv:2505.18878, May 2025), TheAgentCompany (arXiv:2412.14161), and the customer-support field study Generative AI at Work, NBER Working Paper 31161, 5,179 agents, later published in the Quarterly Journal of Economics. The 14% to 15%, 26% and 50% task-level gains are as summarised by the Stanford AI Index economy chapter.
- All prices were read on the vendors' own pricing pages on 21 August 2026 and change often: ChatGPT, OpenAI API, Claude, Gemini API, Google Workspace, Microsoft 365 plans with Copilot and Microsoft 365 business plans with Teams, GitHub Copilot, Cloudflare Workers, Cloudflare Workers AI, Cloudflare AI Gateway, Netlify and Vercel. Two top tiers have no printed price at all: both Claude Max and ChatGPT Pro read "From $100" with the higher usage multiple unpriced on the page, so I have not guessed at it.
- Agency and IT comparables are WebFX's published SEO pricing and VC3's managed IT pricing guide, both read 21 August 2026. I deliberately did not print an "average small-business retainer" figure. Every such number I found traces back to an agency's own marketing content citing a report I could not open, which makes it unusable.
Related reading
- The $50-a-month AI stack for a small business: the concrete tool list that sits under the budget in this post.
- Running AI agents without a surprise bill: the spend controls to put in place before day one of any trial.
- Build the guardrail before you mandate AI on your team: what to do about the gap between what your firm says and what your staff already do.
- A study like this quotes at $15,000 and ran in an afternoon: a worked example of measuring the cost difference instead of arguing about it.
- I cut a recurring AI bill by more than half in an afternoon: what to do after the experiment says yes and the bill starts growing.
This post is informational and is not legal, financial or investment advice. Survey figures are as published on the dates cited, several of the underlying series are revised or re-fielded regularly, and prices change frequently. No affiliation with any vendor, publisher or research organisation mentioned is implied, and mentions are nominative fair use.