- AI pays back most reliably on repetitive text work — drafting, summarising, rewriting, note-taking. That is where controlled trials consistently find real gains.
- It pays back least on judgement, relationships and anything you are accountable for, because checking the output costs more than the drafting saved.
- Failure is normal rather than shameful: S&P Global found 42% of companies scrapped most of their AI initiatives in 2025, up from 17% the year before.
- Across a whole workforce, self-reported time savings run at about 3% of working hours — real, but nothing like the numbers in the marketing.
- The only test that settles it is arithmetic: hours returned each week, times what your hour is worth, minus the subscription. Run it on one task for two weeks.
Sometimes. AI is worth it for a narrow, well-defined slice of small business work — mostly repetitive writing and reading — and it is not yet worth it for most of the rest. The gap between those two categories is much wider than the marketing suggests, and knowing which side a task falls on is most of the skill.
That answer will annoy both camps. It is also what the evidence supports, so here it is with the numbers attached.
What does the research actually say about AI paying off?
Two things, consistently: adoption is high and measured returns are low. Those are not contradictory. It is what an early technology looks like when everyone tries it and only some uses stick.
S&P Global Market Intelligence surveyed over 1,000 organisations in North America and Europe and found the share scrapping most of their AI initiatives had risen to 42% in 2025, up from 17% a year earlier. The average organisation abandoned 46% of its AI proofs of concept before they ever reached production. S&P Global, via CIO Dive
You have probably also met the claim that 95% of enterprise AI pilots deliver no measurable impact on profit and loss. It comes from MIT's NANDA initiative, based on 150 leader interviews, 350 employee surveys and 300 public deployments. Fortune Treat it with care: the headline number rests on a subset of 52 interviews the report itself calls "directionally accurate based on individual interviews rather than official company reporting", and success was defined narrowly as measurable P&L impact six months after the pilot. Marketing AI Institute The direction is probably right. The decimal point is not.
The most careful measurement of what AI does to real working lives comes from Denmark, where economists Anders Humlum and Emilie Vestergaard matched survey answers to payroll records covering 25,000 workers across 7,000 workplaces. Users reported average time savings of about 3% of work hours. Fortune On pay and recorded hours the researchers found precise null effects, ruling out changes larger than 2% two years after ChatGPT launched. Humlum and Vestergaard, NBER
Where does AI demonstrably pay for a small business?
On repetitive text: drafting, summarising, rewriting, transcribing and first-pass research. These are the tasks where controlled experiments find large, repeatable gains — and they happen to make up a great deal of a small business owner's week.
In a randomised experiment with 453 college-educated professionals — marketers, grant writers, consultants, HR staff, managers — writing the sort of documents that fill a real working day, access to ChatGPT cut average completion time by 40% and raised assessed quality by 18%. MIT News on Noy and Zhang, Science The researchers were careful about the limit: those tasks "didn't require precise factual accuracy or context about things like a company's goals or a customer's preferences". That is exactly the boundary you should be drawing too.
It matches what small business workers actually do with it. In the U.S. Chamber of Commerce Foundation's June 2026 survey of 1,070 people employed at US small businesses, about half already use AI at work — and among those users, 90% use it for writing and editing, the most common application by some distance. U.S. Chamber of Commerce Foundation
| Task | Does it pay? | Why |
|---|---|---|
| Drafting and rewriting text | Yes, reliably | The output is a first draft you were going to edit anyway. Mistakes are visible and cheap to fix. |
| Meeting notes and summaries | Yes, reliably | You were in the room. You can spot what it got wrong in seconds. |
| Sorting, tagging, triaging | Usually | Volume work where a small error rate is survivable and the honest alternative is not doing it at all. |
| First-pass research | With supervision | Fast to produce, slow to verify. Worth it only if you know enough to smell a wrong answer. |
| Numbers you will act on | Rarely | Checking costs about as much as doing it yourself, and the failure mode is silent. |
| Client judgement and relationships | No | The thing being paid for is that a person decided and stands behind it. |
The whole pattern reduces to one question: how expensive is it to spot an error? Where errors are obvious and cheap, AI pays. Where they are silent and expensive, it does not, and no amount of prompt technique changes that.
Where does AI not pay off?
Anywhere you are accountable for the output, anywhere the value is that a human exercised judgement, and — more often than people expect — anywhere you are already very good and very fast.
That last one is the genuine surprise. METR ran a randomised controlled trial with 16 experienced open-source developers on 246 real tasks in repositories they had worked in for years. When AI tools were allowed, they took 19% longer. Beforehand they forecast a 24% speed-up; afterwards they still believed they had been about 20% faster. METR METR says plainly this does not generalise to all developers. The transferable bit is the perception gap: people are unreliable witnesses to their own productivity.
The common mistake is not picking the wrong tool. It is being certain a tool is saving you time without ever having checked.
The accountability problem is more concrete. In Moffatt v Air Canada, a customer asked the airline's website chatbot about bereavement fares and was given the wrong answer. Air Canada argued the chatbot was "a separate entity" responsible for its own statements. The British Columbia tribunal rejected that, found the airline had not taken reasonable care to ensure the chatbot was accurate, and held it liable. McCarthy Tétrault on Moffatt v Air Canada Whatever the tool says to your customer, it says on your behalf.
The verification cost is not theoretical either. A public database of court decisions in which someone filed AI-fabricated case citations had passed 200 entries by July 2025, about two months after it was set up. Forbes These are professionals in the one occupation most obviously trained to check citations. A plausible-looking paragraph about your own numbers deserves the same suspicion.
Why do so many AI projects get abandoned?
Rarely because the tool did not work. Usually because nobody owned it, nobody measured it, and the subscription outlived the enthusiasm. In a small business the pattern is consistent:
- No named owner. A tool that is everyone's job is nobody's. In a business of one to ten people that owner is you, and you have the least spare attention.
- No before-measurement. If you never wrote down how long the task took beforehand, you cannot tell whether it is faster now, and the question collapses into opinion.
- Too many at once. Five new tools in a month is how you end up with five unused logins and a worse month.
- Starting with the hardest problem. The task people most want to hand over is usually the judgement-heavy one AI handles worst — so the first experiment fails and the whole idea gets written off.
None of those are technology problems. They are the reasons any small operational change fails — which is oddly reassuring, because they are fixable without understanding a thing about how the models work.
How do you work out whether it is worth it for your business?
Run the arithmetic on one task, for two weeks, before you buy anything else. Hours returned per week, times what an hour of your time is worth, minus what the tool costs to run. If that number is not clearly positive, the tool is a hobby rather than an investment.
- Pick the most repetitive text-based task in your week. Not the most annoying one — the most repeated one.
- Time it for a week before you change anything. Real minutes, written down as they happen, not estimated on a Friday.
- Use one tool on it for the next fortnight. One. Change three things and you will not know which one worked.
- Time it again, and count the checking. Editing and verification are part of the task, not overhead you get to leave out.
- Decide. If the answer is no, cancel that day — while you still remember why.
[Steve — add a real before/after here from a client or your own week: the task, minutes before, minutes after, which tool.]
Most people who run this properly find one or two tasks where the answer is an obvious yes and several where it is a shrug. That is a good outcome, not a disappointing one. Two tasks genuinely handed over is a couple of hours back every week — a better return than most software you already pay for.
The free 3-minute scorecard covers the three places small business weeks usually leak time: your inbox, your repetitive tasks, and the tools you are already paying for.
Take the free scorecardIf you would rather not run the experiment yourself, that is what an AI tools assessment is for: the same arithmetic done against your actual week, with the shortlist already filtered. The running costs of the tools themselves are broken down in how much AI tools cost for a small business, and there is more about how I work.
Frequently asked questions
Is AI worth it for a one-person business?
Often yes, and for a narrower reason than people assume: a sole trader has no one to delegate the repetitive writing to, so the comparison is not "AI versus a staff member" but "AI versus doing it at 9pm". Start with the drafting, summarising and note-taking. Leave the client judgement alone.
Do most AI projects really fail?
A lot of enterprise ones stall. S&P Global found 42% of organisations abandoned most of their AI initiatives in 2025, and that the average organisation scrapped 46% of its proofs of concept before production. S&P Global, via CIO Dive Small businesses fail differently and much more cheaply — a cancelled subscription rather than a written-off integration programme.
What is a realistic time saving from AI?
Lower than the headlines. Controlled experiments on well-chosen writing tasks show around 40% faster completion, but across a whole workforce — including all the tasks AI is bad at — self-reported savings come out at roughly 3% of working hours. Fortune, on the Danish study Both numbers are true. The difference between them is task selection, which is the part you control.
Can AI make me slower?
Yes, particularly on work you already do quickly and well. In METR's randomised trial, experienced developers took 19% longer on real tasks when allowed to use AI tools, while believing they had gone 20% faster. METR The lesson is not to avoid AI; it is that your feelings about your own speed are not evidence.
Should I wait until the technology settles down?
For text-based tasks there is not much to wait for — the drafting and summarising case has been stable for a couple of years and the tools are cheap enough to test in an afternoon. For anything autonomous that acts on your behalf without a person checking it, waiting is entirely reasonable. You remain liable for what it does either way.
How much should a small business spend on AI tools?
Less than most people expect, and the number should be set by the arithmetic rather than by a budget. Individual tools commonly sit in the $10–30 per month range; the detail is in how much AI tools cost for a small business.
- AI project failure rates are on the rise (S&P Global Market Intelligence survey) — CIO Dive (14 March 2025)
- MIT report: 95% of generative AI pilots at companies are failing — Fortune (18 August 2025)
- That viral MIT study claiming 95% of AI pilots fail? Don't believe the hype — Marketing AI Institute (26 August 2025)
- Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI (NBER Working Paper 33777) — Anders Humlum and Emilie Vestergaard, NBER (May 2025, revised March 2026)
- Study looking at AI chatbots in 7,000 workplaces finds no significant impact on earnings or hours — Fortune (18 May 2025)
- Study finds ChatGPT boosts worker productivity for some writing tasks (Noy and Zhang, Science) — MIT News (14 July 2023)
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR (10 July 2025)
- Half of Small Business Workers Use AI — Most to Boost Productivity, Not Automate Jobs — U.S. Chamber of Commerce Foundation / Ipsos (17 June 2026)
- Moffatt v. Air Canada: A Misrepresentation by an AI Chatbot (2024 BCCRT 149) — McCarthy Tétrault (19 February 2024)
- Attorneys — Track AI Hallucination Case Citations With This New Tool — Forbes (18 July 2025)
