- AI is reliably good at reshaping material you give it — drafting, summarising, extracting, transcribing, translating and sorting. That is the whole dependable list.
- It is unreliable whenever the answer has to come from outside the conversation: facts it was not given, live information without a search step, arithmetic, and your own numbers.
- It is the wrong tool entirely for anything you have to be accountable for. "The AI said so" is not a defence to a client, a regulator or the tax office.
- Even on the easiest task — summarise this document and add nothing — the best models still contradict the source in roughly 2% of summaries.
- Starting now is not late. Fewer than 20% of firms with four or fewer employees report using AI at all.
Today's AI is genuinely reliable at one family of jobs: taking text or speech you already have and changing its shape. Drafting, summarising, extracting, transcribing, translating, sorting. That is a narrower claim than the marketing makes, and for most small businesses it is still worth a few hours a week.
Almost everything outside that family — facts it was not handed, arithmetic, current information, anything with your name on the outcome — it does badly enough that you have to check it, and checking sometimes costs more than the drafting saved. Here is the map, with the failure modes marked.
What is AI genuinely reliable at?
Tasks where you supply the material and the AI reshapes it. The model is not being asked to know anything; it is being asked to condense, re-file or rewrite what you just gave it. That is where error rates are lowest and where checking the output takes seconds rather than an afternoon.
| What it does well | What that looks like in a small business | Checking required |
|---|---|---|
| Drafting | First pass at a quote email, a job ad, a listing, an FAQ page, a follow-up you have written 200 times | Read it before it goes. You would read your own draft too. |
| Summarising | A 40-minute call recording turned into six bullets and three actions | Spot-check names, dates and figures against the source. |
| Extracting | Pulling PO number, date and total out of 30 supplier PDFs into one spreadsheet | Check the totals column. It is the field that costs you if it is wrong. |
| Transcribing | Site visits, client interviews, voice notes from the car | Skim it. Watch for invented lines where the audio went quiet. |
| Translating | Inbound customer emails, product copy for another market | Fine for gist. Use a human for anything contractual. |
| Classifying | Sorting enquiries into quote / support / supplier / bin | Grade 20 by hand first and count the mistakes before you trust it. |
This is not a theory about what AI could do — it is what firms already do with it. A US Census Bureau working paper published in April 2026 found that writing, document analysis and information search are the leading generative-AI tasks inside businesses, and that 65% of AI-using firms confine it to three tasks or fewer. Narrow use is the normal case, not a sign you are behind.
What is AI still unreliable at?
Anything where the answer has to come from outside the conversation, and anything where a person has to be accountable for the result. Those are two different failures and it is worth keeping them apart.
- Facts you did not supply. Ask for a statistic, a case reference, a supplier's phone number or a legal deadline and you may get a confident invention. It will be formatted correctly, which is exactly what makes it dangerous.
- Current information without a search step. If the tool is not visibly retrieving live pages, it is answering from training data of unknown vintage. Prices, opening hours, tax thresholds and rules change; the model does not know that they have.
- Arithmetic. Language models predict text, not sums. Some can now write and run code to calculate properly, which is far more trustworthy — but a chatbot doing mental maths in prose is guessing at the shape of a number. Keep your accounts in accounting software.
- Your own numbers. It does not know your margins, your stock, your capacity or who owes you money unless you paste it in. Any answer about your business that you did not feed it is fiction.
- Anything you have to answer for. Tax positions, employment decisions, safety, medical or legal advice, contract terms. Not because the output is always wrong, but because there is no one to hold responsible when it is.
Notice the pattern. Every item on that list asks the model to supply the truth rather than reshape yours. The reliable list does the opposite. That single distinction will sort most new tasks for you without any technical knowledge at all.
Why does AI make things up?
Because it is built to always produce an answer, and "I don't know" scores badly. Researchers at OpenAI argue that standard training and evaluation procedures reward guessing over acknowledging uncertainty — a model that abstains loses points on the benchmarks it is optimised against, so it learns to guess fluently instead. The industry term is hallucination. In practice it means confident, well-written, wrong.
It is worth knowing how good the floor is. On the easiest possible fact-keeping task — summarise this document and add nothing — the leading models on Vectara's public hallucination leaderboard still contradicted the source in roughly 1.8% to 3.3% of summaries when it was last updated in May 2026. That is the best case, on the task AI is best at, with the source sitting right there in front of it.
Transcription has the same weakness in a quieter form. A study presented at the 2024 ACM FAccT conference found that roughly 1% of audio clips run through OpenAI's Whisper came back with entire fabricated phrases or sentences, typically where the speaker paused, and 38% of those inventions carried some explicit harm. That tested a 2023-era model and transcription has improved since — but the failure mode has not gone away. Silence is where made-up text appears.
Fluent and correct look identical on the page. That is the entire problem, and no amount of clever prompting fixes it.
What should a small business try first?
Pick the one task you repeat every week that is mostly typing, has a predictable shape, and where a mistake costs nothing. For most people that is email replies, meeting notes, or the first draft of a quote or proposal.
- Track five days. Note what you did in 15-minute blocks. Not forever — one working week is enough to find the repeat offenders. There is a longer version of this exercise in how many hours a week do you lose to admin.
- Choose the cheapest failure. Of the repeat tasks, pick the one where a bad output is embarrassing at worst, never expensive.
- Run it in parallel for two weeks. Do the task your normal way and with AI, side by side. Yes, that is slower at first. It is the only way you find out.
- Time it with a clock. Minutes before, minutes after, including the checking. Checking time counts.
- Only then add a second task. Adopting five tools at once is the most common way people end up using none.
How much time does AI actually save a small business?
For most people, somewhere between nothing and a few hours a week, and it depends almost entirely on how much of your work is text. In a US survey analysed by ITIF, 20.5% of weekly generative-AI users said they had saved four hours or more in the previous week, rising to 33.5% among daily users.
Read that the honest way round: nearly four in five weekly users saved less than four hours. The gains are real and they are not transformative, which is roughly what you would expect from a tool that is excellent at one family of tasks and useless outside it. [Steve — add what a typical assessment client actually recovers in the first fortnight.]
If you want the fuller argument about whether the trade is worth making at all, including where it clearly is not, that is is AI actually worth it for a small business.
The free 3-minute scorecard walks through where your week leaks — your inbox, your repeat tasks and the tools you already pay for — and tells you which one to hand over first.
Take the free scorecardDo you need to be technical to start?
No. Nothing on the reliable list requires code, and the tools that do it are ordinary web apps you sign up for. The skill that matters is judgement: knowing which tasks belong on which list, and building the habit of checking the things that count.
It is also worth knowing that you are not behind. Census Bureau figures put overall business AI use at between 17% and 20% from December 2025 to May 2026, and fewer than 20% of firms with four or fewer employees report using it at all. The gap is widest at the smallest end, which is precisely where a couple of hours a week matters most.
If you would rather have the sorting done for you than work through it task by task, that is what an AI tools assessment is. There is more on how I work in about.
Frequently asked questions
Can AI do my bookkeeping?
Partly, and only in the right shape. Accounting software with AI features can extract data from invoices and suggest categories for transactions reliably enough to be useful, because that is extraction and classification. A general chatbot doing sums in prose is a different thing — it predicts text rather than calculating, so treat any figure it produces as a guess until the software confirms it.
Is it safe to put client information into an AI tool?
It depends on the tool and the plan, and you have to check rather than assume. Look for whether your inputs are used to train the model, where the data is stored, and what your own client contracts and privacy obligations say about sharing information with a third party. Some business tiers explicitly exclude your data from training; some free tiers explicitly do not. If you cannot find a clear answer in the terms, treat that as an answer.
Which AI tool should I start with?
For most small businesses it is one general assistant plus one job-specific tool for the task that eats the most time — notes, email or documents. Which brand matters far less than the sequencing. Pick one task, prove it saves measurable time, then add the next.
Will AI replace my staff?
The evidence so far says mostly no, at least at this size. The Census Bureau's April 2026 working paper found 66% of AI-using firms apply it to augment tasks rather than replace workers, and that AI-related employment decreases occurred in only 2% of firms.
How do I tell when an AI answer is wrong?
You often cannot from the text alone, which is why the check is procedural rather than intuitive. Prefer tools that retrieve and cite live sources, ask for the source of any factual claim, and open the ones that matter. If an answer cannot be traced back to a document you can read, treat it as a hypothesis rather than a fact.
- The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks (CES-WP-26-25) — US Census Bureau, Center for Economic Studies (April 2026)
- Large Firms With at Least 20 Employees Biggest AI Users — US Census Bureau (May 2026)
- Hallucination Leaderboard (HHEM summarisation benchmark) — Vectara (Updated 11 May 2026)
- Why Language Models Hallucinate — Kalai, Nachum, Vempala and Zhang (OpenAI / Georgia Tech), arXiv:2509.04664 (4 September 2025)
- Careless Whisper: Speech-to-Text Hallucination Harms — Koenecke et al., ACM FAccT 2024, arXiv:2402.08021 (June 2024)
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR (10 July 2025)
- Fact of the Week: 20.5 Percent of Frequent Generative AI Users Report Saving Four or More Hours Weekly at Work — ITIF, citing Federal Reserve Bank of St. Louis survey data (9 May 2025)
