How to choose AI tools for your business

Start from the work, not the tool. Most wasted AI subscriptions can be traced back to a decision that started with a demo instead of a task.

Key takeaways
  • Choose the task first, tool second. Name one specific thing you repeat at least weekly before you open a single pricing page.
  • Quantify it before you shop: minutes per run × runs per month, written down. After you have seen the demo you will not be objective about that number.
  • The tool has to reach the software the work already lives in. If it needs you to work somewhere new, assume it will be abandoned.
  • Trial on a job you have already completed, so you know what the right answer looks like. Vendor demos are built to succeed.
  • Re-measure at two weeks against the number from step two. Feelings about time saved are unreliable — measured results and perceived results diverge badly.

Choose the task first, then the tool. The process that works is five steps: name something you repeat at least weekly, count what it actually costs you in minutes, check the tool reaches into software you already use, trial it on a real job you have already done, then measure the difference two weeks later. Most bad AI purchases skip steps one and five.

The usual sequence is the opposite: see a demo, sign up, use it twice, forget it, notice the charge four months later. That is what happens when a decision starts with a tool that has no job to do.

What is the five-step framework for choosing an AI tool?

Five steps, in order. Each is cheap, and each is designed to kill candidates before you spend money on them.

  1. Name the repeated task. Not "marketing" or "admin" — one thing you do over and over: follow-up emails after a call, voice notes into quotes, chasing unpaid invoices.
  2. Quantify it. Minutes per run × runs per month, written down before you look at any software.
  3. Check it reaches your stack. It has to touch the inbox, calendar or system where the work already happens. A tool that needs a new place to work needs a new habit.
  4. Trial it on real work you have already done. Not the vendor's sample data. Three jobs from last fortnight, where you know what a good answer looks like.
  5. Measure again at two weeks. Same arithmetic as step two. If the number has not moved, cancel while it is still one small line on the card.

Steps one and two take ten minutes with a pen. Step four takes half an hour. Step five takes five minutes, and is the one almost nobody does.

Why should you start with the task instead of the tool?

Because the size of the benefit is decided by the task, not the software. Point the same model at two different jobs and the returns are an order of magnitude apart.

Stanford's 2026 AI Index puts measured productivity gains at roughly 14–15% in customer support, 26% in software development and about 50% for marketing output, and notes that gains shrink on work needing deeper reasoning. Stanford HAI, 2026 AI Index

14–50%
The spread in measured productivity gains between task types. Which task you point a tool at matters more than which tool you point at it.

Getting that order wrong is also where failures start. MIT's Project NANDA report found only about 5% of custom enterprise AI tools reach production, and put the cause in workflow integration rather than model quality. Preliminary findings from a corporate sample, not peer-reviewed research, so hold the figure loosely — but nothing there failed because the model was not clever enough. Reported by Virtualization Review, August 2025

How do you work out how much time a task is really costing you?

Minutes per run times runs per month, timed once rather than guessed. Ten minutes a day is about 3.5 hours a month; at $80 an hour that is $280 of your time, which comfortably justifies a $20 subscription. Ten minutes a week is 45 minutes a month, and does not.

  • Minutes per run. Time it once with a stopwatch. Memory runs high on tasks you dislike and low on the ones you do on autopilot.
  • Runs per month. Count from your sent folder or calendar, not from your impression of a busy week.
  • What it costs when it is late. Some tasks are cheap in minutes and expensive in consequences — quotes that go out three days after the call.

Then set a realistic ceiling. In the Federal Reserve Bank of St. Louis survey of US workers, 20.5% of people who had used generative AI at least once in the past week reported saving four hours or more; among daily users that rose to 33.5%. ITIF, May 2025

Most people save less than four hours a week, and savings scale with frequency — another argument for a task you do daily over an impressive one you do quarterly. Business.com's 2026 survey of 1,009 people at US businesses with 2–250 employees found the same by role: managers saved 7.2 hours a week against 3.4 hours for individual contributors, on the same tools. Business.com

How do you check an AI tool fits the tools you already use?

Open the integrations page before the pricing page. If your email, calendar, accounting software or CRM is not named on it, the tool means copying and pasting — and copying and pasting is where adoption quietly dies in week three.

  • Is your actual software named? "Integrates with 5,000+ apps via Zapier" is not a built connector, and usually means another subscription.
  • Where does the output land? A tool that writes a good draft into its own dashboard has moved the work, not removed it.
  • Does it work where the task happens? If that is between jobs or in the car, a desktop-only tool solves nothing.
  • Can you get your data out today? Not on a higher tier, not by emailing support. Today, in a format something else reads.

Be honest about which stack you actually have. The same MIT research found only around 40% of companies had bought an official subscription to a large language model, while staff at more than 90% were already using personal AI tools for work. What a business really runs on is whatever people open without being asked.

What should you actually score an AI tool on?

Five things, each with a condition that ends the trial rather than starting a debate. Write them down before you sign up, while you can still be unimpressed.

CriterionWhat to checkEnds the trial if
Fit with the taskIt does the thing you named in step oneIt does something adjacent and you would have to change the task to suit it
Reaches your stackA named connector for the app the work lives in"API available" with no ready-made integration
Time to first useful outputSignup to one usable result, timedMore than about 30 minutes
Cost against hours returnedMonthly cost versus hours saved × your rate, on a flat price or a cap you setUnder roughly 3×, or per-credit billing with no ceiling
Getting outExport everything, today, in a readable formatNo export, or export locked behind a higher plan
Five criteria, and the condition that ends the trial

How should you trial an AI tool?

On work you have already completed, so you can compare the output against an answer you know is correct. Pick three jobs from the last fortnight and mark the results against what you actually sent. You cannot judge quality on a task you have never done — anything fluent reads as competent when there is nothing to check it against, which is precisely the condition a demo creates.

A demo shows you the tool at its best on the vendor's data. You need to see it at its worst on yours.

How do you know whether an AI tool is actually working?

Re-measure the number from step two, two weeks in. Do not rely on how it feels, because how it feels is measurably unreliable.

The sharpest evidence on this comes from a METR study of 16 experienced open-source developers working on 246 real issues in codebases they knew well. With AI tools they took 19% longer. They had forecast a 24% speed-up beforehand — and afterwards, having actually been slower, they still believed the tools had sped them up by about 20%. METR, July 2025

METR is careful about the limits: small sample, expert developers on familiar code, tools from early 2025 the researchers now call out of date. It is not evidence that AI makes people slower in general. It is evidence that a 39-point gap can open between how fast you feel and how fast you are, which is why step five needs a number rather than an impression.

  • Minutes per run, timed again. Same task, same stopwatch.
  • How many days out of ten you opened it. Fewer than five is the answer on its own.
  • What you stopped doing. If nothing came off the list, the time went somewhere else rather than back to you.

Cancelling at two weeks is not failure, it is the framework working. A tool you drop after a fortnight cost you $20 and half an hour. A tool you keep out of optimism costs that every month for two years.

Where does this framework fall down?

  • It only evaluates work you already do. If the real opportunity is a service you have never offered, task-first filtering will not surface it.
  • Two weeks is too short for anything involving other people. A tool needing three colleagues to change a habit needs six weeks and a named owner, or a no.
  • It biases towards small, measurable wins. Mostly a feature — small wins stick — but it will pass over a tool that pays off slowly.

[Steve — add a short example of a tool that looked like a failure at two weeks and came good later, or the reverse.]

The version I run for clients is in the five filters I use, the money side in what AI tools cost a small business, and more about how I work on the about page.

Not sure which task to start with?

The free 3-minute scorecard walks through the same opening questions — what you repeat, where your week leaks, and what you already pay for.

Take the free scorecard
FAQ

Frequently asked questions

How long should I trial an AI tool before deciding?

Two weeks of real use, then re-measure. Two weeks is long enough for novelty to wear off and short enough that a wrong call costs one billing cycle. The exception is anything that needs several people to change how they work — give that six weeks and a named owner, or do not start it.

Is it better to use one general AI tool or several specialist ones?

Start general. A single well-known assistant will handle drafting, summarising and rewriting well enough to prove whether the task is worth automating at all. Only move to a specialist tool when you can name the specific thing the general one does badly — otherwise you are paying twice for overlapping capability.

What if I can't measure how long a task takes?

Use a proxy you can count. Number of emails sent, quotes issued, invoices chased, meetings written up. Anything countable beats an estimate, because the point of the number is not precision — it is having something fixed to compare against in a fortnight, when your memory of how bad it used to be has already softened.

How do I choose between two AI tools that both pass?

Take the one that is easier to leave. Compare exports, contract length and how many other things would break if you stopped paying. When two tools do the same job, the tiebreaker is not features — it is which one costs less to be wrong about.

Should I involve my team in choosing the tool?

Involve whoever does the task, at step one and step four. They know the minutes and they will spot the failure cases in a trial faster than you will. Do not run the selection by committee though — a tool chosen by five people to offend nobody usually solves nothing in particular.

Sources
  1. The 2026 AI Index Report — Economy — Stanford HAI (2026)
  2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR (10 July 2025)
  3. Fact of the Week: 20.5 Percent of Frequent Generative AI Users Report Saving Four or More Hours Weekly at Work — ITIF, reporting the Federal Reserve Bank of St. Louis survey by Bick, Blandin and Deming (9 May 2025)
  4. 2026 Small Business AI Outlook Report — Business.com (2026)
  5. MIT Report Finds Most AI Business Investments Fail, Reveals 'GenAI Divide' — Virtualization Review, reporting MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (19 August 2025)
Free · 3 minutes
Find out where your week actually goes.
Take the scorecard
Go deeper
The full AI Tools Assessment.
See how it works