← Insights
Insight · insight · ai · decision-frameworks

The AI investment framework nobody uses

The best AI investment most businesses will make is the one they decide not to make. Six questions, a scoring system, and the restraint to follow what the numbers say.

Published 24 February 2026 · 5 min read

The best AI investment most businesses will make is the one they decide not to make.

The most valuable output of an AI evaluation isn't a shortlist of things to build. It's the "don't build" list — the line items you would have spent $20K, $50K, $100K on if you'd followed enthusiasm instead of evidence. Knowing what not to invest in saves more than anything you build.

But you only get that list if you evaluate AI the right way. Most businesses evaluate it like a technology purchase. What does it do? How much does it cost? Can we get a demo? Reasonable questions for buying a printer. Wrong questions for a capital investment that redirects labour, shifts decision-making, and alters what's possible at a given headcount.

Six dimensions separate good AI investments from expensive distractions. None of them are about the technology. And one of them matters more than the rest.

The one that overrides arithmetic: pain

Pain is the dimension nobody scores and everyone feels.

A function that scores moderately on every other dimension, but is the thing the owner complains about every Monday morning — that one gets built. Because the person who writes the cheques will champion it. Without that champion, even high-scoring projects stall at month four.

Inversely: a project the spreadsheet loves but nobody in the business cares about will never ship. It'll pass every gate review and die quietly in the last mile.

So the first question I ask isn't about volume, or cost, or data. It's: what does the owner complain about on a Monday? That's where to look first. Score the rest against that one.

The other five

Can you measure the output?

The hard constraint. If the function you're automating doesn't have a clear pass/fail — if the quality of the output is a matter of opinion — you cannot prove the system works. No amount of enthusiasm overrides this.

A compliance assessment either meets the building code or it doesn't. An invoice either matches the purchase order or it doesn't. A support response either resolves the ticket or it doesn't. These are measurable. "Write better marketing copy" is not. If you can't define what good looks like before you build, stop here.

How much judgment is involved?

Some functions are mechanical — identical every time. A script handles those. You don't need AI. At the other end, some functions require years of experience per decision. AI can assist, but it can't replace the judgment.

The sweet spot is in the middle — work that follows a methodology but where inputs vary. That's where AI earns its keep.

Is the data already digital?

If the information the function needs lives in someone's head, in handwritten notes, or in formats no system can read — you have a data capture problem, not an AI problem. Fix that first. It'll be clearer what AI can do once the data is clean.

AI built on bad data doesn't produce bad results. It produces confident bad results, which is worse.

Volume and cost.

Together these are the arithmetic. Eight compliance assessments a month at six hours each is 576 hours a year of senior time. The ROI calculation writes itself. Two a quarter, and the system costs more than it saves. The fully loaded cost — hourly rate plus the work you're turning away because the bottleneck exists — is the number that matters, not the subscription fee.

Both are scoreable in an afternoon. They're commodity dimensions. They don't need much airtime — but they're the dimensions where the evidence either confirms the case or kills it.

How the scoring works

Each dimension scores 1 to 5. Total out of 30.

  • 20+: High priority. Design the system.
  • 14–19: Check the flags. One low score in the wrong place kills the case.
  • Under 14: Put it on the "don't build" list.

The flags matter more than the total. Measurability at 1 or 2 is a hard stop regardless of everything else. Judgment at 1 or 2 means a script will do — you're over-engineering with AI. Pain at 5 means the owner will champion it even if the total is moderate.

What to do with what's left

Every business I look at has functions that score well on one or two dimensions but fall apart on the others. The natural instinct is to focus on what scored high. The discipline is what you do with the rest.

The "don't build" list.

Six questions, a scoring system, and the restraint to follow what the numbers say. No technology required to start using it. You can score your own operations this afternoon.

Run the scores. Build the don't-build list first. If something survives it, then we talk about what to build.

Karl Howard · Reforged · 24 February 2026