discover · framework July 08, 2026 11 min

How I Evaluate a New AI Tool in 20 Minutes

My framework for screening AI tools without the marketing demo: real data, deliberate breaking, the cost of wiring it in. 20 minutes, not a two-week pilot.

AI toolsevaluation frameworkchoosing AIproductivityAI for business

Andrew Maryasov, AI consultant. Over three years I’ve been through more than 80 AI tools on real projects. A handful made it into my daily stack. The rest were filtered out by a procedure that takes twenty minutes and not a single vendor call.

A 20-minute framework for evaluating an AI tool, in three blocks: run it on real data, break it on purpose, count the cost of wiring it in — instead of the vendor demo

TL;DR (in 30 seconds)

Why I stopped watching demos

A few years ago my routine was this: find a tool, read the landing page, watch the demo video, book a call. On the call everything worked. The model pulled fields out of a PDF, the agent answered a customer, the dashboard drew charts. I’d sign up for the monthly plan, and three weeks later I couldn’t remember the password.

It’s not that vendors lie. A demo is simply a controlled environment: the vendor chose the file, the vendor wrote the prompt, the vendor knows the question where his model loses its footing. And that’s exactly the question where the call quietly moves on to the next slide.

The corporate world learned to live with this long ago. Its answer looks like this: a weighted scorecard across eleven categories, a paid two-week proof of concept, a legal department, SOC 2 Type II, quarterly governance reviews. It works, and for a five-hundred-person company that’s exactly what I’d do. The central claim of those checklists is honestly hard to argue with: the only number that matters is how the model performed on your data.

But if you’re picking a tool for yourself or for a team of eight, you don’t have two weeks. No legal department, no one to assemble an evaluation matrix in Excel. You have a Tuesday evening and a wish to work out whether this thing deserves your attention.

So I keep a short procedure. Twenty minutes, three blocks, one sheet of paper.

The twenty-minute rule

Twenty minutes don’t answer the question “should we roll this out.” They answer the only question that belongs to this stage: is this tool worth another two days?

Most tools die right here, and that’s the whole point of the exercise.

I really do set a timer, and it isn’t a pose. Without one it’s easy to sink into configuring integrations, burn half a day and come out with a sense that “it seems interesting” instead of an answer. A week later that sense turns into one more subscription you only ever see on your card statement.

Block 1. Real data, minutes 0–7

The first thing I do after signing up is close the onboarding tab along with its test file sample_invoice.pdf.

Instead I take three cases of my own. Not average ones. The worst I can find:

Everything else is a hygienic version of reality, and in it any tool looks smart.

Then I measure one number: how many of the three outputs I’d accept with a light edit rather than rewriting from scratch. My bar is two cases out of three, and the bar is deliberately low here — with three cases there’s no more precise arithmetic to be had.

When a tool gives you a weak result, your hand reaches out on its own to rewrite the prompt, then again, then once more. On the seventh iteration you finally get something good, except you got it, not the tool. If a decent output takes seven passes and half a page of instructions, the tool has lost. I just don’t want to admit it yet, because I’ve already spent an hour on it.

Block 2. Break it on purpose, minutes 7–14

The happy path shows you nothing — everything interesting starts on the bad road.

An empty field. A request with the key data missing. The model should say “not enough data.” If it confidently invents an answer, in production that costs more than silence: nobody checks an invented answer, because it looks like an answer.

A contradiction. Two documents with different figures on the same deal. Good behavior is to point at the discrepancy. Bad behavior is to quietly pick one of the figures, and you’ll never learn which.

Out of scope. A question about something that isn’t in the data you handed over at all. What’s being tested here isn’t intelligence but the willingness to admit not knowing.

When I was testing agent orchestrators (n8n, Mastra and a few less well-known ones), the scenario where all three steps worked wasn’t what interested me. What interested me was what happens when the second step returns an empty response. Some stop and shout. Others neatly pass the emptiness along, and at the end you get a flawlessly formatted report about nothing: headings in place, conclusions in place, figures pulled out of thin air.

The first kind you can leave unattended. The second one you’ll have to babysit, and the babysitter is you.

Block 3. The cost of wiring it in, minutes 14–20

The subscription price is the smallest and most transparent part of the cost. The rest you have to count yourself, and this is where most “cheap” solutions die.

I add up four numbers:

What I countHow I estimate it in 6 minutes
Hours of integrationIs there a native connector to what I already use? Will I have to write a wrapper?
Cost of relearningHow many people have to change a habit, and how badly
Cost of exitCan I export my data to CSV or JSON without calling support
Cost of an errorWhat happens if the tool is wrong in five percent of cases: an inconvenience or a lost client

A stronger tool that has to be stitched in by hand regularly loses to a weaker one with a native integration. Not because quality doesn’t matter. Because forty hours of integration is forty hours I don’t have, and “I’ll finish it later” has never once arrived.

What survived all this and what it costs is in the breakdown of my working stack; that one also covers why its most expensive part isn’t the subscriptions.

The twenty-minute scorecard

BlockMinutesQuestionPassing grade
Real data0–7How many of the three dirty cases got through without a rewrite?2 of 3
Deliberate breaking7–14Does it say “I don’t know”? Does it stop after an error?Yes to both
Cost of wiring it in14–20How many hours to the first useful result inside my own process?no more than 8

Save this table. Next time you open a landing page with “AI-powered” in the headline, you’ll have something to look at besides the “Start free trial” button.

One failed block means no. Just no, without “let’s give it another chance” and “let’s try it on a different case.”

It took me a long time to learn not to hand out second chances, and every lesson came priced as an annual subscription.

Red flags that make me close the tab on the spot

Sometimes the twenty minutes never even start. Here’s what stops me:

What happens in the twenty-first minute

There are exactly three ways out here, and none of them is “it’ll sort itself out somehow.”

The tool passed all three blocks — I give it two days on a real flow of work, with the same criteria, just at a bigger volume. Two-thirds of what makes it this far stays in the stack for good.

The tool failed the cost block but won the first two — I don’t throw it out, I write it into a separate file with a date. Six months on, connectors appear, prices drop, and half of those entries come back to life.

The tool failed on real data or on breaking — I close the tab and don’t come back. No discount and no update ever cures a tool that invents answers.

The hardest part, as always, is me rather than the tools. Twenty minutes is exactly as long as it takes not to have time to fall in love.

FAQ

Is 20 minutes enough to choose an AI tool for a company?

No, and that isn’t the goal. Twenty minutes filter out candidates. For a company-level decision, what comes next is a two-to-four-week pilot with real users, written pass/fail criteria and a measured baseline. The framework saves you ten pointless pilots, not the decision stage itself — it saves you ten pointless pilots.

What if the tool won’t give me access without a sales call?

Take the call, but not for a demo. Ask for a sandbox with your data and your prompt. A vendor who agrees is usually confident in the product. A vendor who insists on a guided presentation is telling you something important before you’ve signed anything.

How many tools should I test in parallel?

Three to five. Fewer and there’s nothing to compare. More and by day four you’re mixing up which one could handle tables. The main thing: they all get the same dirty cases and the same criteria, otherwise you’re comparing your own mood rather than the tools.

How do I test on real data without breaking confidentiality?

De-identify the sample: swap names, contract numbers and amounts, keeping the structure and the “dirt.” For sensitive data, look at self-hosted or enterprise plans with zero retention. The data retention policy belongs in the contract, not in the FAQ on the website.

How is this framework different from a corporate vendor evaluation?

Scale and cost of error. A corporate procedure protects a multi-month contract and several stakeholders, so it comes with governance, audit and weighted scorecards. A personal framework protects your time and your evening. The logic is the same: test on your own data, fix the criteria in advance, count the full cost rather than the price of the license.

The tool passed all three blocks. What now?

Two days on a real flow of work and one question at the end: what breaks if this tool disappears tomorrow. If the answer is “nothing,” you don’t need it. If the answer is “half my work,” you’ve just found a dependency, and now it’s worth going back over the exit terms more carefully.

How many tools survive into daily use in the end?

Out of the 80-plus I’ve tested, nine are left, and that’s a normal ratio. Most of the rest weren’t bad. They just solved a problem I don’t have.

What’s next

The framework answers the question “is this worth a closer look.” It doesn’t answer “which three tools will move your business specifically” — that needs the context of your processes, not a timer.

You can run this screen yourself in one evening. If you’d rather do it with a second pair of eyes — get in touch; how long it takes depends on how well your process is already documented.

See also:

Andrew Maryasov

Founder of Auspex and Grow2.ai. 28 AI projects across 14 industries and 7 countries, 8 proprietary AI systems in production.

Newsletter

A letter on AI — 2–3 emails a month

Personal essays on how AI changes the way we think and work: what I tried myself, what worked and what didn't. At the end, links to the best of what I read.

No spam. Unsubscribe in one click.

Free 30-min discovery →