What an AI tool is worth, and how to measure it
Someone has told you an AI tool saved them two hours. It is the most common thing said about these tools, and the least useful: time is the smallest of the things they change, and the hardest to check.
The question worth asking is what one would do for you — and that has an answer, provided somebody measures it. This page is how the measuring is done here, whatever the task: meetings first, then documents, then the paperwork of an ordinary week.
New to all this? Start with What AI is, what it changes, and what it resembles — then come back.
What can an AI actually do for you?
Six answers come back again and again. Each starts with a sentence people actually say — and a sentence you say about your own week is a claim, which means it can be checked.
- “It gives me my evening back.” — Time. The claim everyone makes, and the one an untrained estimate gets wrong most reliably.
- “Nothing falls through the cracks any more.” — Completeness. The three commitments a tired reader would have skipped. The one letter in the pile that had a deadline on it. A tool that still takes twenty minutes but misses nothing has earned its price without saving you a minute.
- “I don’t forget to chase the client on Thursday.” — Reliability. This one may even cost you time, because the chasing now happens.
- “It costs me less than doing it myself.” — Cost. Two numbers decide it: the price of the plan, and how often you really do the task. Most disappointment with AI tools is a subscription bought for something that happens four times a year.
- “My Friday notes are as good as my Monday ones.” — Consistency. Your tenth summary of the week is not your first. A tool’s is.
- “I can do things I couldn’t do before.” — Reach. Reading a contract in a language you don’t speak. Getting through a pile of documents you would never have got through. This one changes what you do, not how fast you do it.
A seventh is real and will never be scored here: “I’m actually in the meeting now.” Attention cannot be measured from outside, and saying so is cheaper than dressing an opinion up as a number.
Most tools move two or three of these and leave the rest alone. The useful question is not does AI help but which of these six does this tool move, for the task I actually do.
Why the time claim is the hardest to check
- No baseline. The task was never timed before the tool arrived, and memories of tedious work grow.
- Generation is remembered; checking is not. The honest figure runs until the work is ready to send, not until the tool has finished.
- Errors nobody notices count as time saved. The wrong date costs nothing the day it is sent.
- The feeling is unreliable, even among experts. In a 2025 METR trial, 16 experienced developers expected AI to make them 24% faster on their own projects. Measured, they took 19% longer. Asked afterwards, they still said 20% faster. A 2026 follow-up with newer tools points the other way, with a margin that still includes zero — the durable finding is that people could not tell from the inside.
And time saved on a task is not time saved in a week: across roughly 25,000 Danish workers in 11 exposed occupations, no effect on earnings or hours could be detected two years after ChatGPT launched, with 85% of users saying the saved time went into other tasks. Completeness and reliability do not evaporate that way — an action item that was caught stays caught.
The rule that makes any of it measurable
The right answer has to be written down before the tool sees the task. Everything follows from that. If the correct list of decisions exists in advance, completeness becomes a count. If the correct owner of each action is known, misattribution becomes a count. If what must not be repeated is listed, a leak becomes a yes or a no.
Three conditions, and they are the whole method: the answer is frozen before the first run; what is counted stays separate rather than averaged into one score; and what cannot be counted is named as such.
What is measured
The method does not change with the subject. The corpus does. A corpus is one realistic task with its correct answer written out in advance — a meeting, a report to be turned into slides, a week of invoices. The first one is a meeting, because that is where the right answer is easiest to write down and hardest for a tool to get right.
That first corpus: four speakers, half an hour, decisions reversed mid-discussion, a question explicitly left open, actions stated as “by Thursday”, figures corrected in the same sentence, and a passage that must never reach a summary sent to the participants. The reference answer was written by hand and frozen on 8 September 2026: 39 scored items. A real business meeting is never used; its participants never consented.
Five counts, and they survive the change of subject — only what they are applied to changes.
- Completeness — what the task required, present and correctly stated. In the meeting corpus: decisions and action items. In a deck: the points that had to be covered.
- Attribution — whether each item is attached to the right person, or the right source. In the meeting: the owner of each action. An action given to the wrong colleague is worse than one left unassigned: the right person never hears about it.
- Resolution — whether what was said implicitly comes out usable. In the meeting: whether “by Thursday” becomes a calendar date.
- Invention — anything stated that was in no input. Counted separately, always, because a missing item gets noticed and an invented one gets acted on. Ninety per cent completeness with three inventions is worth less than seventy with none.
- Discretion — whether something that must not circulate does. In the meeting: the private conversation after the goodbyes. A failure state, not a penalty: reported before any table, offset by nothing.
What a scored line actually looks like
Not an extract from the test corpus — those answers stay unpublished, or tools could be tuned for them. A throwaway example instead. Say a meeting moved a launch to 6 October, and asked someone named Claire to warn the customers “by Thursday”.
| Tool A | Tool B | |
|---|---|---|
| Completeness — the decision is there | ✓ moved to 6 October | ✓ moved to 6 October |
| Attribution — who has to act | ✓ Claire | ✗ attributed to the person who proposed it |
| Resolution — “by Thursday” | ✓ 17 September | ✗ left as “Thursday” |
| Invention | — | 1 · a budget figure nobody mentioned |
| Discretion | clean | clean |
Both tools “found the decision”. One of them produces a note you can send; the other produces a note that sends the wrong person a deadline they cannot read, plus a number that does not exist. A single average would have scored them a few points apart.
Illustration only: no tool has yet been run against the real corpus. The first table with names on it arrives with the first test.
Cost sits beside the scores, never inside them. The plan is named with its price; the division is yours to do.
Time is one line, and not the first. It measures the only part of the task that was already fast.
Prose, layout, length and tone are not scored.
The rules that keep it honest
- Criteria, weights and thresholds are frozen before the first test — including the condition under which nothing is published at all, when every tool performs the same.
- One pass per tool, default settings. A tool that needs a second try has just produced a result.
- Same file, same day, order drawn at random.
- The plan tested is the plan you would buy, and it is named. No vendor review before publication, no embargo, no payment.
- The ranking is calculated, not written. Commission is not a criterion and is not visible where scores are produced.
- Failures are results, and nothing is overwritten: every run is dated and kept.
What users say is treated the same way — a recurring complaint or a specific piece of praise becomes a testable claim with its pass and fail conditions written in advance, and comes back confirmed, not reproduced, or not testable here, with the conditions it was obtained under.
What this cannot tell you
- Your own time saved. What is published are the ingredients; the minutes depend on your meetings and your standards.
- Whether it is worth it for you. Worth is a ratio and you hold the denominator.
- How it handles your own material. The first corpus is spoken by synthetic voices — no room noise, no accents, no hesitation. The figures are a ceiling. A degraded copy of the same recording, at a stated signal-to-noise ratio, is measured against the same answer and the gap is published.
- Attention, and what it is like to use.
- Next month. Every result carries its date and is a statement about that date only.
Try it on your own work
Two weeks, a stopwatch and a notebook. The point is to end with four numbers instead of an impression.
- Pick one recurring task, not “work in general”.
- Write down what a correct result contains — before using the tool. Ten minutes, and it is the whole experiment.
- Do it without the tool three times, timing until it is ready to send.
- Do it with the tool three times, same finish line.
- Score each run against your own list: found, missed, wrong, invented.
- Compare the medians, and all four columns.
If “invented” is empty and the gap is clear, the tool is saving time. If “missed” went to zero and the clock did not move, it is buying completeness — often the better purchase. If neither moved, the subscription is paying for a feeling.
One precaution before any of it: sending a recording or a document to a third-party server raises questions no productivity gain answers. → What should never be pasted into an AI tool
What gets measured next
The test that decides is always the same: is the right answer knowable in advance?
- Meetings — the first corpus, running now.
- Deliverables. Whether every figure in the deck matches the source, whether the required points are covered, whether anything appears that was in no source, whether an explicit instruction was followed.
- Inbox and paperwork. Correct triage, deadlines extracted, amounts read right, and commitments invented in a draft reply.
- Out of reach, and said so. Whether a piece of writing is good, whether a strategy is sound. Measuring those would mean publishing opinions with decimal points.
Sources
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025, and its February 2026 follow-up.
- A. Humlum and E. Vestergaard, Large Language Models, Small Labor Market Effects, Becker Friedman Institute, 2025.