Kalytia tests AI tools on real work and publishes the numbers.
Who this site is for
Anyone curious about AI: what it actually is, what it does well, and how to use it in an ordinary week — at work and at home. No technical background is assumed, and nothing here requires one.
The tests start where the stakes are sharpest, with people who choose and pay for their own software — freelancers, consultants and independent professionals — because they carry the cost of a bad tool directly and cannot pass the question to an IT department. One subject is tested at a time, meetings first, then documents, then the paperwork of an ordinary week. The articles are written for everyone; the test bench simply had to start somewhere.
Why this site exists
Search for the best AI tool in almost any category and the same article comes back fifteen times: a list of products, a paragraph of marketing copy for each, a table of features copied from vendor websites, and a recommendation that correlates suspiciously well with commission rates. Almost none of it involves anyone actually using the products.
The boring part gets done here instead. One realistic task, the same input for every tool, a scoring grid written before the first test, and the raw outputs kept so that anyone can check the arithmetic.
How the tests work
Every comparison runs on a fixed corpus. One realistic task whose correct answer is written out before any tool sees it. The method does not change with the subject — only the corpus does. The first corpus is a meeting, because that is where a correct answer is easiest to write down and hardest for a tool to get right: four speakers, half an hour, decisions reversed mid-discussion, one question deliberately left open, action items of which one has no owner. Documents and paperwork follow.
All tools receive the same input on the same day, in a randomised order. Timing measurements are repeated three times and the spread is kept. Qualitative scores are given blind, against a scale with worked examples, and any criterion where the ratings disagree too much is excluded and flagged rather than quietly averaged.
Every article ends with what was not measured. That section is not modesty — it is the part that tells you how far to trust the rest.
The full method — what gets counted, and the rules fixed before the first test — is set out in What an AI tool is worth, and how to measure it.
How to read a test
Every comparison published here follows the same structure, and each element has a reason.
- A band under the introduction gives the test date, the material used, and the number of criteria and runs. Without it, the article is not a test.
- The results table comes first, above the fold, not after two thousand words of preamble.
- The plan tested is named. Quality is measured on the plan a reader would actually buy.
- Inventions are reported separately from omissions. An output that recovers 90% of what mattered but invents three things can be worth less than one that recovers 70% and invents nothing.
- A criterion marked unreliable is one on which independent ratings disagreed too much. It is mentioned as an observation and never scored.
- Every article ends with what was not measured.
Who writes this
Samuel Benzakour, based in Lyon, France. Business analyst in digital transformation, project management officer by trade, and an enthusiast of the technology tested here.
Articles are drafted with the help of software built for the purpose, from measurements run here. Every article is read, corrected and signed by a human before publication, and every number in it is traceable to a dated test run.
How this site is funded
How this site earns money, and every rule that keeps it away from the scores, is set out in full on the Affiliate disclosure page.
What will never be done
- Publish a comparison that was not run.
- Let a commission rate influence a score.
- Send a draft to a vendor for approval.
- Recommend a tool when the honest answer is that the differences do not matter.