The AI feasibility diagnostic, and a sample report
One week, a fixed price, and a measured answer on your own data. Below is the offer, how the week runs, and a full sample of the report you receive.
AI feasibility diagnostic
Before anyone spends money on a build, find out whether AI can do the job on your data, how well, and at what cost.
What you get
- A review of the problem, the process around it and the data you have.
- A measured test on a sample of your real data, run the way the Lab runs its experiments: what works, what fails, and how often.
- A written recommendation: build, change the process, buy something off the shelf, or do nothing.
- If building makes sense, two or three fixed-price options for the build, with timelines.
€1,900
Fixed, final price.
Order the build within 60 days and the full €1,900 is deducted from its price.
Send a project briefBuilds are quoted per stage at a fixed price, typically from €8,000.
How the week runs
- Before it starts: you send a project brief. I reply within two working days with questions by email and a short written proposal. If you need one, I sign your NDA before you share anything.
- Days 1–2: I review the process, the systems around it and a sample of your data, usually 50–200 documents or records, anonymised if you prefer.
- Days 3–4: a measured test on that sample, run the way every Lab on this site is run: a fixed method, recorded results, and the failures counted as carefully as the successes.
- Day 5: you receive the written report, and we go through it by email, or in a short call if you want one.
What the report looks like
The sample below answers a typical client question, “can our receipt and invoice entry be automated?”, using the public data and measurements from LAB/002. It has the same structure and level of detail as a real report. Only the client is missing: no client data appears anywhere on this site.
SAMPLE REPORT · AI feasibility diagnostic · Receipt and invoice entry
1. Summary and recommendation
Recommendation: build, with checks. About seven in ten documents can be processed with no person involved, and in this test every one of those had the correct total. The rest go to a person, and the system tells them which check failed. Line-item names are the weak point: the checks cannot see a misread item name, so if item names drive anything downstream (stock, cost centres), they need their own check.
2. The question
Staff retype the key fields of every receipt by hand. Can a system do it, how much of the work can it take over safely, and how often will it be wrong on the part it takes over?
3. Data and method
200 receipts (the CORD-v2 public dataset, CC BY 4.0), photographed, with checked ground truth. An open-weight vision-language model, running locally on one GPU (no data leaves the machine), extracts a fixed set of fields. Plain code then checks the arithmetic every receipt must satisfy, and only receipts that pass every applicable check are accepted automatically. The prompt was written before the run and not tuned on these receipts.
4. Results
| All receipts | Auto-accepted | Sent to review | |
|---|---|---|---|
| Receipts | 200 | 145 | 55 |
| Fully correct | 65.5% | 80.7% | 25.5% |
| Total correct | 97.4% | 100.0% | 90.6% |
| Line items F1 | 0.88 | 0.94 | 0.76 |
Of the receipts accepted automatically, about 81% were entirely correct and all had the correct total. Of those sent to a person, only about a quarter were entirely correct: the checks concentrate the problems where someone looks.
- total n=19397%
- subtotal n=13299%
- cash n=13196%
- change n=12095%
- tax n=8883%
- service n=2564%
- discount n=1292%
- line items (F1) n=47288%
5. What fails, and why

| Field | Model output | Ground truth |
|---|---|---|
| subtotal | 18,181 | 18,181 |
| tax | 10% Tax Included | 1,818 |
| discount | 50% | — |
| total | 20,000 | 20,000 |
| cash | 100,000 | 100,000 |
| change | 80,000 | 80,000 |
| items | S-Ovaltine: 20,000 | |
- Failed: sum of line prices = subtotal (or total if no subtotal)
- Failed: subtotal + tax + service − |discount| = total
- Passed: cash − change = total
This receipt looked fine at a glance: well-formed output, plausible values. The model copied “10% Tax Included” into the tax field and read a sugar level in an item name as a discount. Both arithmetic checks failed, so it went to a person. Of the receipts accepted automatically, 28 still had an error, 23 of them in line items: a misread name, or a line split or merged while the prices still added up.
6. Limits of this test
One document type from one country, one model and one prompt. The labels themselves contain mistakes, and with these rules even a perfect extractor would be accepted automatically on only about 82% of receipts; loosening a rule raises automation and lets more errors through. On your documents, these numbers will differ; that is why the test is run on your sample.
7. Options
| Option | What it does | When it makes sense |
|---|---|---|
| A. Assisted entry | The system pre-fills every field; a person confirms each document. | Low volume, or when every field must be confirmed for compliance. |
| B. Automated with checks (recommended) | Documents that pass every check are posted automatically; the rest go to a review queue with the failed check shown. | Steady volume, totals matter more than item names. |
| C. Don’t build | Use an off-the-shelf OCR product, or keep manual entry. | Few documents a month, or one fixed template where a simple parser is enough. |
In a real report each option comes with a fixed price, a timeline and the running cost per document. Builds are quoted per stage at a fixed price, typically from €8,000, and the diagnostic fee is deducted if you order within 60 days.
END OF SAMPLE
Every number in this sample comes from the recorded run behind LAB/002, with its code and data published. Your report would be about your process and your data, and stay private.