DATIMORE / WEB-086 / v1.0
Read the articleDocument AI evaluation pack
Open the six source documents · Download the bilingual pack from the website (ZIP) · العربية
1. Agree what a useful record means
The record needs the delivery reference, date and received quantity shown in the source. All three are required and critical in this example. For this exercise, any wrong or missing answer to a readable field, or any guess the document cannot support, means the test fails. This is a teaching rule, not a universal business standard.
| Delivery reference / delivery_ref | Match the source text exactly, including capital letters. Ignore spaces only at the start and end. |
|---|---|
| Date / delivery_date | Use year-month-day, for example 2026-09-01, only when the date is clear. Do not guess whether the source puts the day or month first. |
| Received quantity / received_units | Whole units received, not units ordered. Treat Arabic-Indic digits such as ١٢ as the same value as 12. |
2. Compare the response with the source
The source documents, expected answers and responses written for this example are separate files in the ZIP. In a real trial, set aside the final test documents and expected answers. Do not use them to train or adjust the tool. Read each source before checking its expected answer.
D01
Expected: R-101, 2026-09-01, 8 units. The authored response matches all three.
D02
Expected: R-102, 2026-09-02, 12 units. The example response uses Arabic-Indic digits. All three values match once those digits are written as 0–9. One Arabic layout does not establish Arabic performance.
D03
Expected: R-103, unresolved date, 6 units. 03/04/2026 could mean 3 April or 4 March. The response 2026-03-04 is an unsupported guess; seek clarification from the document source.
D04
Expected: R-104, 2026-09-04, quantity absent. The response leaves the quantity blank, correctly avoiding a guess. The record still needs clarification before use.
D05
Expected: R-105, 2026-09-05, 7 units. The response says 10, the ordered quantity. Received quantity is essential in this example, so this wrong answer fails the test.
D06
Expected: R-106, unreadable date, 4 units. A whole-document processing failure is simulated with no returned values. Count the known reference and quantity as two missed values. Keep the unreadable date separate, and record the document failure.
3. Count errors without hiding difficult cases
The six documents contain 18 fields. The source gives a clear answer for 15: the example responses get 12 right, get 1 wrong and miss 2. Keep the other three fields separate: one date is ambiguous, one quantity is absent and one date is unreadable. The responses guess the ambiguous date, correctly leave the absent quantity blank and return nothing for the unreadable date because that whole document failed. None of these three counts as a correctly extracted value.
| Reference | 6 known: 5 correct, 0 wrong, 1 missed. |
|---|---|
| Date | 4 known: 4 correct, 0 wrong, 0 missed. Also report the ambiguous and unreadable dates separately. |
| Received units | 5 known: 3 correct, 1 wrong, 1 missed. Also report the absent quantity separately. |
| Decision | This practice test fails. In a real trial, address the errors and missing evidence, then test with fresh documents that were not used to train or adjust the tool. |
Keep the failed document in the submitted set. Record unsupported files and retries too. Report results by language, layout and copy quality. There are too few documents here to judge performance by language, layout or copy quality.
4. Account for staff time and cost
Practice assumption: every one of the six documents is checked for two minutes. D03–D06 each require two extra minutes for correction or follow-up. Total: 6 × 2 + 4 × 2 = 20 minutes. Assume manual entry takes three minutes per document, or 18 minutes. This example shows no saving.
At an assumed SAR 90 per staff hour, that is SAR 30 for checking and follow-up versus SAR 27 for manual entry. No person was timed. These figures exclude waiting for clarification, tool usage, retries, storage, setup, labeling, integration and support. Add these in a real comparison; report waiting time separately from active work.
| Coverage and errors | Document types, languages, critical fields and acceptable error limits: __________ |
|---|---|
| Review capacity | Reviewer and backup, minutes per document, exceptions and daily capacity: __________ |
| Cost | Volume, currency, usage, one-time work and support costs: __________ |
| Operating fit | Access, retention, integration, next decision and owner: __________ |
5. Save enough detail to repeat the test
The ZIP contains the worked run log and a blank log for your own trial. Record the tool and model version, document set and expected answers. Save the settings, rules for matching answers and what happens if processing is tried again. If the tool uses a confidence score to decide which answers to return, record that cutoff too. That score is the tool’s estimate, not a guarantee of a correct answer. Keep the original responses and failures. Save a hash for each file: a verification code that changes when its contents change. This helps confirm that later comparisons use the same files. If a setting or expected answer changes, repeat the test and save a new result. Explain any differences in test conditions before comparing tools.
To repeat the calculation, you or a technical colleague can unzip and run python score.py with Python 3. It makes no network calls and performs no extraction; it only compares the example responses with the expected answers. See README.txt for the file list.
Before a real trial, add authorized documents representative of everyday work, including poor scans and unsupported files. Keep the final test documents separate from those used to train or adjust the tool. Agree the expected answers and acceptable errors with the people responsible for the process. Passing this small exercise does not show how accurate a vendor’s tool is or whether it is ready for everyday use.