Start with the record people need
For a receiving team, the useful result is a delivery reference, date and received quantity that agree with the original receipt. Staff can check the prepared record, resolve questions and continue their work. A tool that fills every box but confuses ordered and received quantities would undermine that improvement.
Choose a small set of document types first. Include the languages, layouts and unclear copies your team actually receives. Have people agree the expected answers before comparing tools. Set aside some documents until the final test. Do not use them to train or adjust the tool.
Compare each important field, not just one overall score. Count correct values, wrong values and missing answers separately. If the document does not give a clear answer, flag the value for review instead of guessing. Google's evaluation guidance explains that a score depends on which fields are checked and what counts as a match.
See what one wrong quantity would change
Synthetic example: an invented receiving team wants prepared delivery records it can check quickly. Receipt R-105 says 10 units ordered and 7 received. We wrote a sample answer that incorrectly lists 10 units received. No AI tool produced this response, and no vendor accuracy was measured.
Expected: 7 received. Sample response: 10. The evaluation records one wrong quantity. In this example, received quantity is essential, so that error blocks a move to routine use until the cause is addressed and a fresh test passes.
The comparison worksheet and synthetic test pack let you repeat these checks. Six invented receipts include Arabic text, an ambiguous date, a missing quantity and a simulated processing failure. Expected answers stay separate from sample responses. The pack shows how to count failures without treating an unreadable value as a known answer.
Leave room for the people doing the checking
Better prepared records only help if checking them fits the working day. During a real trial, time the reading, correction and follow-up work as well as the software run. Compare that with the current manual routine. Include setup, usage charges and continuing support. Correctly copying most values does not prove that the tool will cost less to use.
Agree acceptable errors and available review time before testing. Keep results by language and document type so a strong result on clear English receipts cannot hide weak Arabic or damaged-copy results. If an essential value is wrong or missing, or staff cannot keep up with the checking, test fewer document types or continue entering those records by hand.
Save the document set, expected answers, tool version and settings with each result. If you change a setting or correct an expected answer, repeat the test and save it as a new run. Our bilingual pack to download includes a blank run log and a worked cost worksheet. Its figures are assumptions for practice, not observed performance. Six invented receipts cannot show whether a tool is ready for everyday use.
Datimore can help connect an agreed document process with your systems through a focused integration project. Bring one recurring document type, the record staff need and the people who would review it. The goal is a useful starting record and manageable checking work, backed by evidence from your own trial.