- Exact match accuracy is the share of values that match the correct answer exactly. There is no partial credit: a name with one wrong letter, or an amount one cent off, counts as wrong.1
- Most mistakes are in contact details - addresses and emails - not in amounts.
- Anyone can repeat this test with public data (section 9).
Terms used on this page
- Field
- A type of information on an invoice, such as the invoice number or the total.
- Value
- One field on one invoice - for example, the total on one invoice.
- Answer key
- The correct value of every field, taken from the answers published with the dataset.
- Exact match accuracy
- Correct values ÷ all values checked. A value is correct only if it matches the answer key exactly. Letter capitalization and extra spaces are ignored, and
$andUSDare treated as the same.1 Any other difference, even one character, makes the value wrong. - Invoice layout
- One invoice design (template). FATURA has 50 layouts, with 200 invoices each.
- Test set
- The 1,250 invoices the dataset's creators set aside for testing.
- AI Judge
- A second AI pass that checks the first reading against the invoice image and corrects any errors that it discovered.
- Postprocessing rules
- Fixed formatting rules with no AI, such as turning
(-) 9.39into a discount of 9.39.
1. Why we publish this
Most vendors quote an accuracy number without disclosing how it was arrived at. “99% accurate” can mean 99% of characters, 99% of fields, or 99% of documents.
A number is only useful with three things attached: a public dataset, a test set the vendor did not choose, and a written scoring rule. This page gives all three.
2. The test invoices
We used FATURA (paper (opens in a new tab) · dataset (opens in a new tab)), a public research collection of 10,000 invoice images: 50 invoice layouts, 200 invoices each. The correct value of every field is published with each invoice. We do not decide what the right answer is - the dataset does.
FATURA's creators split the invoices into three groups: 7,500 for training, 1,250 for development
and 1,250 for testing. We scored the whole test group and nothing else: the
1,250 invoices listed in the dataset's own file strat1_test.csv. The
99.1% comes from these 1,250 invoices only. FigureIQ was never trained or tuned on them, so it saw
them for the first time in this test.
Figure 1 - How FATURA's 10,000 invoices are divided. Every layout contributes 150 training, 25 development and 25 test invoices.
What the test invoices look like
The 50 layouts differ the way real suppliers' invoices do: the seller block sits at the top on some and at the bottom on others, some print one tax line and some several, and some show “Balance due” instead of “Total”.
Figure 2 - Eight test invoices, each from a different layout.
3. How FigureIQ reads an invoice
Each invoice goes through three steps, much like an accounts team where one person keys in the invoice, a second checks it against the paper, and house rules put everything in the same format:
- AI Extraction reads the invoice image and fills in the fields.
- AI Judge, a second AI pass, checks those fields against the image and corrects what looks wrong.
- Postprocessing rules (no AI) put values into a standard format - for example, a printed
$becomes the currency codeUSD, and(-) 9.39becomes a discount of 9.39. The same input always gives the same output.
Figure 3 - The three steps. You receive only the final result.
4. What we checked
Figure 4 - The 26 fields checked. Tax lines are checked only on invoices with more than one tax rate.
How we matched FATURA's answers to FigureIQ's fields
FATURA and FigureIQ name the same information differently. For example, what FATURA labels
BUYER, FigureIQ calls Customer. Some FATURA labels also hold several values at once,
such as a whole tax line. So before we can compare anything, we need a mapping: a
list of which FATURA label is compared with which FigureIQ field.
- One-to-one mapping. One FATURA label maps to
one FigureIQ field. Example:
NUMBER→ Invoice number. - One-to-many mapping. One FATURA label maps to
several FigureIQ fields. Example:
TAX→ tax name, tax rate, tax amount and currency (Figure 6). - Conditional mapping. Used only in some cases.
BILL TOmaps to the customer only when there is noBUYER.GST(n%)maps to the tax-line list only when several GST lines are printed.
Figure 5 - Mapping of FATURA labels (left) to FigureIQ fields (right). Currency has no FATURA label of its own; it is taken from the symbol printed with the amounts.
For technical reviewers: the same mapping as a table, and why it is 26 checks
| FATURA label | FigureIQ fields it checks |
|---|---|
NUMBER | Invoice number |
PO NUMBER | PO number |
DATE | Invoice date, invoice date (calendar) |
DUE DATE | Due date, due date (calendar) |
TOTAL | Total, currency |
AMOUNT DUE | Total, currency Some layouts print only a TOTAL, some only an AMOUNT DUE, and some print both. If a layout prints both, we count the total only once, so the same field isn't marked twice. |
SUB TOTAL | Subtotal, currency |
TAX | Tax name, tax rate, tax amount, currency |
GST(n%) | Tax name, tax rate, tax amount; as a list of tax lines when several are printed |
DISCOUNT | Discount rate, discount amount |
SELLER | Seller name, address, email |
BUYER | Customer name, address, email, phone |
BILL TO | Customer name, address, email, phone - only when there is no BUYER |
SEND TO | Ship-to name, address, email, phone |
FigureIQ returns 72 fields from an invoice. FATURA's answers use
41 labels. This is how the two line up:
- Matched: 25 FATURA labels match 39 FigureIQ fields. In FATURA, each party (
SELLER,BUYER,BILL TO,SEND TO) is one row. FATURA labels its name, address, email and phone separately. FigureIQ returns each address in 5 parts. - FigureIQ fields with no FATURA answer: 33, so they are not tested. 13 of them are line-item details, such as the description, quantity and price of each item.
- FATURA labels with no FigureIQ field: 16, such as bank details, tax IDs (GSTIN), websites and amount in words.
Of the 39 matched fields, each address counts as one check, with all 5 parts compared together as one block, and the tax lines count as one list. That gives the 26 checks used in the scoring (Figure 4).
5. How we decide right or wrong
Each value is marked either correct or incorrect. There is no partial credit: a value that is nearly right is marked incorrect.
- Amounts and rates must match to the cent.
4.7and4.70match;589.02and589.03do not. - Names, emails, invoice numbers and phone numbers: one wrong character makes the whole value wrong.
- Dates: the printed date must match as printed, and the calendar date must be the same day.
- Addresses are compared as one whole block. Only commas, semicolons and full stops are ignored.
- Currency: a printed
$and the codeUSDcount as the same.1 - Letter capitalization and extra spaces are ignored everywhere.
One printed line often holds several values. The answer key splits it using the invoice's own punctuation:
Figure 6 - One printed tax line is split into four values: tax name, tax rate, tax amount and currency.
The invoice prints only $, so the answer key keeps $. We accept FigureIQ's USD as a match.
Figure 7 - How an address is compared (illustrative). Letter capitalization, extra spaces, , ; and . are ignored; every other character must match.
For technical reviewers: comparison rule for each type of field
| Field type | How it is compared |
|---|---|
| Amounts | as numbers at two decimal places, no tolerance; a discount is a positive amount on both sides |
| Percentages | as numbers at two decimal places |
| Printed dates | as text |
| Calendar dates | read as a date and compared as a date |
| Invoice and PO numbers | as text |
| Emails | as text |
| Phones | as printed, ignoring spaces; brackets and hyphens must match |
| Names | as text |
| Addresses | parts joined, then compared with the printed block, ignoring , ; . and nothing else |
Text clean-up procedure: Before comparing text, we tidy up the answer and FigureIQ's value the same way: letter capitalization is ignored, extra spaces count as one.
6. Results
We scored all 1,250 test invoices on the final result you receive - after AI Extraction, the AI Judge and the postprocessing rules. Of the 17,499 values checked, 17,343 were exactly right and 156 were wrong.
| Measure | Values correct |
|---|---|
| Exact match accuracy | 99.1% |
The result splits clearly into two groups of fields:
| Group of fields | Values checked | Wrong | Exact match accuracy |
|---|---|---|---|
| Accounting fields - amounts, tax, currency, dates, invoice and PO numbers | 10,474 | 9 | 99.91% |
| Contact details - names, addresses, emails, phone numbers | 7,025 | 147 | 97.91% |
| All fields | 17,499 | 156 | 99.1% |
7. Where the mistakes are
The amounts were the most reliable part. Out of 10,474 accounting values, such as amounts, tax, currency, dates, and invoice and PO numbers, only 9 were wrong. Most mistakes were in contact details, mainly addresses and emails.
Figure 8 - Mistakes per 100 values, by type of field. For example, about 4 in every 100 addresses had a mistake.
Every subtotal, tax amount, tax rate, discount, currency and date was correct on every invoice. Only 4 of 1,050 invoice totals were wrong.
Figure 9 - Mistakes per 100 values, by field, worst first.
What a typical mistake looks like
Most wrong values are off by just one character. Here are five real examples, each checked against the invoice image. Everything else in the value was right, but the whole value still counts as wrong:
| Field | Invoice prints | FigureIQ returned | Difference |
|---|---|---|---|
| Invoice number | 4Y3M1d-603 | 4Y3Mld-603 | digit 1 read as letter l |
| Customer email | billyhaley@example.net | billyharley@example.net | extra r |
| Customer phone | +(479)757-7330 | +(479)757-7333 | last digit 0 read as 3 |
| Customer name | Shawn White | Shaun White | w read as u |
| Ship-to address | … WY 75625 US | … WY 75624 US | last postcode digit 5 read as 4 |
55 of the 89 address mistakes (62%) are only USA written where the
invoice says US. The street, city, state and postcode were all correct. This is a
formatting issue, not a reading problem. Fixing it would raise exact match accuracy to 99.42%.
8. What the AI Judge adds
We also scored the result after each step, to see how many mistakes each step removes.
| Step | Mistakes left, of 17,499 | Exact match accuracy |
|---|---|---|
| AI Extraction | 207 | 98.82% |
| + AI Judge - what you receive | 156 | 99.1% |
The table shows how many mistakes were left after each step. The first AI extraction left 207. The AI Judge brought that down to 156.
9. Check it yourself
Download FATURA
It is free to download from Zenodo (opens in a new tab), with the invoice images, the answer files and the test list
strat1_test.csv.Take the 1,250 invoices in the test list
You do not need any list from us. Each row names an image, such as
Template10_Instance175.jpg.Run them through FigureIQ
Use the app and save the results. Compare amounts and rates as numbers, not as text.
Build the answer key and compare
Follow the rules in sections 4 and 5. Accuracy is the number of correct values divided by all values checked. A value left blank when the invoice prints one counts as wrong.
Answer-key rules and common pitfalls
- Split each
TAXanswer into name, rate, amount and currency (Figure 6). - Split each
DISCOUNTinto rate and amount, and store the amount as a positive number. If you store-12.33, almost every discount will be marked wrong. - Each
DATEgives two values: the date as printed and the calendar date. - Count
$andUSDas the same when comparing currency. - Keep each address as one whole block. Comparing street, city and so on separately marks almost every address wrong.
- Compare amounts as numbers, not text. Otherwise
4.7and4.70are marked different.
- Split each
Compare your numbers with ours
Checkpoint (final result) Expected Invoices scored 1,250 Values checked 17,499 Fields checked (including the tax-line list) 26 Exact match accuracy 99.1%
10. Summary
- The result: FigureIQ got 99.1% of 17,499 values exactly right on 1,250 invoices. The dataset's creators chose these test invoices, not us.
- The answers were mapped in advance: 25 FATURA labels were mapped to 39 FigureIQ fields, giving 26 checks. We fixed this mapping before seeing any results and never changed FATURA's answers.
- The scoring was strict: a value counted only if it matched exactly. One wrong character or one cent off made it wrong.
- Amounts were the most reliable part: only 9 of 10,474 accounting values were wrong. Every subtotal, tax amount, tax rate, discount, currency and date was correct. 4 of 1,050 totals were wrong.
- Most mistakes were in contact details: mainly addresses and emails, and usually off
by one small detail. More than half of the address mistakes were just
USAwritten where the invoice saysUS. - The AI Judge reduced the number of errors from 207 to 156.
- Anyone can check this: the data is public, and section 9 explains how to repeat the test.
1 FigureIQ returns currency as a code, so a printed $ comes back as USD. We count these as the same. This is the only exception in the scoring: for example, US returned as USA in an address still counts as a mistake.
Tested in October 2026 on the 1,250 invoices in FATURA's official test set, covering 50 invoice layouts. The 99.1% figure is for the full FigureIQ process (AI Extraction, AI Judge and postprocessing rules), using the methodology defined in section 5.