Benchmark report · FATURA dataset

99.1% accuracy on the FATURA benchmark, and where the other 0.9% went

Published October 8, 2026  ·  12 min read

We ran FigureIQ on 1,250 invoices from FATURA (paper (opens in a new tab) · dataset (opens in a new tab)), a public invoice benchmark, and checked 17,499 extracted values against the dataset's published answers. This page covers what we checked, how we scored it, and where the mistakes were.

At a glance
99.1%
of 17,499 values exactly right
1,250
test invoices
  • Exact match accuracy is the share of values that match the correct answer exactly. There is no partial credit: a name with one wrong letter, or an amount one cent off, counts as wrong.1
  • Most mistakes are in contact details - addresses and emails - not in amounts.
  • Anyone can repeat this test with public data (section 9).
Terms used on this page
Field
A type of information on an invoice, such as the invoice number or the total.
Value
One field on one invoice - for example, the total on one invoice.
Answer key
The correct value of every field, taken from the answers published with the dataset.
Exact match accuracy
Correct values ÷ all values checked. A value is correct only if it matches the answer key exactly. Letter capitalization and extra spaces are ignored, and $ and USD are treated as the same.1 Any other difference, even one character, makes the value wrong.
Invoice layout
One invoice design (template). FATURA has 50 layouts, with 200 invoices each.
Test set
The 1,250 invoices the dataset's creators set aside for testing.
AI Judge
A second AI pass that checks the first reading against the invoice image and corrects any errors that it discovered.
Postprocessing rules
Fixed formatting rules with no AI, such as turning (-) 9.39 into a discount of 9.39.

1. Why we publish this

Most vendors quote an accuracy number without disclosing how it was arrived at. “99% accurate” can mean 99% of characters, 99% of fields, or 99% of documents.

A number is only useful with three things attached: a public dataset, a test set the vendor did not choose, and a written scoring rule. This page gives all three.

2. The test invoices

We used FATURA (paper (opens in a new tab) · dataset (opens in a new tab)), a public research collection of 10,000 invoice images: 50 invoice layouts, 200 invoices each. The correct value of every field is published with each invoice. We do not decide what the right answer is - the dataset does.

FATURA's creators split the invoices into three groups: 7,500 for training, 1,250 for development and 1,250 for testing. We scored the whole test group and nothing else: the 1,250 invoices listed in the dataset's own file strat1_test.csv. The 99.1% comes from these 1,250 invoices only. FigureIQ was never trained or tuned on them, so it saw them for the first time in this test.

10,000 invoices: 50 layouts × 200 each Training: 7,500 invoices (75%) 7,500 75% Development: 1,250 invoices (12.5%) 1,250 12.5% Test: 1,250 invoices (12.5%) - the only group we score 1,250 12.5% Training Development Test The only group used for the 99.1%

Figure 1 - How FATURA's 10,000 invoices are divided. Every layout contributes 150 training, 25 development and 25 test invoices.

What the test invoices look like

The 50 layouts differ the way real suppliers' invoices do: the seller block sits at the top on some and at the bottom on others, some print one tax line and some several, and some show “Balance due” instead of “Total”.

Figure 2 - Eight test invoices, each from a different layout.

3. How FigureIQ reads an invoice

Each invoice goes through three steps, much like an accounts team where one person keys in the invoice, a second checks it against the paper, and house rules put everything in the same format:

  1. AI Extraction reads the invoice image and fills in the fields.
  2. AI Judge, a second AI pass, checks those fields against the image and corrects what looks wrong.
  3. Postprocessing rules (no AI) put values into a standard format - for example, a printed $ becomes the currency code USD, and (-) 9.39 becomes a discount of 9.39. The same input always gives the same output.
Invoice1AI ExtractionAI2AI JudgeAI3Postprocessing RulesNo AIYour result

Figure 3 - The three steps. You receive only the final result.

4. What we checked

Invoice details
Invoice numberPO numberInvoice dateInvoice date (calendar)Due dateDue date (calendar)
Amounts & tax
SubtotalDiscount rateDiscount amountTax nameTax rateTax amountTotalCurrencyTax lines (list)
Seller
NameAddressEmail
Customer
NameAddressEmailPhone
Ship-to
NameAddressEmailPhone
Not checked
Bank detailsTax IDs (GSTIN)WebsitesAmount in wordsLine items

Figure 4 - The 26 fields checked. Tax lines are checked only on invoices with more than one tax rate.

How we matched FATURA's answers to FigureIQ's fields

FATURA and FigureIQ name the same information differently. For example, what FATURA labels BUYER, FigureIQ calls Customer. Some FATURA labels also hold several values at once, such as a whole tax line. So before we can compare anything, we need a mapping: a list of which FATURA label is compared with which FigureIQ field.

FATURA answer labelFigureIQ field it checksNUMBER → Invoice number (one to one)PO NUMBER → PO number (one to one)SELLER → Seller: name, address, email (one to one)BUYER → Customer: name, address, email, phone (one to one)SEND TO → Ship-to: name, address, email, phone (one to one)DATE → Invoice date (one answer split into several values)DATE → Invoice date (calendar) (one answer split into several values)DUE DATE → Due date (one answer split into several values)DUE DATE → Due date (calendar) (one answer split into several values)TOTAL → Total (one answer split into several values)TOTAL → Currency (one answer split into several values)SUB TOTAL → Subtotal (one answer split into several values)SUB TOTAL → Currency (one answer split into several values)TAX → Tax name (one answer split into several values)TAX → Tax rate (one answer split into several values)TAX → Tax amount (one answer split into several values)TAX → Currency (one answer split into several values)GST(n%) → Tax name (one answer split into several values)GST(n%) → Tax rate (one answer split into several values)GST(n%) → Tax amount (one answer split into several values)DISCOUNT → Discount rate (one answer split into several values)DISCOUNT → Discount amount (one answer split into several values)AMOUNT DUE → Total (only in certain cases)AMOUNT DUE → Currency (only in certain cases)GST(n%) → Tax lines (list) (only in certain cases)BILL TO → Customer: name, address, email, phone (only in certain cases)NUMBERPO NUMBERDATEDUE DATETOTALAMOUNT DUESUB TOTALTAXGST(n%)DISCOUNTSELLERBUYERBILL TOSEND TOInvoice numberPO numberInvoice dateInvoice date (calendar)Due dateDue date (calendar)TotalSubtotalTax nameTax rateTax amountDiscount rateDiscount amountTax lines (list)CurrencySeller: name, address, emailCustomer: name, address, email, phoneShip-to: name, address, email, phone
  • One-to-one mapping. One FATURA label maps to one FigureIQ field. Example: NUMBER → Invoice number.
  • One-to-many mapping. One FATURA label maps to several FigureIQ fields. Example: TAX → tax name, tax rate, tax amount and currency (Figure 6).
  • Conditional mapping. Used only in some cases. BILL TO maps to the customer only when there is no BUYER. GST(n%) maps to the tax-line list only when several GST lines are printed.

Figure 5 - Mapping of FATURA labels (left) to FigureIQ fields (right). Currency has no FATURA label of its own; it is taken from the symbol printed with the amounts.

For technical reviewers: the same mapping as a table, and why it is 26 checks
FATURA labelFigureIQ fields it checks
NUMBERInvoice number
PO NUMBERPO number
DATEInvoice date, invoice date (calendar)
DUE DATEDue date, due date (calendar)
TOTALTotal, currency
AMOUNT DUETotal, currency

Some layouts print only a TOTAL, some only an AMOUNT DUE, and some print both. If a layout prints both, we count the total only once, so the same field isn't marked twice.
SUB TOTALSubtotal, currency
TAXTax name, tax rate, tax amount, currency
GST(n%)Tax name, tax rate, tax amount; as a list of tax lines when several are printed
DISCOUNTDiscount rate, discount amount
SELLERSeller name, address, email
BUYERCustomer name, address, email, phone
BILL TOCustomer name, address, email, phone - only when there is no BUYER
SEND TOShip-to name, address, email, phone

FigureIQ returns 72 fields from an invoice. FATURA's answers use 41 labels. This is how the two line up:

  • Matched: 25 FATURA labels match 39 FigureIQ fields. In FATURA, each party (SELLER, BUYER, BILL TO, SEND TO) is one row. FATURA labels its name, address, email and phone separately. FigureIQ returns each address in 5 parts.
  • FigureIQ fields with no FATURA answer: 33, so they are not tested. 13 of them are line-item details, such as the description, quantity and price of each item.
  • FATURA labels with no FigureIQ field: 16, such as bank details, tax IDs (GSTIN), websites and amount in words.

Of the 39 matched fields, each address counts as one check, with all 5 parts compared together as one block, and the tax lines count as one list. That gives the 26 checks used in the scoring (Figure 4).

5. How we decide right or wrong

Each value is marked either correct or incorrect. There is no partial credit: a value that is nearly right is marked incorrect.

  • Amounts and rates must match to the cent. 4.7 and 4.70 match; 589.02 and 589.03 do not.
  • Names, emails, invoice numbers and phone numbers: one wrong character makes the whole value wrong.
  • Dates: the printed date must match as printed, and the calendar date must be the same day.
  • Addresses are compared as one whole block. Only commas, semicolons and full stops are ignored.
  • Currency: a printed $ and the code USD count as the same.1
  • Letter capitalization and extra spaces are ignored everywhere.

One printed line often holds several values. The answer key splits it using the invoice's own punctuation:

Printed on the invoice VAT (3.85%): 10.26 $ Tax name “VAT” Tax rate “3.85” Tax amount “10.26” Currency “$”

Figure 6 - One printed tax line is split into four values: tax name, tax rate, tax amount and currency.

The invoice prints only $, so the answer key keeps $. We accept FigureIQ's USD as a match.

FigureIQ returns five parts street 882 Oak Lane city Tammyland state SD postcode 42587 country US joined into one line 882 oak lane tammyland sd 42587 us The invoice prints 882 Oak Lane Tammyland, SD 42587 US comma removed 882 oak lane tammyland sd 42587 us Same text → correct Hyphens, slashes, # and & still count: “lot 1851-a” is not the same as “lot 1851 a”

Figure 7 - How an address is compared (illustrative). Letter capitalization, extra spaces, , ; and . are ignored; every other character must match.

For technical reviewers: comparison rule for each type of field
Field typeHow it is compared
Amountsas numbers at two decimal places, no tolerance; a discount is a positive amount on both sides
Percentagesas numbers at two decimal places
Printed datesas text
Calendar datesread as a date and compared as a date
Invoice and PO numbersas text
Emailsas text
Phonesas printed, ignoring spaces; brackets and hyphens must match
Namesas text
Addressesparts joined, then compared with the printed block, ignoring , ; . and nothing else

Text clean-up procedure: Before comparing text, we tidy up the answer and FigureIQ's value the same way: letter capitalization is ignored, extra spaces count as one.

6. Results

We scored all 1,250 test invoices on the final result you receive - after AI Extraction, the AI Judge and the postprocessing rules. Of the 17,499 values checked, 17,343 were exactly right and 156 were wrong.

MeasureValues correct
Exact match accuracy99.1%

The result splits clearly into two groups of fields:

Group of fieldsValues checkedWrongExact match accuracy
Accounting fields - amounts, tax, currency, dates, invoice and PO numbers10,474999.91%
Contact details - names, addresses, emails, phone numbers7,02514797.91%
All fields17,49915699.1%

7. Where the mistakes are

The amounts were the most reliable part. Out of 10,474 accounting values, such as amounts, tax, currency, dates, and invoice and PO numbers, only 9 were wrong. Most mistakes were in contact details, mainly addresses and emails.

0246810Addresses89 wrong of 2,175Addresses: 89 wrong of 2,1754.09Emails39 wrong of 1,700Emails: 39 wrong of 1,7002.29Names14 wrong of 2,000Names: 14 wrong of 2,0000.70Phone numbers5 wrong of 1,150Phone numbers: 5 wrong of 1,1500.43Accounting fields9 wrong of 10,474Accounting fields: 9 wrong of 10,4740.09wrong values per 100 checked

Figure 8 - Mistakes per 100 values, by type of field. For example, about 4 in every 100 addresses had a mistake.

Every subtotal, tax amount, tax rate, discount, currency and date was correct on every invoice. Only 4 of 1,050 invoice totals were wrong.

Contact detailsAccounting fields0246810Ship-to addressShip-to address: 15 wrong of 2256.67Seller addressSeller address: 47 wrong of 1,0254.59Ship-to emailShip-to email: 7 wrong of 2253.11Customer addressCustomer address: 27 wrong of 9252.92Seller emailSeller email: 12 wrong of 5502.18Customer emailCustomer email: 20 wrong of 9252.16Ship-to phoneShip-to phone: 3 wrong of 2251.33Seller nameSeller name: 9 wrong of 8501.06Tax nameTax name: 4 wrong of 5500.73Ship-to nameShip-to name: 1 wrong of 2250.44Customer nameCustomer name: 4 wrong of 9250.43Invoice totalInvoice total: 4 wrong of 1,0500.38Customer phoneCustomer phone: 2 wrong of 9250.22Invoice numberInvoice number: 1 wrong of 1,1000.09Currencynone wrongDiscount amountnone wrongDiscount ratenone wrongDue datenone wrongInvoice datenone wrongPO numbernone wrongSubtotalnone wrongTax amountnone wrongTax ratenone wrongwrong values per 100 of that field

Figure 9 - Mistakes per 100 values, by field, worst first.

What a typical mistake looks like

Most wrong values are off by just one character. Here are five real examples, each checked against the invoice image. Everything else in the value was right, but the whole value still counts as wrong:

FieldInvoice printsFigureIQ returnedDifference
Invoice number4Y3M1d-6034Y3Mld-603digit 1 read as letter l
Customer emailbillyhaley@example.netbillyharley@example.netextra r
Customer phone+(479)757-7330+(479)757-7333last digit 0 read as 3
Customer nameShawn WhiteShaun Whitew read as u
Ship-to address… WY 75625 US… WY 75624 USlast postcode digit 5 read as 4
The biggest single fix

55 of the 89 address mistakes (62%) are only USA written where the invoice says US. The street, city, state and postcode were all correct. This is a formatting issue, not a reading problem. Fixing it would raise exact match accuracy to 99.42%.

8. What the AI Judge adds

We also scored the result after each step, to see how many mistakes each step removes.

StepMistakes left, of 17,499Exact match accuracy
AI Extraction20798.82%
+ AI Judge - what you receive15699.1%

The table shows how many mistakes were left after each step. The first AI extraction left 207. The AI Judge brought that down to 156.

9. Check it yourself

  1. Download FATURA

    It is free to download from Zenodo (opens in a new tab), with the invoice images, the answer files and the test list strat1_test.csv.

  2. Take the 1,250 invoices in the test list

    You do not need any list from us. Each row names an image, such as Template10_Instance175.jpg.

  3. Run them through FigureIQ

    Use the app and save the results. Compare amounts and rates as numbers, not as text.

  4. Build the answer key and compare

    Follow the rules in sections 4 and 5. Accuracy is the number of correct values divided by all values checked. A value left blank when the invoice prints one counts as wrong.

    Answer-key rules and common pitfalls
    • Split each TAX answer into name, rate, amount and currency (Figure 6).
    • Split each DISCOUNT into rate and amount, and store the amount as a positive number. If you store -12.33, almost every discount will be marked wrong.
    • Each DATE gives two values: the date as printed and the calendar date.
    • Count $ and USD as the same when comparing currency.
    • Keep each address as one whole block. Comparing street, city and so on separately marks almost every address wrong.
    • Compare amounts as numbers, not text. Otherwise 4.7 and 4.70 are marked different.
  5. Compare your numbers with ours

    Checkpoint (final result)Expected
    Invoices scored1,250
    Values checked17,499
    Fields checked (including the tax-line list)26
    Exact match accuracy99.1%

10. Summary

  • The result: FigureIQ got 99.1% of 17,499 values exactly right on 1,250 invoices. The dataset's creators chose these test invoices, not us.
  • The answers were mapped in advance: 25 FATURA labels were mapped to 39 FigureIQ fields, giving 26 checks. We fixed this mapping before seeing any results and never changed FATURA's answers.
  • The scoring was strict: a value counted only if it matched exactly. One wrong character or one cent off made it wrong.
  • Amounts were the most reliable part: only 9 of 10,474 accounting values were wrong. Every subtotal, tax amount, tax rate, discount, currency and date was correct. 4 of 1,050 totals were wrong.
  • Most mistakes were in contact details: mainly addresses and emails, and usually off by one small detail. More than half of the address mistakes were just USA written where the invoice says US.
  • The AI Judge reduced the number of errors from 207 to 156.
  • Anyone can check this: the data is public, and section 9 explains how to repeat the test.

1 FigureIQ returns currency as a code, so a printed $ comes back as USD. We count these as the same. This is the only exception in the scoring: for example, US returned as USA in an address still counts as a mistake.

Tested in October 2026 on the 1,250 invoices in FATURA's official test set, covering 50 invoice layouts. The 99.1% figure is for the full FigureIQ process (AI Extraction, AI Judge and postprocessing rules), using the methodology defined in section 5.