Media & publishingNine fields per invoice in 1 day: how LabelFort annotated 1,489 invoice images
How LabelFort delivered value-only bounding box annotations across nine invoice fields on 1,489 images - tight rectangular boxes for OCR and document extraction model training.
- Finance & accounting
- Image · bounding box annotation

Contents
The client builds document intelligence workflows and OCR model training pipelines that depend on precise, value-only bounding boxes - not loose regions that include label words and add noise to extraction models. Their invoice automation stack requires tight rectangular annotations around each field value across highly variable templates.
LabelFort processed 1,489 invoice images across nine core invoice fields - 13,401 labeled fields - within 1 day, using a single-pass workflow with standardized value-only boxing rules across 22 annotators.
The challenge
Invoice images vary in layout, font, and field placement. Annotators must draw boxes around values only - excluding label words like “Invoice Date:” or “Tax ID:” - while covering multi-line addresses completely and skipping unreadable text to avoid noisy labels.
The client needed consistent boxing behavior at speed, with rules strict enough for distributed teams to maintain uniform quality.
- Tight boxes around target values only - label words excluded unless inseparable
- Multi-line values such as addresses captured in one complete bounding box
- Unreadable text skipped to prevent noisy training data
- Consistent field mapping across 22 annotators on varied templates
- Delivery within 1 day without sacrificing boxing precision
What we delivered
Each invoice image was annotated using bounding boxes for the following nine structured fields:
- Invoice Issue Date
- Invoice No
- Seller Name
- Seller Address
- Client Name
- Client Address
- Seller Tax ID
- Client Tax ID
- Total Amount
The annotation goal was a tightly fitted rectangular bounding box around each field value, producing model-ready training data for invoice extraction.
Our approach
LabelFort enforced a single-pass workflow with global quality rules: one tight box per field value, standardized field definitions, and consistent interpretation of multi-line content.
Step 1
Locate
Identify the target field value on the invoice.
Step 2
Box
Draw one tight rectangle around the value text only.
Step 3
Verify
Confirm complete coverage without label-word noise.
Step 4
Deliver
Nine fields per image, model-ready.
- One tight bounding box per field - label words excluded unless inseparable from the value.
- Multi-line addresses captured in a single box covering the full text block.
- Unreadable values skipped rather than guessed.
- Tight and complete coverage enforced - no cut-off characters or excessive whitespace.
- Standardized field definitions applied across the team to reduce label drift.
Results
The engagement delivered a high-precision invoice bounding box dataset on schedule - consistent value-only boxing applied across the full dataset.
One reporting note: the label total counts the nine-field framework applied across all 1,489 images (1,489 × 9 = 13,401 labeled fields). Values that were unreadable were skipped per guideline rather than boxed, so the delivered box count is the annotated subset of that total.
- 1,489
- invoice images annotated
- 13,401
- labeled fields covered
- 1 day
- end to end delivery
Nine fields per image
Value-only tight boxing
37 hours 39 minutes of effort
- 1,489 invoice images annotated
- 13,401 labeled fields covered across nine invoice fields per image
- 37 hours 39 minutes of total agent effort
- 22 annotators on a single-pass workflow
- Consistent value-only boxing rules applied across the full dataset
Key insights
- Value-only boxing improves dataset quality for OCR and extraction models by reducing label noise.
- Strict exclusion of label words increases generalizability across diverse invoice templates.
- Multi-line address boxing requires consistent interpretation to maintain uniform training signals.
- Clear global rules and field definitions enable distributed teams to maintain consistent quality at scale.
Impact
- OCR model training and evaluation for invoice field extraction
- Automated invoice processing workflows for finance and accounting operations
- Vendor and buyer identity resolution systems
- Tax ID and payment reconciliation use cases
Key takeaways
- Covered 13,401 labeled fields across 1,489 invoice images in 1 day with value-only boxing.
- Nine-field framework covered identity, address, tax ID, and total amount extraction.
- Tight boxing rules excluded label words and kept multi-line addresses in single complete boxes.
- Standardized field definitions held 22 annotators to consistent quality across varied templates.
More case studies
Media & publishing
E-commerce & consumer
Trust & safetyWant labels your auditors can read?
Start with a one hour Compliance Review on your own data.




