Nine fields per invoice in 1 day: how LabelFort annotated 1,489 invoice images

How LabelFort delivered value-only bounding box annotations across nine invoice fields on 1,489 images - tight rectangular boxes for OCR and document extraction model training.

  • Finance & accounting
  • Image · bounding box annotation
Invoice bounding box annotation case study
Contents

The client builds document intelligence workflows and OCR model training pipelines that depend on precise, value-only bounding boxes - not loose regions that include label words and add noise to extraction models. Their invoice automation stack requires tight rectangular annotations around each field value across highly variable templates.

LabelFort processed 1,489 invoice images across nine core invoice fields - 13,401 labeled fields - within 1 day, using a single-pass workflow with standardized value-only boxing rules across 22 annotators.

The challenge

Invoice images vary in layout, font, and field placement. Annotators must draw boxes around values only - excluding label words like “Invoice Date:” or “Tax ID:” - while covering multi-line addresses completely and skipping unreadable text to avoid noisy labels.

The client needed consistent boxing behavior at speed, with rules strict enough for distributed teams to maintain uniform quality.

  • Tight boxes around target values only - label words excluded unless inseparable
  • Multi-line values such as addresses captured in one complete bounding box
  • Unreadable text skipped to prevent noisy training data
  • Consistent field mapping across 22 annotators on varied templates
  • Delivery within 1 day without sacrificing boxing precision

What we delivered

Each invoice image was annotated using bounding boxes for the following nine structured fields:

  1. Invoice Issue Date
  2. Invoice No
  3. Seller Name
  4. Seller Address
  5. Client Name
  6. Client Address
  7. Seller Tax ID
  8. Client Tax ID
  9. Total Amount

The annotation goal was a tightly fitted rectangular bounding box around each field value, producing model-ready training data for invoice extraction.

Our approach

LabelFort enforced a single-pass workflow with global quality rules: one tight box per field value, standardized field definitions, and consistent interpretation of multi-line content.

  1. Step 1

    Locate

    Identify the target field value on the invoice.

  2. Step 2

    Box

    Draw one tight rectangle around the value text only.

  3. Step 3

    Verify

    Confirm complete coverage without label-word noise.

  4. Step 4

    Deliver

    Nine fields per image, model-ready.

  • One tight bounding box per field - label words excluded unless inseparable from the value.
  • Multi-line addresses captured in a single box covering the full text block.
  • Unreadable values skipped rather than guessed.
  • Tight and complete coverage enforced - no cut-off characters or excessive whitespace.
  • Standardized field definitions applied across the team to reduce label drift.

Results

The engagement delivered a high-precision invoice bounding box dataset on schedule - consistent value-only boxing applied across the full dataset.

One reporting note: the label total counts the nine-field framework applied across all 1,489 images (1,489 × 9 = 13,401 labeled fields). Values that were unreadable were skipped per guideline rather than boxed, so the delivered box count is the annotated subset of that total.

1,489
invoice images annotated

Nine fields per image

13,401
labeled fields covered

Value-only tight boxing

1 day
end to end delivery

37 hours 39 minutes of effort

  • 1,489 invoice images annotated
  • 13,401 labeled fields covered across nine invoice fields per image
  • 37 hours 39 minutes of total agent effort
  • 22 annotators on a single-pass workflow
  • Consistent value-only boxing rules applied across the full dataset

Key insights

  • Value-only boxing improves dataset quality for OCR and extraction models by reducing label noise.
  • Strict exclusion of label words increases generalizability across diverse invoice templates.
  • Multi-line address boxing requires consistent interpretation to maintain uniform training signals.
  • Clear global rules and field definitions enable distributed teams to maintain consistent quality at scale.

Impact

  • OCR model training and evaluation for invoice field extraction
  • Automated invoice processing workflows for finance and accounting operations
  • Vendor and buyer identity resolution systems
  • Tax ID and payment reconciliation use cases

Key takeaways

  • Covered 13,401 labeled fields across 1,489 invoice images in 1 day with value-only boxing.
  • Nine-field framework covered identity, address, tax ID, and total amount extraction.
  • Tight boxing rules excluded label words and kept multi-line addresses in single complete boxes.
  • Standardized field definitions held 22 annotators to consistent quality across varied templates.

More case studies

See all case studies

Want labels your auditors can read?

Start with a one hour Compliance Review on your own data.

Certifications & readiness

  • ISO 27001:2022 - CERTIFIED
  • SOC 2 - ALIGNED
  • HIPAA - COMPLIANT
  • GDPR - COMPLIANT
  • DPDP - READY