Skip to content

03 — Extraction & Classification

Document Intelligence Pipeline

Turning PDFs, scans and photographs into structured, verified data at scale.

A document ingestion and extraction service. It classifies an incoming file, reads it with OCR and vision models, extracts structured fields against an enforced schema, and hands clean typed data to whatever needs it.

The hard part is never the happy path. It is the same document producing two different answers on two runs, the field that exists in one country and not another, and the cost of getting it wrong at volume.

What I built

  • 01Document classification and routing across invoices, POs, contracts and general files
  • 02Schema-enforced extraction — the response shape is guaranteed by the API, not requested in prose
  • 03Country-aware extraction modules so one pipeline handles multiple tax regimes
  • 04Prompt and context caching, with per-document cost tracking by pipeline phase
  • 05Deduplication, re-extraction and append-only revision history
  • 06Streaming progress to the browser over server-sent events

Outcomes

  • —Extraction made repeatable — same input, same output
  • —One pipeline extended to new countries without touching existing logic
  • —Per-document model cost made visible and then substantially reduced