System 03

Tournament Document Intelligence

Messy documents in, validated records out. 412 records merged from HTML, text PDFs and scans with 0 validation errors in the audited production run.

412records, 0 validation errors
Hover a step for what it does

Text version: Upload → Extract → Resolve → Validate → Merge

Problem

Tournament data arrived as HTML pages, text PDFs and scanned PDFs with inconsistent names and formats.

What I built

Multi-format intake (vision model for scans) feeding an 8-tier fuzzy entity resolver built for OCR noise, with a hard validation gate before anything reaches the merged dataset.

Key engineering decisions

  • Deterministic validation after AI extraction: the model proposes, rules decide.
  • An 8-tier resolver escalates from exact match to fuzzy strategies instead of one similarity threshold.
  • Nothing is written to the dataset until it passes the gate; failures are queued with the reason.

Validation and QA

Audited production run: 412 records, 286 fights, 89 categories, 0 validation errors.

Result

Clean, correct data from documents that previously required manual retyping. 412 records, 0 errors.