Problem
Tournament data arrived as HTML pages, text PDFs and scanned PDFs with inconsistent names and formats.
What I built
Multi-format intake (vision model for scans) feeding an 8-tier fuzzy entity resolver built for OCR noise, with a hard validation gate before anything reaches the merged dataset.
Key engineering decisions
- Deterministic validation after AI extraction: the model proposes, rules decide.
- An 8-tier resolver escalates from exact match to fuzzy strategies instead of one similarity threshold.
- Nothing is written to the dataset until it passes the gate; failures are queued with the reason.
Validation and QA
Audited production run: 412 records, 286 fights, 89 categories, 0 validation errors.
Result
Clean, correct data from documents that previously required manual retyping. 412 records, 0 errors.