Technology / Dataset
Building a Provenance-First Computer Vision Corpus
Film becomes training data only through a rigorous, auditable process.
SECTION
Provenance by Design
Every bounding box, track, label, and homography stores:
- Source master file ID
- Exact frame number
- Model/version that produced it
- Confidence score
- Schema version
Future models can be run retroactively over the entire historical corpus because raw evidence (tracks, homography, audio features) is stored at ingest.
SECTION
Rights Policy
Only first-party film (your own team's Hudl exports, self-filmed 4K sideline) or properly licensed commercial film is ingested. No NFHS streams, no scraped content. The pipeline is deliberately source-agnostic but the business rule is strict.
SECTION
What Is Measured So Far
- Restoration round trip, truth known: SSIM 0.8617 to 0.9708, PSNR 21.20 to 33.43 dB
- Effective-scale check flags upscaled masters: caught a 540p source sold as 1080p (SSIM 0.9976 against a 2x round trip, threshold 0.985)
- Helmet training corpus: 9,947 frames / 193,736 boxes, from the public-domain NFL helmet dataset (5 helmet classes, no player bodies or play labels)
- No temporal validation yet: detection is scored per frame, not across a play
The dataset grows with every game ingested, whether or not the film receives restoration or clipping. This ensures that when advanced route/coverage models arrive, they have the full historical base to train on.