Technology / Dataset

Building a Provenance-First Computer Vision Corpus

Film becomes training data only through a rigorous, auditable process.

SECTION

Provenance by Design

Every bounding box, track, label, and homography stores:

  • Source master file ID
  • Exact frame number
  • Model/version that produced it
  • Confidence score
  • Schema version

Future models can be run retroactively over the entire historical corpus because raw evidence (tracks, homography, audio features) is stored at ingest.

SECTION

Rights Policy

Only first-party film (your own team's Hudl exports, self-filmed 4K sideline) or properly licensed commercial film is ingested. No NFHS streams, no scraped content. The pipeline is deliberately source-agnostic but the business rule is strict.

SECTION

What Is Measured So Far

  • Restoration round trip, truth known: SSIM 0.8617 to 0.9708, PSNR 21.20 to 33.43 dB
  • Effective-scale check flags upscaled masters: caught a 540p source sold as 1080p (SSIM 0.9976 against a 2x round trip, threshold 0.985)
  • Helmet training corpus: 9,947 frames / 193,736 boxes, from the public-domain NFL helmet dataset (5 helmet classes, no player bodies or play labels)
  • No temporal validation yet: detection is scored per frame, not across a play

The dataset grows with every game ingested, whether or not the film receives restoration or clipping. This ensures that when advanced route/coverage models arrive, they have the full historical base to train on.