IP Library Granted Patent US 12,579,602
Granted Patent B2
US 12,579,602 · App. 18/513,966 · Granted Mar 17, 2026

End-to-end camera calibration for broadcast video

Inventors: Long Sha (Brisbane, AU); Sujoy Ganguly (Chicago, IL); Patrick Joseph Lucey (Chicago, IL)
Assignee: Stats LLC
G06T3/00G06N3/08G06T3/18G06T7/11G06T7/80G06V10/24G06V10/82G06V20/49G06V30/19173G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,602
App. No.
18/513,966
Granted
Mar 17, 2026
Kind
B2
Abstract

A system and method of calibrating a broadcast video feed are disclosed herein. A computing system retrieves a plurality of broadcast video feeds that include a plurality of video frames. The computing system generates a trained neural network, by generating a plurality of training data sets based on the broadcast video feed and learning, by the neural network, to generate a homography matrix for each frame of the plurality of frames. The computing system receives a target broadcast video feed for a target sporting event. The computing system partitions the target broadcast video feed into a plurality of target frames. The computing system generates for each target frame in the plurality of target frames, via the neural network, a target homography matrix. The computing system calibrates the target broadcast video feed by warping each target frame by a respective target homography matrix.

Claims (46)

1 . A method of calibrating a broadcast video feed, comprising:

receiving, by a computing system, a broadcast video feed for a sporting event, the broadcast video feed including a plurality of frames;

inputting, by the computing system, the plurality of frames into a trained neural network, wherein the trained neural network is configured to generate a corresponding homography matrix for each of the plurality of frames;

calibrating, by the computing system, the broadcast video feed, wherein calibrating the broadcast video feed includes:

warping each of the plurality of frames based on the corresponding homography matrix to a high occupancy perspective and a low occupancy perspective, wherein warping each of the plurality of frames includes generating one or more images and one or more semantic labels; and

computing a loss function between the plurality of warped frames, wherein computing the loss function includes weighting the plurality of warped frames according to the high occupancy perspective and the low occupancy perspective; and

utilizing, by the computing system, the one or more images, the one or more semantic labels, and the loss function to further train the trained neural network.

2 . The method of claim 1 , wherein the trained neural network comprises:

a semantic segmentation module;

a camera pose initialization module; and

a homography refinement module.

3 . The method of claim 2 , wherein the semantic segmentation module is configured to generate a semantic map.

4 . The method of claim 3 , wherein the camera pose initialization module is configured to select a template from a set of templates using the semantic map to generate the homography matrix.

5 . The method of claim 4 , wherein the homography refinement module is configured to generate the homography matrix based on the template and the semantic map.

6 . The method of claim 4 , wherein the camera pose initialization module utilizes a Siamese network to select the template from the set of templates.

7 . The method of claim 1 , wherein the homography matrix registers a target ground-plane surface of any of the plurality of frames with a top view field model.

8 . The method of claim 1 , wherein the broadcast video includes an overhead view model, and wherein the overhead view model includes projected one or more three-dimensional locations of at least one of: one or more players or a ball onto a two-dimensional overhead view of a court of the sporting event.

9 . The method of claim 8 , wherein the warping each of the plurality of frames includes warping the overhead view model with the homography matrix.

10 . A system for calibrating a broadcast video feed, comprising:

a processor; and

a memory having programming instructions stored thereon, which, when executed by the processor, performs one or more operations, comprising:

receiving a broadcast video feed for a sporting event, the broadcast video feed including a plurality of frames;

inputting the plurality of frames into a trained neural network, wherein the trained neural network is configured to generate a corresponding homography matrix for each of the plurality of frames;

calibrating the broadcast video feed, wherein calibrating the broadcast video feed includes:

warping each of the plurality of frames based on the corresponding homography matrix to a high occupancy perspective and a low occupancy perspective, wherein warping each of the plurality of frames includes generating one or more images and one or more semantic labels; and

computing a loss function between the plurality of warped frames, wherein computing the loss function includes weighting the plurality of warped frames according to the high occupancy perspective and the low occupancy perspective; and

utilizing the one or more images, the one or more semantic labels, and the loss function to further train the trained neural network.

11 . The system of claim 10 , wherein the trained neural network comprises:

a semantic segmentation module;

a camera pose initialization module; and

a homography refinement module.

12 . The system of claim 11 , wherein the semantic segmentation module is configured to generate a semantic map.

13 . The system of claim 12 , wherein the camera pose initialization module is configured to select a template from a set of templates using the semantic map to generate the homography matrix.

14 . The system of claim 13 , wherein the camera pose initialization module utilizes a Siamese network to select the template from the set of templates.

15 . The system of claim 10 , wherein the homography matrix registers a target ground-plane surface of any of the plurality of frames with a top view field model.

16 . The system of claim 10 , wherein the broadcast video includes an overhead view model, and wherein the overhead view model includes projected one or more three-dimensional locations of at least one of: one or more players or a ball onto a two-dimensional overhead view of a court of the sporting event.

17 . The system of claim 16 , wherein the warping each of the plurality of frames includes warping the overhead view model with the homography matrix.

18 . A non-transitory computer readable medium including one or more sequences of instructions that, when executed by one or more processors, causes:

receiving, by a computing system, a broadcast video feed for a sporting event, the broadcast video feed including a plurality of frames;

inputting, by the computing system, the plurality of frames into a trained neural network, wherein the trained neural network is configured to generate a corresponding homography matrix for each of the plurality of frames;

calibrating, by the computing system, the broadcast video feed, wherein calibrating the broadcast video feed includes:

warping each of the plurality of frames based on the corresponding homography matrix to a high occupancy perspective and a low occupancy perspective, wherein warping each of the plurality of frames includes generating one or more images and one or more semantic labels; and

computing a loss function between the plurality of warped frames, wherein computing the loss function includes weighting the plurality of warped frames according to the high occupancy perspective and the low occupancy perspective; and

utilizing, by the computing system, the one or more images, the one or more semantic labels, and the loss function to further train the trained neural network.

19 . The non-transitory computer readable medium of claim 18 , wherein the homography matrix registers a target ground-plane surface of any of the plurality of frames with a top view field model.

20 . The non-transitory computer readable medium of claim 18 , wherein the broadcast video includes an overhead view model, and wherein the overhead view model includes projected one or more three-dimensional locations of at least one of: one or more players or a ball onto a two-dimensional overhead view of a court of the sporting event.

Assignments (2)
SECURITY INTEREST Recorded Apr 14, 2026
From: STATS LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 075390/0491 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2023
From: SHA, LONG; GANGULY, SUJOY; LUCEY, PATRICK JOSEPH
To: STATS LLC
Reel/Frame 065798/0783 →
Continuity (3)
Continuation 17226205 · Apr 9, 2021
Provisional Application 63008184 · Apr 10, 2020
Related Publication 20240095874A1 · Mar 21, 2024
References Cited (19)
US 20130063549A1 · Schnyder et al. · 2013 [cited by applicant]
US 20170255828A1 · Chang et al. · 2017 [cited by applicant]
US 20180336704A1 · Javan Roshtkhari · 2018 [cited by examiner]
US 20190034739A1 · Stein et al. · 2019 [cited by applicant]
US 20190087661A1 · Lee et al. · 2019 [cited by applicant]
US 20190147341A1 · Rabinovich et al. · 2019 [cited by applicant]
US 20200074682A1 · Sunkavalli et al. · 2020 [cited by applicant]
US 20200279398A1 · Sha · 2020 [cited by examiner]
US 20200372679A1 · Roshtkhari et al. · 2020 [cited by applicant]
CN 110782483A · 2020 [cited by applicant]
CN 110798673A · 2020 [cited by applicant]
WO 2010127418A1 · 2010 [cited by applicant]
WO 2020176872A1 · 2020 [cited by applicant]
Homayounfar, Namdar, Sanja Fidler, and Raquel Urtasun. “Soccer field localization from a single image.” arXiv preprint arXiv: 1604.02715 (2016). (Year: 2016). [cited by examiner]
Homayounfar, Namdar, Sanja Fidler, and Raquel Urtasun. “Sports field localization via deep structured models.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017. (Year: 2017). [cited by examiner]
Homayounfar Namdar et al: “Sports Field Localization via Deep Structured Models”,2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, US, Jul. 21, 2017, pp. 4012-4020, XP0332497… [cited by applicant]
PCT International Application No. PCT/US2021/026517, International Search Report and Written Opinion of the International Searching Authority, dated Jul. 16, 2021, 12 pages. [cited by applicant]
Chen, J. et al., “Sports Camera Calibration via Synthetic Data”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Jun. 16, 2019, pp. 2497-2504, XP033747032. [cited by applicant]
Detone, D. et al., “Deep Image Homography Estimation”, Jun. 1, 2016, total 5 pages, XP093261713, Retrieved from the Internet: URL:https://www.researchgate.net/publication/305881252_Deep_Image_Homography_Estimation. [cited by applicant]