IP Library Patent Application 18425760
Patent Application
App. No. 18/425,760

END-TO-END CAMERA CALIBRATION FOR BROADCAST VIDEO

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/425,760
Abstract

A system and method of calibrating a broadcast video feed are disclosed herein. A computing system retrieves a plurality of broadcast video feeds that include a plurality of video frames. The computing system generates a trained neural network, by generating a plurality of training data sets based on the broadcast video feed and learning, by the neural network, to generate a homography matrix for each frame of the plurality of frames. The computing system receives a target broadcast video feed for a target sporting event. The computing system partitions the target broadcast video feed into a plurality of target frames. The computing system generates for each target frame in the plurality of target frames, via the neural network, a target homography matrix. The computing system calibrates the target broadcast video feed by warping each target frame by a respective target homography matrix.

Claims (44)

1 - 20 . (canceled)

21 . A method of generating a fully trained calibration model, comprising:

retrieving, by a computing system, one or more training data sets from one or more data stores, wherein the one or more training data sets include a plurality of images captured by a camera system during a sporting event;

generating, by the computing system, a plurality of camera pose templates from the one or more training data sets;

training, by the computing system, a neural network to calibrate a camera based on the one or more training data sets and the plurality of camera pose templates; and

outputting, by the computing system, a trained prediction model based on the trained neural network.

22 . The method of claim 21 , wherein the plurality of camera pose templates are generated based on a high grid resolution approach.

23 . The method of claim 22 , wherein the high grid resolution approach comprises setting a pan resolution, a tilt resolution, and a focal length resolution.

24 . The method of claim 22 , wherein the neural network comprises:

a semantic segmentation module;

a camera pose initialization module; and

a homography refinement module.

25 . The method of claim 24 , wherein the training the neural network is performed module-by-module.

26 . The method of claim 24 , wherein the semantic segmentation module is configured to generate a semantic map.

27 . The method of claim 26 , wherein the camera pose initialization module is configured to determine a template of the plurality of camera pose templates for generating a homography matrix based on the semantic map.

28 . The method of claim 26 , wherein the homography refinement module is configured to generate a homography matrix based on a template of the plurality of camera pose templates and the semantic map.

29 . A system for generating a fully trained calibration model, comprising:

a processor; and

a memory having programming instructions stored thereon, which, when executed by the processor, performs one or more operations, comprising:

retrieving one or more training data sets from one or more data stores, wherein the one or more training data sets include a plurality of images captured by a camera system during a sporting event;

generating a plurality of camera pose templates from the one or more training data sets;

training a neural network to calibrate a camera based on the one or more training data sets and the plurality of camera pose templates; and

outputting a trained prediction model based on the trained neural network.

30 . The system of claim 29 , wherein the plurality of camera pose templates are generated based on a high grid resolution approach.

31 . The system of claim 30 , wherein the high grid resolution approach comprises setting a pan resolution, a tilt resolution, and a focal length resolution.

32 . The system of claim 30 , wherein the trained neural network comprises:

a semantic segmentation module;

a camera pose initialization module; and

a homography refinement module.

33 . The system of claim 32 , wherein the training the neural network is performed module-by-module.

34 . The system of claim 32 , wherein the semantic segmentation module is configured to generate a semantic map.

35 . The system of claim 34 , wherein the camera pose initialization module is configured to determine a template of the plurality of camera pose templates for generating a homography matrix based on the semantic map.

36 . The system of claim 34 , wherein the homography refinement module is configured to generate a homography matrix based on a template of the plurality of camera pose templates and the semantic map.

37 . A non-transitory computer readable medium including one or more sequences of instructions that, when executed by one or more processors, causes:

retrieving, by a computing system, one or more training data sets from one or more data stores, wherein the one or more training data sets include a plurality of images captured by a camera system during a sporting event;

generating, by the computing system, a plurality of camera pose templates from the one or more training data sets;

training, by the computing system, a neural network to calibrate a camera based on the one or more training data sets and the plurality of camera pose templates; and

outputting, by the computing system, a trained prediction model based on the trained neural network.

38 . The non-transitory computer readable medium of claim 37 , wherein the plurality of camera pose templates are generated based on a high grid resolution approach.

39 . The non-transitory computer readable medium of claim 38 , wherein the high grid resolution approach comprises setting a pan resolution, a tilt resolution, and a focal length resolution.

40 . The non-transitory computer readable medium of claim 38 , wherein the neural network comprises:

a semantic segmentation module;

a camera pose initialization module; and

a homography refinement module.

Assignments (2)
SECURITY INTEREST Recorded Apr 14, 2026
From: STATS LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 075390/0491 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2024
From: SHA, LONG; GANGULY, SUJOY; LUCEY, PATRICK JOSEPH
To: STATS LLC
Reel/Frame 066315/0623 →