IP Library Granted Patent US 12,412,392
Granted Patent B2
US 12,412,392 · App. 18/050,331 · Granted Sep 9, 2025

Sports neural network codec

Inventors: Valerio Colamatteo (Milan, IT); Christopher Evi-Parker (London, GB); Sateesh Padagadi (London, GB); Patrick Joseph Lucey (Chicago, IL)
Assignee: STATS LLC
G06V20/42G06V10/778
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,392
App. No.
18/050,331
Granted
Sep 9, 2025
Kind
B2
Abstract

A computing system receives a broadcast video stream of a game. A codec module of the computing system extracts image level features from the broadcast video stream. The codec module includes an object detection portion configured to detect players in the broadcast video stream and a subnet portion attached to the object detection portion. The subnet portion is configured to identify foreground information of the detected players. The codec module provides the image level features to a plurality of task specific modules for analysis. The plurality of task specific modules generates a plurality of outputs based on the image level features.

Claims (52)

1. A method comprising:

receiving, by a computing system, a broadcast video stream of a game;

extracting, via a codec module of the computing system, image level features from the broadcast video stream, the codec module comprising an object detection portion that detects players in the broadcast video stream and a subnet portion attached to and downstream of the object detection portion, wherein the subnet portion identifies foreground information of the detected players;

providing, by the codec module, the image level features to a plurality of task specific modules for analysis; and

generating, by the plurality of task specific modules, a plurality of outputs based on the image level features.

2. The method of claim 1 , wherein the object detection portion comprises:

a backbone that extracts image level features from the broadcast video stream;

a neck downstream of the backbone, wherein the neck aggregates the extracted image level features; and

a head downstream of the neck, wherein the head identifies locations of players in the broadcast video stream based on the extracted image level features.

3. The method of claim 2 , wherein the head comprises a plurality of convolutions, wherein each convolution identifies a location of a player at varying resolutions.

4. The method of claim 3 , wherein the codec module further comprises:

a non-maximum suppression function downstream of the head, wherein the non-maximum suppression function combines the identified locations of the player at varying resolutions to generate a single location for the player.

5. The method of claim 2 , wherein the subnet portion is attached to the neck.

6. The method of claim 2 , wherein the subnet portion receives input from the neck, wherein the input from the neck is output generated by the neck, the output comprising floating point values indicating a likely position of players in the broadcast video stream.

7. The method of claim 1 , further comprising:

training, by the computing system, the object detection portion independent of the subnet portion; and

after training the object detection portion, training, by the computing system, the object detection portion with the subnet portion attached thereto.

8. A non-transitory computer readable medium comprising one or more sequences of instructions, which, when executed by a processor, causes a computing system to perform operations comprising:

receiving, by the computing system, a broadcast video stream of a game;

extracting, via a codec module of the computing system, image level features from the broadcast video stream, the codec module comprising an object detection portion that detects players in the broadcast video stream and a subnet portion attached to and downstream of the object detection portion, wherein the subnet portion identifies foreground information of the detected players;

providing, by the codec module, the image level features to a plurality of task specific modules for analysis; and

generating, by the plurality of task specific modules, a plurality of outputs based on the image level features.

9. The non-transitory computer readable medium of claim 8 , wherein the object detection portion comprises:

a backbone that extracts image level features from the broadcast video stream;

a neck downstream of the backbone, wherein the neck aggregates the extracted image level features; and

a head downstream of the neck, wherein the head identifies locations of players in the broadcast video stream based on the extracted image level features.

10. The non-transitory computer readable medium of claim 9 , wherein the head comprises a plurality of convolutions, wherein each convolution identifies a location of a player at varying resolutions.

11. The non-transitory computer readable medium of claim 10 , wherein the codec module further comprises:

a non-maximum suppression function downstream of the head, wherein the non-maximum suppression function combines the identified locations of the player at varying resolutions to generate a single location for the player.

12. The non-transitory computer readable medium of claim 9 , wherein the subnet portion is attached to the neck.

13. The non-transitory computer readable medium of claim 9 , wherein the subnet portion receives input from the neck, wherein the input from the neck is output generated by the neck, the output comprising floating point values indicating a likely position of players in the broadcast video stream.

14. The non-transitory computer readable medium of claim 8 , further comprising:

training, by the computing system, the object detection portion independent of the subnet portion; and

after training the object detection portion, training, by the computing system, the object detection portion with the subnet portion attached thereto.

15. A system comprising:

a processor; and

a memory having programming instructions stored thereon, which, when executed by the processor, cause the system to perform operations comprising:

receiving a broadcast video stream of a game;

extracting, via a codec module, image level features from the broadcast video stream, the codec module comprising an object detection portion that detects players in the broadcast video stream and a subnet portion attached to and downstream of the object detection portion, wherein the subnet portion identifies foreground information of the detected players;

providing, by the codec module, the image level features to a plurality of task specific modules for analysis; and

generating, by the plurality of task specific modules, a plurality of outputs based on the image level features.

16. The system of claim 15 , wherein the object detection portion comprises:

a backbone that extracts image level features from the broadcast video stream;

a neck downstream of the backbone, wherein the neck aggregates the extracted image level features; and

a head downstream of the neck, wherein the head identifies locations of players in the broadcast video stream based on the extracted image level features.

17. The system of claim 16 , wherein the head comprises a plurality of convolutions, wherein each convolution identifies a location of a player at varying resolutions.

18. The system of claim 17 , wherein the codec module further comprises:

a non-maximum suppression function downstream of the head, wherein the non-maximum suppression function combines the identified locations of the player at varying resolutions to generate a single location for the player.

19. The system of claim 16 , wherein the subnet portion receives input from the neck, wherein the input from the neck is output generated by the neck, the output comprising floating point values indicating a likely position of players in the broadcast video stream.

20. The system of claim 15 , further comprising:

training the object detection portion independent of the subnet portion; and

after training the object detection portion, training the object detection portion with the subnet portion attached thereto.

Assignments (2)
SECURITY INTEREST Recorded Apr 14, 2026
From: STATS LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 075390/0491 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: COLAMATTEO, VALERIO; EVI-PARKER, CHRISTOPHER; PADAGADI, SATEESH; LUCEY, PATRICK JOSEPH
To: STATS LLC
Reel/Frame 062135/0034 →
Continuity (2)
Provisional Application 63263189 · Oct 28, 2021
Related Publication 20230148112A1 · May 11, 2023
References Cited (21)
US 20080192116A1 · Tamir et al. · 2008 [cited by applicant]
US 20110013836A1 · Gefen et al. · 2011 [cited by applicant]
US 20110268320A1 · Huang et al. · 2011 [cited by applicant]
US 20160292510A1 · Han et al. · 2016 [cited by applicant]
US 20170344884A1 · Lin et al. · 2017 [cited by applicant]
US 20190221001A1 · Dassa · 2019 [cited by examiner]
US 20200234051A1 · Lee · 2020 [cited by examiner]
US 20200394413A1 · Bhanu et al. · 2020 [cited by applicant]
US 20220292311A1 · Song · 2022 [cited by examiner]
US 20240161461A1 · Zu · 2024 [cited by examiner]
CN 111951249A · 2020 [cited by examiner]
CN 112115914A · 2020 [cited by examiner]
WO WO2020237215A1 · 2020 [cited by examiner]
WO 2021016901 · 2021 [cited by applicant]
PCT International Application No. PCT/US22/78794, International Search Report and Written Opinion of the International Searching Authority, dated Feb. 2, 2023, 11 pages. [cited by applicant]
He Kaiming et al: “Mask R-CNN”, Jan. 24, 2018 (Jan. 24, 2018), XP055930853, Retrieved from the Internet:URL: https://arxiv.org/pdf/1703.06870.pdf [retrieved on Jun. 14, 2022]. [cited by applicant]
Jacek Komorowski et al: “FootAndBall: Integraed player and ball detector”, ARXIV.ORG, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 26, 2020 (Oct. 26, 2020), XP081799873. [cited by applicant]
Pobar Miran et al: “Mask R-CNN and Optical Flow Based Method for Detection and Marketing of Handball Actions”, 2018 11TH International Congress on Image and Signal Processing, Biomedical Engineering and Informatics (CIS… [cited by applicant]
Roman Voeikov et al: “TTNet: Real-time temporal and spatial video analysis of table tennis”, ARXIV.ORG, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 21, 2020 (Apr. 21, 2020), XP… [cited by applicant]
Sah Melike et al: “Review and evaluation of player detection methods in field sports: Comparing conventional and deep learning based methods”, Multimedia Tools and Applications., [Online] vol. 82, No. 9, Jun. 3, 2021 (J… [cited by applicant]
Zhang Yiqing et al: “Mask-Refined R-CNN: A Network for Refining Object Details in Instancce Segmentation”, Sensors, vol. 20, No. 4, Feb. 13, 2020 (Feb. 13, 2020), p. 1010, XP093266346. [cited by applicant]