IP Library › Granted Patent US 12,651,456
Granted Patent B1
US 12,651,456 · App. 19/290,059 · Granted Jun 9, 2026

System for generating a real-time object-focused video

Inventor: Michael Burnett (Barcelona, ES)
Assignee: MXV Inc.
G06V20/20G06T7/251G06T7/292G06T11/00G06T13/40G06T13/80G06V10/273G06V10/82G06T2207/10016G06T2207/30196G06T2207/30221G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,456
App. No.
19/290,059
Granted
Jun 9, 2026
Kind
B1
Abstract

A system for generating a real-time object-focused video using minimal camera arrays with pre-computed sports-field-optimized spatial mapping. The system positions virtual cameras to maintain tracked objects in focused, straight-ahead orientations while supporting one-dimensional movement between two cameras using geometric interpolation and two-dimensional movement within three-camera triangular configurations using barycentric coordinates. Computer spatial mapping with discretized depth information optimized for fast-moving object tracking in sports environments eliminates real-time depth calculation overhead, avoiding latency bottleneck and enabling ultra-low latency processing suitable for live sports broadcasting. The system includes predictive camera set switching using mathematical positioning variables, multi-object tracking capabilities with distributed processing frameworks, and intelligent 2D occlusion handling optimized for broadcast video output with parallax-induced occlusion management. Video synthesis techniques including adaptive geometric transformation, and multi-resolution image processing achieve rapid processing performance for live broadcasting applications while preventing discrete camera switching artifacts through continuous interpolation coefficients that eliminate abrupt perspective transitions.

Claims (48)

1 . A system for generating real-time object-focused video, comprising:

a processor; and

a non-transitory, processor-readable medium storing instructions that, when executed by the processor, cause the processor to:

identify at least one moving object within a coverage area using a convolutional neural network,

represent spatial data for positions within the coverage area as a plurality of depth vectors that defines a resolution that is based on the at least one moving object,

automatically determine virtual camera positions to track the at least one moving object from a desired orientation, based on a weighted combination of a plurality of cameras associated with the coverage area, and

generate virtual camera viewpoints based on the virtual camera positions by interpolating video feeds from the plurality of cameras based on the plurality of depth vectors.

2 . The system of claim 1 , wherein;

the plurality of cameras defines at least a first set of cameras comprising at least a first camera and a second set of cameras comprising at least a second camera, wherein the instructions to cause the processor to track the at least one moving object include instructions to cause the processor to track a moving object from the at least one moving object within (1) a coverage area defined by a first set of cameras from the plurality of cameras and (2) a coverage area defined by the second set of cameras from the plurality of cameras; and

the instructions to cause the processor to automatically determine the virtual camera positions include instructions to cause the processor to determine the virtual camera positions based on an alpha value associated with at least one of the coverage area defined by the first set of cameras or the coverage area defined by the second set of cameras.

3 . The system of claim 1 , wherein the instructions to cause the processor to determine the virtual camera positions include instructions to cause the processor to calculate a switching timing as a function of object velocity associated with the at least one moving object, a current position parameter, and a configurable boundary margin.

4 . The system of claim 1 , wherein;

the plurality of cameras comprises a first camera and a second camera; and

a virtual camera position from the virtual camera positions is (1) associated with a line defined by the first camera and the second camera and (2) defined by (Ox−C1x)/(C2x−C1x), where Ox is an object x-coordinate associated with the at least one moving object, C1x is a camera x-coordinate associated with the first camera, and C2x is a camera x-coordinate associated with the second camera.

5 . The system of claim 1 , wherein;

the at least one moving object includes a plurality of objects; and

the instructions to cause the processor to track the plurality of objects include instructions to cause the processor to concurrently track the plurality of objects by generating an independent virtual camera viewpoint for each tracked object from the plurality of objects.

6 . The system of claim 1 , wherein the instructions to cause the processor to track the at least one moving object include instructions to cause the processor to track the at least one moving object in configurable viewing orientations relative to a camera configuration geometry associated with the plurality of cameras.

7 . The system of claim 1 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to:

receive a dynamic angle adjustment parameter during operation, the virtual camera viewpoints depicting dolly-style lateral movement and angular perspective changes based on the dynamic angle adjustment parameter.

8 . The system of claim 1 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to correct a parallax error based on the video feeds, to generate the virtual camera viewpoints.

9 . The system of claim 1 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to smoothly vary continuous interpolation coefficients based on the at least one moving object, to reduce at least one of discrete camera switching artifacts or abrupt perspective transitions.

10 . The system of claim 1 , wherein the instructions to cause the processor to represent the spatial data as the plurality of depth vectors include instructions to cause the processor to generate the plurality of depth vectors with sub-50 ms processing latency, using sports-optimized spatial reference data via discretized lookup tables.

11 . The system of claim 1 , wherein the plurality of cameras comprises three or more cameras arranged to define a multi-dimensional area enabling two-dimensional virtual camera movement within an area defined by positions of the three or more cameras using barycentric coordinates.

12 . The system of claim 11 , wherein;

the three or more cameras are arranged in a triangular configuration; and

the instructions to cause the processor to cause the processor to determine the virtual camera position s include instructions to cause the processor to:

generate a plurality of camera weights associated with the weighted combination of the three or more cameras, based on the barycentric coordinates, and

determine the virtual camera positions based on the plurality of camera weights.

13 . The system of claim 11 , wherein;

the three or more cameras are arranged in a rectangular configuration; and

the instructions to cause the processor to determine the virtual camera position s include instructions to cause the processor to determine the virtual camera positions based on bilinear interpolation performed via parallel processing hardware.

14 . The system of claim 11 , wherein the instructions to cause the processor to generate the virtual camera viewpoints include instructions to cause the processor to apply, to the video feeds, a smoothing operation that balances predictive positioning with reactive adjustments, implementing human-like delay characteristics for abrupt object movements to reduce computational overhead while maintaining viewing comfort.

15 . A method, comprising:

receiving, via a processor and from a plurality of cameras, image data that depicts an object;

providing the image data as input to a convolutional neural network to identify the object within the image data;

generating, via the processor, a plurality of depth vectors that (1) represents a plurality of depths for a coverage area associated with the plurality of cameras and (2) defines a resolution that is based on motion of the object;

determining, via the processor, a plurality of virtual camera positions to track the object, based on the plurality of depth vectors and a plurality of weights associated with the plurality of cameras; and

generating, via the processor, video data that represents a plurality of virtual camera viewpoints, by interpolating, based on the plurality of virtual camera positions, the image data.

16 . The method of claim 15 , further comprising:

defining, via the processor, a grid having a plurality of cells, based on the coverage area associated with the plurality of cameras, each depth vector from the plurality of depth vectors being associated with a different cell from the plurality of cells.

17 . The method of claim 15 , wherein the resolution is a first resolution, the method further comprising:

defining, via the processor, a grid that represents a plurality of resolutions that includes the first resolution and a second resolution that is (1) different from the first resolution and (2) associated with an area of interest within the coverage area, the generating the plurality of depth vectors being based on the grid.

18 . The method of claim 15 , wherein the object includes at least one of a game player or a gameplay object.

19 . The method of claim 15 , wherein the generating the video data includes:

applying, via the processor, to the image data, and based on the plurality of depth vectors, at least one of a mesh warping operation, a dense correspondence mapping operation, an optical flow operation, or a temporal consistency filtering operation, to generate the video data.

20 . The method of claim 15 , wherein the generating the video data includes:

applying, via the processor, a mesh warping operation to the image data based on a mesh that is defined based on at least one of (1) a scene complexity associated with the coverage area or (2) a proximity of the object to at least one camera from the plurality of cameras.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2025
From: BURNETT, MICHAEL
To: MXV INC
Reel/Frame 071927/0918 →
Continuity (1)
Continuation 19264027 · Jul 9, 2025
References Cited (54)
US 8170277B2 · Michimoto et al. · 2012 [cited by applicant]
US 11983927B1 · Goyal et al. · 2024 [cited by applicant]
US 12033348B1 · Cao et al. · 2024 [cited by applicant]
US 12039781B2 · Goyal et al. · 2024 [cited by applicant]
US 12067755B1 · Nan et al. · 2024 [cited by applicant]
US 12254686B2 · Mwaura et al. · 2025 [cited by applicant]
US 12260491B2 · Fukuyasu · 2025 [cited by applicant]
US 20080192116A1 · Tamir et al. · 2008 [cited by applicant]
US 20080303901A1 · Variyath et al. · 2008 [cited by applicant]
US 20090041298A1 · Sandler et al. · 2009 [cited by applicant]
US 20110228092A1 · Park · 2011 [cited by applicant]
US 20120162436A1 · Cordell et al. · 2012 [cited by applicant]
US 20150063775A1 · Nakamura et al. · 2015 [cited by applicant]
US 20150147047A1 · Wang et al. · 2015 [cited by applicant]
US 20170094259A1 · Kouperman et al. · 2017 [cited by applicant]
US 20180167553A1 · Yee · 2018 [cited by examiner]
US 20180359427A1 · Choi · 2018 [cited by applicant]
US 20190083885A1 · Yee · 2019 [cited by examiner]
US 20190132529A1 · Ito · 2019 [cited by examiner]
US 20190174122A1 · Besley · 2019 [cited by examiner]
US 20190259199A1 · Yee · 2019 [cited by examiner]
US 20210241518A1 · Tong · 2021 [cited by examiner]
US 20220417441A1 · Voelker et al. · 2022 [cited by applicant]
US 20230186628A1 · Li · 2023 [cited by examiner]
CN 209181784U · 2019 [cited by applicant]
CN 112565630A · 2021 [cited by applicant]
CN 113114950A · 2021 [cited by applicant]
CN 114004773A · 2022 [cited by examiner]
CN 114401378A · 2022 [cited by applicant]
CN 115174850A · 2022 [cited by applicant]
DE 102005033853B3 · 2006 [cited by applicant]
EP 4109893A1 · 2022 [cited by applicant]
JP H06325180A · 1994 [cited by applicant]
JP 2011130323A · 2011 [cited by examiner]
JP 2016219968A · 2016 [cited by examiner]
JP 2017163245A · 2017 [cited by examiner]
RU 2706576C1 · 2019 [cited by applicant]
WO WO2019021375A1 · 2019 [cited by applicant]
WO WO2023196203A1 · 2023 [cited by applicant]
Adzemovic, M., “Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art,” arXiv:2506.13457v1 [cs.CV], (Jun. 16, 2025); 39 pages. [cited by applicant]
Audi, “Meet the Audi RS 3 Sedan,” Youtube.com, (Nov. 18, 2024) [online]. Retrieved on Oct. 10, 2025 from the Internet at URL: https://www.youtube.com/watch?v=YRR25-QgoQc&t=53s, PDF of Video Screenshot Provided; 6 pages. [cited by applicant]
Chen, J. et al., “A Robust Billboard-based Free-viewpoint Video Synthesizing Algorithm for Sports Scenes,” arXiv:1908.02446v1 [cs.MM], (Aug. 7, 2019); 10 pages. [cited by applicant]
Disney Research, “Algorithm combines videos from unstructured camera arrays into panoramas,” Phys.org, (May 4, 2015) [online]. Retrieved on Mar. 8, 2025 from the Internet at URL: https://phys.org/news/2015-05-algorithm-… [cited by applicant]
EP Application No. 25194315.5, Extended European Search Report mailed Oct. 22, 2025; Applicant MXV Inc.; 11 pages. [cited by applicant]
EP Application No. 25382556.6, Extended European Search Report mailed Nov. 13, 2025; Applicant MXV Inc.; 10 pages. [cited by applicant]
Formula One World Championship Limited, “Highlights: Watch the Action As Piastri Takes Chinese Grand Prix Victory in McLaren 1-2,” Formula1.com, (Mar. 23, 2025) [online]. Retrieved on Oct. 10, 2025 from the Internet at … [cited by applicant]
Lin, L. et al., “Line-preserving video stitching for asymmetric cameras,” Multimedia Tools and Applications, [Epub Nov. 12, 2018]; (Jun. 15, 2019), 78:14591-14611. [cited by applicant]
Olympics, “Alpine Skiing Beijing 2022 | Men's downhill highlights,” Youtube.com, (Feb. 7, 2022) [online]. Retrieved on Oct. 10, 2025 from the Internet at URL: https://www.youtube.com/watch?v=iZ-W2oEZfVg, PDF of Video Sc… [cited by applicant]
Park, K-W. et al., “Multi-Frame Based Homography Estimation for Video Stitching in Static Camera Environments,” Sensors, [Epub Dec. 22, 2019]; (Jan. 1, 2020), 20(1):92; 17 pages. [cited by applicant]
Red Bull, “World's Fastest Camera Drone Vs F1 Car (ft. Max Verstappen),” Youtube.com, (Feb. 27, 2024) [online]. Retrieved on Oct. 10, 2025 from the Internet at URL: https://www.youtube.com/watch?v=9pEqyr_uT-k&t=573s, PD… [cited by applicant]
Wikipedia, “Camera dolly,” Wikipedia.org, first publication date unknown [online], [last edited on Jan. 15, 2025, at 19:04 (UTC)]. Retrieved on Oct. 6, 2025 from the Internet at URL: https://en.wikipedia.org/wiki/Camera… [cited by applicant]
Wikipedia, “Cut (transition),” Wikipedia.org, first publication date unknown [online], [last edited on Aug. 4, 2025, at 03:14 (UTC)]. Retrieved on Oct. 6, 2025 from the Internet at URL: https://en.wikipedia.org/wiki/Cut… [cited by applicant]
Wikipedia, “Panning (camera),” Wikipedia.org, first publication date unknown [online], [last edited on Feb. 27, 2023, at 00:15 (UTC)]. Retrieved on Oct. 6, 2025 from the Internet at URL: https://en.wikipedia.org/wiki/Pa… [cited by applicant]
U.S. Appl. No. 19/289,932, Non-Final Office Action mailed Feb. 27, 2026, Inventor: Burnett, 12 pages. [cited by applicant]