IP Library › Granted Patent US 12,602,927
Granted Patent B2
US 12,602,927 · App. 18/421,772 · Granted Apr 14, 2026

Pre-processing image frames based on camera statistics

Inventors: Naveen Thumpudi (Redmond, WA); Louis-Philippe Bourret (Redmond, WA); Christian Palmer Larson (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06V20/40G06N20/00G06V20/41G06V20/46H04N23/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,927
App. No.
18/421,772
Granted
Apr 14, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and computer-readable media for selectively identifying image frames from an input video to provide to an image processing model based on camera statistics. For example, systems disclosed herein include receiving an input video and associated camera statistics from a video capturing device. The systems disclosed herein further include identifying select image frames to provide to the image processing model based on the camera statistics and based on an application of the image processing model. The systems disclosed herein further include selectively identifying and providing camera statistics to the image processing model. By selectively providing data to the image processing model based on camera statistics, the systems disclosed herein can leverage capabilities of video capturing devices to significantly reduce the expense of processing resources when utilizing a variety of image processing models.

Claims (43)

1 . A method, comprising:

receiving, from one or more video capturing devices, input video content and a set of camera statistics, wherein the set of camera statistics includes data obtained by the one or more video capturing devices in conjunction with generating the input video content;

identifying an application for an image processing model to perform with respect to the input video content;

based on the identified application for the image processing model:

pre-processing the input video content using the set of camera statistics to generate transformed video content; and

identifying a subset of camera statistics from the set of camera statistics that are relevant to the identified application for the image processing model; and

providing a first input and a second input to the image processing model, the first input including the transformed video content and the second input including the subset of camera statistics, wherein the image processing model comprises a deep learning model trained to generate an output based on video content input and associated camera statistics inputs.

2 . The method of claim 1 , wherein the deep learning model is trained based on training data including both training video content and associated camera statistics.

3 . The method of claim 1 , wherein pre-processing the input video content includes one or more of refining image frames, modifying color or brightness, or down-sampling resolution.

4 . The method of claim 1 , wherein pre-processing the input video content includes identifying a subset of the input video content based on content of interest included in the subset of the input video content.

5 . The method of claim 4 , wherein identifying the subset of the input video content includes identifying cropped portions of image frames of the input video content, and wherein providing the first input and the second input to the image processing model comprises providing the cropped portions of the image frames and the subset of camera statistics as inputs to the deep learning model.

6 . The method of claim 4 , wherein identifying the subset of the input video content includes identifying a subset of image frames from a plurality of image frames within which the content of interest appears, and wherein providing the first input and the second input to the image processing model comprises providing the subset of image frames and the subset of camera statistics as inputs to the deep learning model.

7 . The method of claim 4 , wherein identifying the subset of the input video content is based on the subset of camera statistics indicating the content of interest included in the subset of the input video content.

8 . The method of claim 1 , wherein the input video content comprises captured video footage that has been locally refined by the one or more video capturing devices based on the set of camera statistics.

9 . The method of claim 1 , wherein the deep learning model is implemented on one or more of a cloud computing system or a computing device that receives the input video content from the one or more video capturing devices.

10 . A system, comprising:

one or more processors;

memory in electronic communication with the at one or more processors; and

instructions stored in the memory, the instructions being executable by the one or more processors to:

receive, from one or more video capturing devices, input video content and a set of camera statistics, wherein the set of camera statistics includes data obtained by the one or more video capturing devices in conjunction with generating the input video content;

identify an application for an image processing model to perform with respect to the input video content;

based on the identified application for the image processing model:

pre-processing the input video content using the set of camera statistics to generate transformed video content; and

identify a subset of camera statistics from the set of camera statistics that are relevant to the identified application for an image processing model; and

provide a first input and a second input to the image processing model, the first input including the transformed video content and the second input including the subset of camera statistics, wherein the image processing model comprises a deep learning model trained to generate an output based on video content input and associated camera statistics.

11 . The system of claim 10 , wherein the deep learning model is trained based on training data including both training video content and associated camera statistics.

12 . The system of claim 10 , wherein pre-processing the input video content includes one or more of refining image frames, modifying color or brightness, or down-sampling resolution.

13 . The system of claim 10 , wherein pre-processing the input video content includes identifying a subset of the input video content based on content of interest included in the subset of the input video content.

14 . The system of claim 13 , wherein identifying the subset of the input video content includes identifying cropped portions of image frames of the input video content, and wherein providing the first input and the second input to the image processing model comprises providing the cropped portions of the image frames and the subset of camera statistics as inputs to the deep learning model.

15 . The system of claim 13 , wherein identifying the subset of the input video content includes identifying a subset of image frames from a plurality of image frames within which the content of interest appears, and wherein providing the first input and the second input to the image processing model comprises providing the subset of image frames and the subset of camera statistics as inputs to the deep learning model.

16 . The system of claim 13 , wherein identifying the subset of the input video content is based on the subset of camera statistics indicating the content of interest included in the subset of the input video content.

17 . The system of claim 10 , wherein the input video content comprises captured video footage that has been locally refined by the one or more video capturing devices based on the set of camera statistics.

18 . The system of claim 10 , wherein the deep learning model is implemented on one or more of a cloud computing system or a computing device that receives the input video content from the one or more video capturing devices.

19 . A method, comprising:

receiving, from a plurality of video capturing devices, input video content and a set of camera statistics, wherein the set of camera statistics includes data obtained by the plurality of video capturing devices in conjunction with generating a plurality of input video streams;

identifying an application for an image processing model to perform with respect to the input video content;

based on the identified application for the image processing model:

pre-processing the input video content using the set of camera statistics to generate transformed video content including transforming the plurality of input video streams; and

identifying a subset of camera statistics from the set of camera statistics that are relevant to the identified application for the image processing model; and

providing a first input and a second input to the image processing model, the first input including the transformed video content and the second input including the subset of camera statistics, wherein the image processing model comprises a deep learning model trained to generate an output based on video content input and associated camera statistics inputs, and wherein providing the first input and the second input involves providing at least some transformed input video streams from each of the plurality of video capturing devices to generate the output.

20 . The method of claim 19 , wherein pre-processing the input video content includes identifying a subset of input video content based on content of interest included in the subset of input video content, and wherein identifying the subset of input video content includes one or more of:

identifying cropped portions of image frames of the input video content; or

identifying a subset of image frames from a plurality of image frames within which the content of interest appears.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2024
From: THUMPUDI, NAVEEN; BOURRET, LOUIS-PHILIPPE; LARSON, CHRISTIAN PALMER
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 066245/0962 →
Continuity (2)
Division 16298841 · Mar 11, 2019
Related Publication 20240320971A1 · Sep 26, 2024
References Cited (32)
US 10691925B2 · Ng · 2020 [cited by examiner]
US 20060109902A1 · Yu · 2006 [cited by applicant]
US 20060136575A1 · Payne · 2006 [cited by examiner]
US 20100118114A1 · Hosseini · 2010 [cited by examiner]
US 20160148650A1 · Laksono · 2016 [cited by applicant]
US 20180061459A1 · Song · 2018 [cited by examiner]
US 20180157939A1 · Butt · 2018 [cited by examiner]
US 20190005653A1 · Choi · 2019 [cited by examiner]
US 20190035047A1 · Lim · 2019 [cited by examiner]
US 20190042900A1 · Smith et al. · 2019 [cited by applicant]
US 20190156124A1 · Singhal · 2019 [cited by examiner]
US 20190213474A1 · Lin · 2019 [cited by examiner]
US 20190220698A1 · Pradeep · 2019 [cited by examiner]
US 20200293782A1 · Thumpudi · 2020 [cited by examiner]
US 20210176405A1 · Ishii · 2021 [cited by examiner]
CN 101223786A · 2008 [cited by applicant]
CN 101663676A · 2010 [cited by applicant]
JP 2006352879A · 2006 [cited by applicant]
WO 2018230294A1 · 2018 [cited by applicant]
Intimation of Grant received for IN Application No. 202147037878, mailed on Jul. 7, 2024, 1 page. [cited by applicant]
Invitation To Pay Additional Fees received for PCT Application No. PCT/US20/020574, mailed on Jun. 26, 2020, 34 Pages. [cited by applicant]
Invitation to Pay Additional Fees received for PCT Application No. PCT/US20/020575, mailed on May 12, 2020, 29 pages. [cited by applicant]
Second Office Action Received for Chinese Application No. 202080020602.7, mailed on Jan. 13, 2025, 07 pages. (English Translation Provided). [cited by applicant]
Notice of Allowance mailed on Sep. 28, 2023, in US Application No. 16/298, 841, 10 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC Received for European Application No. 20714432.0, mailed on Jul. 22, 2024, 09 pages. [cited by applicant]
Office Action Received for Chinese Application No. 202080020602.7, mailed on Jun. 1, 2024, 25 pages. (English Translation Provided). [cited by applicant]
Office Action Received for India Application No. 202147037878, mailed on Apr. 25, 2024, 03 pages. [cited by applicant]
Communication pursuant to Article 94(3) received in European Application No. 20715241.4, mailed on Jul. 15, 2025, 04 pages. [cited by applicant]
Huang, et al., “A Novel Key-Frames Selection Framework for Comprehensive Video Summarization”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 2, Feb. 2020, pp. 577-589. [cited by applicant]
Non-Final Office Action mailed on Jul. 23, 2025, in U.S. Appl. No. 18/070,236, 19 pages. [cited by applicant]
Notice of Allowance Received for Chinese Application No. 202080020602.7, mailed on Jun. 9, 2025, 8 pages. (English Translation Provided). [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 11 pages. [cited by applicant]