IP Library › Granted Patent US 11,902,532
Granted Patent B2
US 11,902,532 · App. 17/488,944 · Granted Feb 13, 2024

Video encoding optimization for machine learning content categorization

Inventors: Sunil Gopal Koteyar (Markham, CA); Mingkai Shao (Richmond Hill, CA)
Assignee: ATI Technologies ULC
H04N19/139G06N20/00G06T3/40G06V20/49H04N19/126H04N19/142H04N19/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,902,532
App. No.
17/488,944
Granted
Feb 13, 2024
Kind
B2
Abstract

Systems, apparatuses, and methods for performing machine learning content categorization leveraging video encoding pre-processing are disclosed. A system includes at least a motion vector unit and a machine learning (ML) engine. The motion vector unit pre-processes a frame to determine if there is temporal locality with previous frames. If the objects of the scene have not changed by a threshold amount, then the ML engine does not process the frame, saving computational resources that would typically be used. Otherwise, if there is a change of scene or other significant changes, then the ML engine is activated to process the frame. The ML engine can then generate a QP map and/or perform content categorization analysis on this frame and a subset of the other frames of the video sequence.

Claims (52)

1. An apparatus comprising:

a motion estimation unit comprising circuitry configured to:

preprocess an input frame; and

generate an indication, based at least in part on a comparison of the input frame to a previous frame; and

machine learning (ML) engine circuitry configured to:

responsive to the indication meeting a condition:

process the input frame; and

provide one or more outputs of the ML engine circuitry to an encoder for encoding the input frame; and

responsive to the indication not meeting the condition, prevent processing of the input frame by the ML engine circuitry.

2. The apparatus as recited in claim 1 , wherein the motion estimation unit is further configured to identify one or more objects in the input frame, and wherein the ML engine circuitry is further configured to process only the one or more objects identified by the motion estimation unit, responsive to the indication meeting the condition.

3. The apparatus as recited in claim 1 , wherein:

responsive to the indication meeting a condition, the input frame is processed by the ML engine circuitry to determine a bit budget for the input frame; and

responsive to the indication not meeting the condition, the bit budget for the frame is calculated based on a bit budget of a previous frame.

4. The apparatus as recited in claim 1 , wherein the one or more outputs provided to the encoder comprise a quantization parameter (QP) map.

5. The apparatus as recited in claim 1 , wherein the indication comprises one or more motion vectors.

6. The apparatus as recited in claim 3 , wherein the condition comprises the indication indicating the input frame has a threshold number of changes compared to the previous frame.

7. The apparatus as recited in claim 1 , wherein the ML engine is further configured to process a subset of objects in the input frame, wherein the subset of objects are identified by the motion estimation unit.

8. A method comprising:

preprocessing, by a motion estimation unit comprising circuitry, an input frame;

generating an indication, by the motion estimation unit, based at least in part on a comparison of the input frame to a previous frame;

responsive to the indication meeting a condition:

processing, by a machine learning (ML) engine comprising circuitry, the input frame; and

providing one or more outputs of the ML engine to an encoder for encoding the input frame; and

responsive to the indication not meeting the condition, preventing processing of the input frame by the ML engine.

9. The method as recited in claim 8 , wherein responsive to the indication meeting the condition, the method further comprises:

identifying, by the motion estimation unit, one or more objects in the input frame; and

processing, by the machine learning engine, only the one or more objects identified by the motion estimation unit.

10. The method as recited in claim 8 , further comprising

responsive to the indication meeting a condition, processing the input frame by the ML engine to determine a bit budget for the input frame; and

responsive to the indication not meeting the condition, determining the bit budget for the input frame based on a bit budget of a previous frame.

11. The method as recited in claim 8 , wherein the one or more outputs provided to the encoder comprise a quantization parameter (QP) map.

12. The method as recited in claim 8 , wherein the indication comprises one or more motion vectors.

13. The method as recited in claim 10 , wherein the condition comprises the indication indicating the input frame has a threshold number of changes compared to the previous frame.

14. The method as recited in claim 8 , further comprising processing a subset of objects in the input frame, wherein the subset of objects are identified by the motion estimation unit.

15. A system comprising:

a memory storing at least a portion of an input frame; and

a processor comprising circuitry configured to:

preprocess an input frame; and

generate an indication, based at least in part on a comparison of the input frame to a previous frame;

responsive to the indication meeting a condition:

process the input frame; and

provide one or more outputs of the ML engine circuitry to an encoder for encoding the input frame; and

responsive to the indication not meeting the condition, prevent processing of the input frame by the ML engine circuitry.

16. The system as recited in claim 15 , wherein the processor is further configured to:

identify, during preprocessing, one or more objects in the input frame; and

process, by the machine learning engine circuitry, only the one or more objects identified by a motion estimation unit.

17. The system as recited in claim 16 , wherein:

responsive to the indication meeting a condition, the input frame is processed by the ML engine circuitry to determine a bit budget for the input frame; and

responsive to the indication not meeting the condition, the bit budget for the frame is calculated based on a bit budget of a previous frame.

18. The system as recited in claim 15 , wherein the one or more outputs provided to the encoder comprise a quantization parameter (QP) map.

19. The system as recited in claim 15 , wherein the indication comprises one or more motion vectors.

20. The system as recited in claim 19 , wherein the processor is further configured to process a downscaled version of the input frame by the ML engine circuitry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2021
From: KOTEYAR, SUNIL GOPAL; SHAO, MINGKAI
To: ATI TECHNOLOGIES ULC
Reel/Frame 057899/0283 →
Continuity (1)
Related Publication 20230095541A1 · Mar 30, 2023