Context-aware error concealment to improve inference accuracy
In various examples, systems and methods are disclosed relating to context-aware error concealment to improve inference accuracy are disclosed. A system can identify a frame of a video stream and determine that the frame comprises corrupted or lost data. The system can generate a corrected frame by applying an error concealment function selected based at least on a location of the corrupted or lost data in the frame and a region of interest in the frame.
1 . One or more processors comprising:
one or more circuits to:
identify a frame of a video stream, the frame having at least one region of interest;
determine that the frame comprises corrupted or lost data;
select a first error concealment function for a first portion of the corrupted or lost data within the at least one region of interest, and a second error concealment function for a second portion of the corrupted or lost data outside of the at least one region of interest in the frame; and
generate a corrected frame by applying the first error concealment function to the first portion of the corrupted or lost data and by applying the second error concealment function to the second portion of the corrupted or lost data.
2 . The one or more processors of claim 1 , wherein the one or more circuits are to apply the first error concealment function or the second error concealment function by providing the frame as input to a machine-learning model.
3 . The one or more processors of claim 1 , wherein the one or more circuits are to:
receive an encoded bitstream of the video stream; and
generate the frame by decoding the encoded bitstream, wherein decoding the frame indicates a location of the corrupted or lost data.
4 . The one or more processors of claim 1 , wherein the first error concealment function and the second error concealment function are two of three or more error concealment functions.
5 . The one or more processors of claim 1 , wherein the one or more circuits are to:
determine that an object is detected in a predetermined number of prior frames in the video stream; and
select the first error concealment function from a plurality of error concealment functions based at least on the object being detected in the predetermined number of prior frames, the first error concealment function using a greater amount of computing resources relative to the second error concealment function of the plurality of error concealment functions.
6 . The one or more processors of claim 5 , wherein the one or more circuits are to:
select the first error concealment function based on the object being estimated to appear at least in the first portion of the corrupted or lost data.
7 . The one or more processors of claim 1 , wherein the one or more processors are comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing conversational AI operations;
a system for performing generative AI operations using a large language model (LLM);
a system for performing generative AI operations using a vision language model (VLM);
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
8 . A system, comprising:
one or more processors to:
receive a request to process a video stream comprising a plurality of frames;
apply a first error concealment function and a second error concealment function to at least one frame of the plurality of frames, wherein the first error concealment function is selected for a first portion of corrupted or lost data of the at least one frame within at least one region of interest, and the second error concealment function is selected for a second portion of the corrupted or lost data of the at least one frame outside of the at least one region of interest; and
execute a machine-learning model using the at least one frame as input for the request.
9 . The system of claim 8 , wherein the one or more processors are to:
determine that the at least one frame comprises the corrupted or lost data using a decoding process.
10 . The system of claim 8 , wherein the machine-learning model generates an indication of an object in the at least one frame, and wherein the one or more processors are to:
select a third error concealment function for at least one second frame of the plurality of frames based at least on the indication of the object in the at least one frame.
11 . The system of claim 10 , wherein the third error concealment function for the at least one second frame is selected further based on an expected location of the object in the at least one second frame.
12 . The system of claim 8 , wherein a location of the corrupted or lost data is one of a macroblock location or a slice location of the at least one frame.
13 . The system of claim 8 , wherein the one or more processors are to:
apply at least one of the first error concealment function or the second error concealment function by executing a second machine-learning model using the at least one frame as input, the second machine-learning model to generate replacement information for the corrupted or lost data of the at least one frame.
14 . The system of claim 8 , wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing conversational AI operations;
a system for performing generative AI operations using a large language model (LLM);
a system for performing generative AI operations using a vision language model (VLM);
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
15 . A method, comprising:
identifying, using one or more processors, a frame of a video stream, the frame having at least one region of interest;
determining, using the one or more processors, that the frame comprises corrupted or lost data;
selecting, using the one or more processors, a first error concealment function for a first portion of the corrupted or lost data within the at least one region of interest, and a second error concealment function for a second portion of the corrupted or lost data outside of the at least one region of interest in the frame; and
generating, using the one or more processors, a corrected frame by applying the first error concealment function to the first portion of the corrupted or lost data and the second error concealment function to the second portion of the corrupted or lost data.