IP Library Granted Patent US 12694667
Granted Patent B2
US 12694667 · App. 18/505,010 · Granted Jul 28, 2026

Systems and methods for segmentation map error correction

Inventors: Chung-Chi Tsai (San Diego, CA); Hau Hwang (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06V10/98G06T3/40G06T7/12G06V10/764G06V10/774G06T2207/10016G06T2207/20021G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694667
App. No.
18/505,010
Granted
Jul 28, 2026
Kind
B2
Abstract

Imaging systems and techniques are described. In some examples, an imaging system generates a segmentation map of an image by processing image data associated with the image using a segmentation mapper. Different object types in the image are categorized into different regions in the segmentation map. The imaging system generates an augmented segmentation map by processing at least the segmentation map using a segmentation map error correction engine. The imaging system generates processed image data by processing the image using the augmented segmentation map.

Claims (51)

1 . An apparatus for error correction, the apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

generate a segmentation map of a first video frame of a video by processing image data associated with the first video frame using a segmentation mapper, wherein different object types in the first video frame are categorized into different regions in the segmentation map;

generate an augmented segmentation map by processing at least the segmentation map and a second video frame of the video using a segmentation map error correction engine;

generate processed image data by processing the first video frame using the augmented segmentation map; and

output a processed variant of the video that includes the processed image data.

2 . The apparatus of claim 1 , wherein the image data associated with the first video frame includes a downscaled variant of the first video frame, wherein the augmented segmentation map is upscaled relative to the segmentation map.

3 . The apparatus of claim 1 , wherein, to process at least the segmentation map and the second video frame using the segmentation map error correction engine, the at least one processor is configured to:

process the image data, the segmentation map, and the second video frame using the segmentation map error correction engine.

4 . The apparatus of claim 1 , wherein the image is a second video frame is temporally adjacent to the first video frame within the video.

5 . The apparatus of claim 1 , wherein the at least one processor is configured to:

generate an augmented second segmentation map associated with the second video frame of the video by processing at least the image data, the segmentation map, and secondary image data associated with a second frame of the video using the segmentation map error correction engine; and

generate a processed second video frame of the video by processing the second video frame of the video using the augmented second segmentation map, wherein the processed variant of the video includes the processed second video frame.

6 . The apparatus of claim 1 , wherein the augmented segmentation map includes information that is based on the second video frame of the video.

7 . The apparatus of claim 1 , wherein the segmentation map error correction engine includes a trained machine learning model, and wherein, to process at least the segmentation map using the segmentation map error correction engine, the at least one processor is configured to input at least the segmentation map and the second video frame into the trained machine learning model to process at least the segmentation map and the second video frame using the trained machine learning model.

8 . The apparatus of claim 7 , the trained machine learning model having been trained based on training data, the training data including an image, a first segmentation map generated using the image with one or more image processing operations applied, and a second segmentation map generated using the image without the one or more image processing operations applied.

9 . The apparatus of claim 8 , wherein the one or more image processing operations include at least one of a resampling filter, a blur filter, logit degradation, a perspective transform, a flip, a rotation, a shift, or a crop.

10 . The apparatus of claim 7 , wherein the trained machine learning model is an error correction network (ECN).

11 . The apparatus of claim 7 , wherein the at least one processor is configured to:

update the trained machine learning model based on at least the image data, the segmentation map, and the augmented segmentation map.

12 . The apparatus of claim 1 , wherein, to process the first video frame using the augmented segmentation map, the at least one processor is configured to:

process different regions of the first video frame using to different processing settings according to the augmented segmentation map, wherein the different regions of the first video frame are based on different regions of the augmented segmentation map.

13 . The apparatus of claim 12 , wherein the different processing settings indicate different strengths at which to apply a specified image processing function to the different regions of the first video frame.

14 . The apparatus of claim 1 , wherein an edge in the first video frame aligns more closely to a first corresponding edge in the augmented segmentation map than to a second corresponding edge in the segmentation map.

15 . The apparatus of claim 1 , wherein, to output the processed variant of the video, the at least one processor is configured to:

output the processed image data.

16 . The apparatus of claim 1 , wherein, to output the processed variant of the video, the at least one processor is configured to:

display the processed image data.

17 . The apparatus of claim 1 , wherein, to output the processed variant of the video, the at least one processor is configured to:

transmit the processed image data to a recipient device.

18 . The apparatus of claim 1 , wherein the apparatus includes at least one of a head-mounted display (HMD), a mobile handset, or a wireless communication device.

19 . A method of error correction, the method comprising:

generating a segmentation map of a first video frame of a video by processing image data associated with the first video frame using a segmentation mapper, wherein different object types in the first video frame are categorized into different regions in the segmentation map;

generating an augmented segmentation map by processing at least the segmentation map and a second video frame of the video using a segmentation map error correction engine;

generating processed image data by processing the first video frame using the augmented segmentation map; and

output a processed variant of the video that includes the processed image data.

20 . The method of claim 19 , wherein the image data associated with the first video frame includes a downscaled variant of the first video frame, wherein the augmented segmentation map is upscaled relative to the segmentation map.

21 . The method of claim 19 , wherein processing at least the segmentation map and the second video frame using the segmentation map error correction engine includes processing the image data, the segmentation map, and the second video frame using the segmentation map error correction engine.

22 . The method of claim 19 , wherein the second video frame is temporally adjacent to the first video frame within the video.

23 . The method of claim 19 , further comprising:

generating an augmented second segmentation map associated with the second video frame of the video by processing at least the image data, the segmentation map, and secondary image data associated with a second frame of the video using the segmentation map error correction engine; and

generating a processed second video frame of the video by processing the second video frame of the video using the augmented second segmentation map, wherein the processed variant of the video includes the processed second video frame.

24 . The method of claim 19 , wherein the augmented segmentation map includes information that is based on the second video frame of the video.

25 . The method of claim 19 , wherein the segmentation map error correction engine includes a trained machine learning model, and wherein processing at least the segmentation map using the segmentation map error correction engine includes inputting at least the segmentation map and the second video frame into the trained machine learning model to process at least the segmentation map and the second video frame using the trained machine learning model.

26 . The method of claim 25 , the trained machine learning model having been trained based on training data, the training data including an image, a first segmentation map generated using the image with one or more image processing operations applied, and a second segmentation map generated using the image without the one or more image processing operations applied.

27 . The method of claim 25 , further comprising:

updating the trained machine learning model based on at least the image data, the segmentation map, and the augmented segmentation map.

28 . The method of claim 19 , wherein processing the first video frame using the augmented segmentation map includes processing different regions of the first video frame using to different processing settings according to the augmented segmentation map, wherein the different regions of the first video frame are based on different regions of the augmented segmentation map.

29 . The method of claim 19 , wherein an edge in the first video frame aligns more closely to a first corresponding edge in the augmented segmentation map than to a second corresponding edge in the segmentation map.

30 . The method of claim 19 , wherein outputting the processed variant of the video includes outputting the processed image data.