IP Library Granted Patent US 11,983,906
Granted Patent B2
US 11,983,906 · App. 17/704,907 · Granted May 14, 2024

Systems and methods for image compression at multiple, different bitrates

Inventors: Christopher Schroers (Uster, CH); Erika Doggett (Los Angeles, CA); Stephan Mandt (Santa Monica, CA); Jared Mcphillen (Glendale, CA); Scott Labrozzi (Cary, NC); Romann Weber (Uster, CH); Mauro Bamert (Nafels, CH)
Assignee: Disney Enterprises, Inc.
G06T9/20G06T9/002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,983,906
App. No.
17/704,907
Granted
May 14, 2024
Kind
B2
Abstract

Systems and methods for predicting a target set of pixels are disclosed. In one embodiment, a method may include obtaining target content. The target content may include a target set of pixels to be predicted. The method may also include convolving the target set of pixels to generate an estimated set of pixels. The method may include matching a second set of pixels in the target content to the target set of pixels. The second set of pixels may be within a distance from the target set of pixels. The method may include refining the estimated set of pixels to generate a refined set of pixels using a second set of pixels in the target content.

Claims (43)

1. A computer-implemented method for compressing target content, the method comprising:

selecting a target set of pixels from a plurality of pixels included in the target content;

generating, using an auto-encoder included in a prediction neural network, an estimated set of pixels based on one or more convolutions performed on the target set of pixels;

matching the target set of pixels to a second set of pixels included in the plurality of pixels; and

generating, using a decoder included in the prediction neural network based on input that includes the estimated set of pixels and the second set of pixels, a refined set of pixels.

2. The computer-implemented method of claim 1 , further comprising:

generating a residual between the refined set of pixels and the target set of pixels.

3. The computer-implemented method of claim 1 , wherein matching the target set of pixels to the second set of pixels comprises generating displacement data between the target set of pixels and the second set of pixels, wherein the displacement data is encoded using a threshold number of bits.

4. The computer-implemented method of claim 1 , wherein generating the refined set of pixels comprises applying a convolution to the target set of pixels to generate displacement data corresponding to displacement from the second set of pixels to the target set of pixels.

5. The computer-implemented method of claim 1 , wherein generating the estimated set of pixels comprises generating a different weight corresponding to each pixel included in the estimated set of pixels.

6. The computer-implemented method of claim 1 , wherein generating the estimated set of pixels comprises concatenating supplemental information to the target set of pixels.

7. The computer-implemented method of claim 6 , wherein the supplemental information comprises one or more of a mask applied to the target set of pixels, one or more color values corresponding to the target set of pixels, or one or more contours corresponding to the target set of pixels.

8. The computer-implemented method of claim 1 , wherein generating the estimated set of pixels comprises applying a mask to the target set of pixels.

9. The computer-implemented method of claim 1 , further comprising displacing the second set of pixels based on a position of the estimated set of pixels, wherein generating the refined set of pixels is further based on the displaced second set of pixels.

10. The computer-implemented method of claim 1 , wherein generating the refined set of pixels comprises:

encoding the estimated set of pixels using a first encoder included in the prediction neural network to generate an encoded estimated set of pixels;

encoding the second set of pixels using a second encoder included in the prediction neural network to generate an encoded second set of pixels; and

applying the decoder included in the prediction neural network to the encoded estimated set of pixels and the encoded second set of pixels.

11. One or more non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

selecting a target set of pixels from a plurality of pixels included in a target content;

generating, using an auto-encoder included in a prediction neural network, an estimated set of pixels based on one or more convolutions performed on the target set of pixels;

matching the target set of pixels to a second set of pixels included in the plurality of pixels; and

generating, using a decoder included in the prediction neural network based on input that includes the estimated set of pixels and the second set of pixels, a refined set of pixels.

12. The one or more non-transitory computer-readable medium of claim 11 , wherein generating the refined set of pixels comprises:

encoding the estimated set of pixels using a first encoder included in the prediction neural network to generate an encoded estimated set of pixels;

encoding the second set of pixels using a second encoder included in the prediction neural network to generate an encoded second set of pixels;

concatenating the encoded estimated set of pixels and the encoded second set of pixels to generate a concatenated set of pixels; and

applying the decoder included in the prediction neural network to the concatenated set of pixels.

13. The one or more non-transitory computer-readable medium of claim 11 , wherein matching the target set of pixels to the second set of pixels is based on displacement data corresponding to displacement between the target set of pixels and the second set of pixels.

14. The one or more non-transitory computer-readable medium of claim 11 , wherein matching the target set of pixels to the second set of pixels is based on a distance from each pixel included in the second set of pixels to an edge of the target set of pixels.

15. The one or more non-transitory computer-readable medium of claim 11 , wherein matching the target set of pixels to the second set of pixels is based on a similarity between a color of each pixel included in the second set of pixels to one or more colors of the target set of pixels.

16. The one or more non-transitory computer-readable medium of claim 11 , wherein matching the target set of pixels to the second set of pixels is based on a similarity between one or more objects depicted in the second set of pixels and one or more objects depicted in the target set of pixels.

17. The one or more non-transitory computer-readable medium of claim 11 , wherein the plurality of pixels are included in a content frame, and wherein selecting the target set of pixels is based on at least one of: a similarity between the target set of pixels and one or more other pixels included in the plurality of pixels, a similarity between the target set of pixels and one or more other pixels included in a previous content frame, a similarity between the target set of pixels and one or more other pixels included in a subsequent content frame, an order associated with encoding the plurality of pixels, or an order associated with decoding the plurality of pixels.

18. The one or more non-transitory computer-readable medium of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to perform the step of generating a residual between the refined set of pixels and the target set of pixels.

19. The one or more non-transitory computer-readable medium of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to perform the step of compressing the refined set of pixels to generate a compressed target content.

20. A system comprising:

one or more memories storing instructions; and

one or more processors that are coupled to the one or more memories and,

when executing the instructions, are configured to:

select a target set of pixels from a plurality of pixels included in a target content;

generate, using an auto-encoder included in a prediction neural network, an estimated set of pixels based on one or more convolutions performed on the target set of pixels;

match the target set of pixels to a second set of pixels included in the plurality of pixels; and

generate, using a decoder included in the prediction neural network based on input that includes the estimated set of pixels and the second set of pixels, a refined set of pixels.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2022
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 059450/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2022
From: SCHROERS, CHRISTOPHER; BAMERT, MAURO; WEBER, ROMANN
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 059413/0598 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2022
From: DOGGETT, ERIKA; MCPHILLEN, JARED; LABROZZI, SCOTT; MANDT, STEPHAN MARCEL
To: DISNEY ENTERPRISES, INC.
Reel/Frame 059418/0111 →
Continuity (2)
Division 16249861 · Jan 16, 2019
Related Publication 20220215595A1 · Jul 7, 2022