IP Library Granted Patent US 12682500
Granted Patent B2
US 12682500 · App. 18/430,369 · Granted Jul 14, 2026

High-fidelity neural rendering of images

Inventors: Dimitar Petkov Dinev (Sunnyvale, CA); Siddarth Ravichandran (Santa Clara, CA); Hyun Jae Kang (Mountain View, CA); Ondrej Texler (San Jose, CA); Anthony Sylvain Jean-Yves Liot (San Jose, CA); Sajid Sadi (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T9/00G06T3/40G06T11/00G06V10/771
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682500
App. No.
18/430,369
Granted
Jul 14, 2026
Kind
B2
Abstract

Generating images includes generating encoded data by encoding input data into a latent space. The encoded data is decoded through a first decoder having first decoder layers by processing the encoded data through one or more of the first decoder layers. The encoded data is decoded through a second decoder having second decoder layers by processing the encoded data through one or more of the second decoder layers. An updated feature map is generated by replacing at least a portion of a feature map output from a selected layer of the first decoder layers with at least a portion of a feature map output from a selected layer of the second decoder layers. An image is generated by further decoding the updated feature map through one or more additional layers of the first decoder layers.

Claims (43)

1 . A method, comprising:

generating encoded data by encoding input data into a latent space;

decoding the encoded data through a first decoder having a plurality of first decoder layers by processing the encoded data through one or more of the plurality of first decoder layers;

decoding the encoded data through a second decoder having a plurality of second decoder layers by processing the encoded data through one or more of the plurality of second decoder layers;

generating an updated feature map by replacing at least a portion of a first feature map output from a selected layer of the plurality of first decoder layers with at least a portion of a second feature map output from a selected layer of the plurality of second decoder layers; and

generating an image by further decoding the updated feature map through one or more additional layers of the plurality of first decoder layers.

2 . The method of claim 1 , wherein the selected layer of the plurality of first decoder layers is a penultimate layer of the plurality of first decoder layers.

3 . The method of claim 1 , wherein the selected layer of the plurality of second decoder layers is a penultimate layer of the plurality of second decoder layers.

4 . The method of claim 1 , further comprising:

resizing the second feature map to correspond to a size of the portion of the first feature map.

5 . The method of claim 1 , wherein the image is a red, green, blue image.

6 . The method of claim 1 , wherein the first decoder is trained to generate images including a face of a digital human.

7 . The method of claim 6 , wherein the second decoder is trained to generate images including a mouth of the digital human.

8 . The method of claim 1 , further comprising:

decoding the encoded data through one or more additional decoders each having a plurality decoder layers by processing the encoded data through one or more of the plurality of decoder layers of each of the one or more additional decoders;

wherein the updated feature map is generated by replacing one or more additional portions of the first feature map with a further feature map output from a selected layer of the plurality of decoder layers from each of the one or more additional decoders.

9 . The method of claim 1 , wherein an output from a final layer of the plurality of second decoder layers of the second decoder is used only during training of the second decoder.

10 . A system, comprising:

a processor configured to execute operations including:

generating encoded data by encoding input data into a latent space;

decoding the encoded data through a first decoder having a plurality of first decoder layers by processing the encoded data through one or more of the plurality of first decoder layers;

decoding the encoded data through a second decoder having a plurality of second decoder layers by processing the encoded data through one or more of the plurality of second decoder layers;

generating an updated feature map by replacing at least a portion of a first feature map output from a selected layer of the plurality of first decoder layers with at least a portion of a second feature map output from a selected layer of the plurality of second decoder layers; and

generating an image by further decoding the updated feature map through one or more additional layers of the plurality of first decoder layers.

11 . The system of claim 10 , wherein the selected layer of the plurality of first decoder layers is a penultimate layer of the plurality of first decoder layers.

12 . The system of claim 10 , wherein the selected layer of the plurality of second decoder layers is a penultimate layer of the plurality of second decoder layers.

13 . The system of claim 10 , wherein the processor is configured to execute operations comprising:

resizing the second feature map to correspond to a size of the portion of the first feature map.

14 . The system of claim 10 , wherein the image is a red, green, blue image.

15 . The system of claim 10 , wherein the first decoder is trained to generate images including a face of a digital human.

16 . The system of claim 15 , wherein the second decoder is trained to generate images including a mouth of the digital human.

17 . The system of claim 10 , wherein the processor is configured to execute operations comprising:

decoding the encoded data through one or more additional decoders each having a plurality decoder layers by processing the encoded data through one or more of the plurality of decoder layers of each of the one or more additional decoders;

wherein the updated feature map is generated by replacing one or more additional portions of the first feature map with a further feature map output from a selected layer of the plurality of decoder layers from each of the one or more additional decoders.

18 . A system, comprising:

an encoder configured to generate encoded data by encoding input data into a latent space;

a first decoder having a plurality of first decoder layers, wherein the first decoder is configured to decode the encoded data through one or more of the plurality of first decoder layers; and

a second decoder having a plurality of second decoder layers, wherein the second decoder is configured to decode the encoded data through one or more of the plurality of second decoder layers;

wherein the first decoder is configured to generate an updated feature map by replacing at least a portion of a first feature map output from a selected layer of the plurality of first decoder layers with at least a portion of a second feature map output from a selected layer of the plurality of second decoder layers; and

wherein the first decoder is further configured to generate an image by further decoding the updated feature map through one or more additional layers of the plurality of first decoder layers.

19 . The system of claim 18 , wherein the selected layer of the plurality of first decoder layers is a penultimate layer of the plurality of first decoder layers.

20 . The system of claim 18 , further comprising:

a resizer configured to resize the second feature map to correspond to a size of the portion of the first feature map output from the selected layer of the plurality of first decoder layers.