IP Library Granted Patent US 12,602,849
Granted Patent B2
US 12,602,849 · App. 18/342,726 · Granted Apr 14, 2026

Image generation using one-dimensional inputs

Inventors: Hyun Jae Kang (Mountain View, CA); Siddarth Ravichandran (Santa Clara, CA); Ondrej Texler (San Jose, CA); Dimitar Petkov Dinev (Sunnyvale, CA); Anthony Sylvain Jean-Yves Liot (San Jose, CA); Sajid Sadi (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T11/60G06T13/40G06T13/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,849
App. No.
18/342,726
Granted
Apr 14, 2026
Kind
B2
Abstract

Image-to-image translations using 1D inputs includes concatenating multiple 1D vectors forming a concatenated 1D vector. The multiplicity of 1D vectors includes 1D vectors of at least two different modalities. An encoded 1D vector is generated by encoding the concatenated 1D vector. An encoded 2D array of features is generated by reshaping an arrangement of features of the encoded 1D feature vector. An image of a virtual human is generated by decoding the encoded 2D array.

Claims (40)

1 . A computer-implemented method, comprising:

concatenating a plurality of 1D vectors forming a concatenated 1D vector, wherein the plurality of 1D vectors include 1D vectors of at least two different modalities;

generating an encoded 1D feature vector by encoding the concatenated 1D vector;

generating an encoded 2D array of features by reshaping an arrangement of features of the encoded 1D feature vector; and

generating an image of a virtual human by decoding the encoded 2D array.

2 . The computer-implemented method of claim 1 , further comprising:

generating a video rendering of the virtual human, wherein the video rendering includes the image of the virtual human.

3 . The computer-implemented method of claim 1 , wherein each 1D vector of the plurality of 1D vectors is of a different modality.

4 . The computer-implemented method of claim 1 , wherein the plurality of 1D vectors includes a 1D vector having a first modality including head pose information.

5 . The computer-implemented method of claim 1 , wherein the plurality of 1D vectors includes a 1D vector having a second modality including features representing audio data.

6 . The computer-implemented method of claim 5 , wherein the features representing audio data includes a plurality of viseme coefficients.

7 . The computer-implemented method of claim 1 , wherein the plurality of 1D vectors includes a 1D vector with features comprising blend shape coefficients.

8 . A computer-implemented method, comprising:

generating an encoded 2D image template by encoding a 2D static template image;

generating an encoded 1D feature vector by encoding a concatenation of a plurality of 1D vectors;

generating an encoded 2D array of features by reshaping an arrangement of multiple features of the encoded 1D feature vector;

concatenating the encoded 2D image template and the encoded 2D array of features;

generating a fused tensor by fusing the encoded 2D image template with the encoded 2D array of features; and

generating an image of a virtual human by decoding the fused tensor, wherein the image of the virtual human is changeable to a different virtual human in response to encoding a different 2D static template image.

9 . The computer-implemented method of claim 8 , further comprising:

generating a video rendering of the virtual human, wherein the video rendering includes the image of the virtual human.

10 . The computer-implemented method of claim 8 , further comprising:

changing at least one visual characteristic of the image of the virtual human in response to encoding a different concatenation of 1D vectors.

11 . The computer-implemented method of claim 10 , wherein the changing at least one visual characteristic includes changing a direction of illumination of the image of the virtual human.

12 . The computer-implemented method of claim 8 , wherein the concatenating concatenates the encoded 2D image template and the encoded 2D array of features along a channel dimension.

13 . The computer-implemented method of claim 8 , wherein the generating the encoded 1D feature vector includes encoding a concatenation of a plurality of 1D vectors that each comprise multiple features corresponding to different modalities.

14 . The computer-implemented method of claim 8 , wherein the plurality of 1D vectors includes a 1D vector whose features include head pose information.

15 . The computer-implemented method of claim 8 , wherein the plurality of 1D vectors includes a 1D vector whose features comprise blend shape coefficients.

16 . The computer-implemented method of claim 8 , wherein the plurality of 1D vectors includes a 1D vector comprising a plurality of features representing audio data.

17 . The computer-implemented method of claim 16 , wherein the plurality of features representing audio data includes a plurality of viseme coefficients.

18 . A computer-implemented method, comprising:

generating an encoded 2D image template by encoding a 2D static template image;

generating an encoded 1D feature vector by encoding a concatenation of a plurality of 1D vectors;

generating an encoded 2D array of features by reshaping an arrangement of multiple features of the encoded 1D feature vector;

concatenating the encoded 2D image template and the encoded 2D array of features;

generating a fused tensor by fusing the encoded 2D image template with the encoded 2D array of features; and

generating an image of a virtual human by decoding the fused tensor, wherein at least one visual characteristic of the image of the virtual human is changeable in response to encoding a different concatenation of 1D vectors.

19 . The computer-implemented method of claim 18 , further comprising:

generating a video rendering of the virtual human, wherein the video rendering includes the image of the virtual human.

20 . The computer-implemented method of claim 18 , wherein the generating the encoded 1D feature vector includes encoding a concatenation of a plurality of 1D vectors that each comprise multiple features corresponding to different modalities.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: KANG, HYUN JAE; RAVICHANDRAN, SIDDARTH; TEXLER, ONDREJ; DINEV, DIMITAR PETKOV; LIOT, ANTHONY SYLVAIN JEAN-YVES; SADI, SAJID
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064089/0106 →
Continuity (2)
Provisional Application 63436211 · Dec 30, 2022
Related Publication 20240221254A1 · Jul 4, 2024
References Cited (51)
US 5907351A · Chen et al. · 1999 [cited by applicant]
US 6839672B1 · Beutnagel et al. · 2005 [cited by applicant]
US 7260539B2 · Cosatto et al. · 2007 [cited by applicant]
US 9613450B2 · Wang et al. · 2017 [cited by applicant]
US 10304439B2 · Okaniwa et al. · 2019 [cited by applicant]
US 10521946B1 · Roche et al. · 2019 [cited by applicant]
US 11158102B2 · Liu et al. · 2021 [cited by applicant]
US 11195053B2 · Kim et al. · 2021 [cited by applicant]
US 11270487B1 · Steptoe · 2022 [cited by applicant]
US 11410570B1 · Yang et al. · 2022 [cited by applicant]
US 11514634B2 · Liao et al. · 2022 [cited by applicant]
US 20100189342A1 · Parr et al. · 2010 [cited by applicant]
US 20150187112A1 · Rozen · 2015 [cited by applicant]
US 20160170975A1 · Jephcott · 2016 [cited by applicant]
US 20170011279A1 · Soldevila · 2017 [cited by examiner]
US 20170351935A1 · Liu et al. · 2017 [cited by applicant]
US 20190318194A1 · Liu et al. · 2019 [cited by applicant]
US 20200167605A1 · Kim · 2020 [cited by examiner]
US 20200226724A1 · Fang et al. · 2020 [cited by applicant]
US 20200380246A1 · Liu et al. · 2020 [cited by applicant]
US 20210090314A1 · Hussen Abdelaziz et al. · 2021 [cited by applicant]
US 20210166461A1 · Riesen et al. · 2021 [cited by applicant]
US 20210192824A1 · Chen et al. · 2021 [cited by applicant]
US 20210327404A1 · Savchenkov et al. · 2021 [cited by applicant]
US 20220068010A1 · Cambra et al. · 2022 [cited by applicant]
US 20220084273A1 · Pan et al. · 2022 [cited by applicant]
US 20220129689A1 · Kim et al. · 2022 [cited by applicant]
US 20220172462A1 · Wang et al. · 2022 [cited by applicant]
US 20220398794A1 · Lee · 2022 [cited by applicant]
US 20220399025A1 · Chae et al. · 2022 [cited by applicant]
US 20230042654A1 · Zhang · 2023 [cited by applicant]
US 20240013462A1 · Seol · 2024 [cited by applicant]
US 20240221260A1 · Dinev et al. · 2024 [cited by applicant]
CN 111292407A · 2020 [cited by applicant]
CN 113609255A · 2021 [cited by applicant]
CN 115170622A · 2022 [cited by applicant]
WO 2020226785A1 · 2020 [cited by applicant]
WIPO Appln. No. PCT/KR2023/020861, Written Opinion, Mar. 11, 2024, 4 pg. [cited by applicant]
WIPO Appln. No. PCT/KR2023/020861, International Search Report, Mar. 11, 2024, 4 pg. [cited by applicant]
WIPO Appln. No. PCT/KR2023/009802, International Search Report, Oct. 23, 2023, 4 pages. [cited by applicant]
WIPO Appln. No. PCT/KR2023/009802, Written Opinion, Oct. 23, 2023, 5 pg. [cited by applicant]
WIPO Appln. No. PCT/KR2023/017157, International Search Report, Feb. 13, 2024, 4 pg. [cited by applicant]
WIPO Appln. No. PCT/KR2023/017157, Written Opinion, Feb. 13, 2024, 4 pg. [cited by applicant]
Teng, W. et al., “Unimodal Face Classification with Multimodal Training,” with supplementary document, In 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021) (pp. 1-5). IEEE. [cited by applicant]
Ravichandran, S. et al., “Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglement,” arXiv:2209.01320v1, Sep. 3, 2022.10 pg. [cited by applicant]
Suwajanakorn, S, et al., “Synthesizing Obama: learning lip sync from audio,” ACM Transactions on Graphics (ToG), Jul. 20, 2017; vol. 36, No. 4, Art. 95, 13pg. [cited by applicant]
“Synthesia, #1 AI Video Creation Platform,” [online] synthesia.io, [retrieved Jun. 29, 2023], retrieved from the Internet: <https://www.synthesia.io/>, 7 pg. [cited by applicant]
Fruhstuck et al., “InsetGAN for Full-Body Image Generation,” In Proc. Of the IEEE/CVF Conf on Computer Vision and Pattern Recognition, 2022, pp. 7723-7732. [cited by applicant]
EP Appln. No. 23912522, Extended European Search Report, Oct. 13, 2025, 7 pg. [cited by applicant]
Sinha, S. et al., “Emotion-Controllable Generalized Talking Face Generation,” arXiv preprint No., arXiv:2205.01155v1, May 2, 2022, 10 pg. [cited by applicant]
Richard, A. et al., “Audio-and gaze-driven facial animation of codec avatars,” In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision 2021, pp. 41-50. [cited by applicant]