IP Library › Granted Patent US 12,591,947
Granted Patent B2
US 12,591,947 · App. 18/433,133 · Granted Mar 31, 2026

Distortion-based image rendering

Inventors: Sajid Sadi (San Jose, CA); Varun Menon (Mountain View, CA); Siddarth Ravichandran (Santa Clara, CA); Chuhua Wang (Sunnyvale, CA); Hyun Jae Kang (Mountain View, CA); Rahul Lokesh (Sunnyvale, CA); Vignesh Gokul (Mountain View, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T3/18G06T11/60G06V10/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,947
App. No.
18/433,133
Granted
Mar 31, 2026
Kind
B2
Abstract

Synthesizing high-resolution input for rendering a digital human includes generating, with a generative artificial intelligence (AI) model, a distorted image of the digital human by enhancing a region of interest (ROI) within the distorted image relative to other regions of the distorted image. The generative AI model is previously trained against a distorted control image generated using a distortion function to distort a control image used to guide image generation by the generative AI model. The distorted control image is generated by reconfiguring and augmenting pixels of the control image. An undistorted image of the digital human is generated using a reverse distortion function to reverse distortion of the distorted image.

Claims (36)

1 . A computer-implemented method of rendering a digital human using generative artificial intelligence (AI), the method comprising:

generating, with a generative AI model, a distorted image of the digital human by enhancing a region of interest (ROI) within the distorted image relative to other regions of the distorted image;

wherein the generative AI model is trained against a distorted control image generated using a distortion function to distort a control image used to guide image generation by the generative AI model, the distorted control image generated by reconfiguring pixels of the control image with the distortion function and augmenting the pixels with pixels generated by interpolation; and

generating an undistorted image of the digital human using a reverse distortion function to reverse distortion of the distorted image.

2 . The computer-implemented method of claim 1 , wherein the generative AI model learns to generate the distorted image from data input to the generative AI model, wherein the data comprises a distorted contour image generated from landmarks of the control image in which the landmarks are distorted by the distortion function.

3 . The computer-implemented method of claim 1 , wherein the generative AI model learns to generate the distorted image from data input to the generative AI model, wherein the data comprises a distorted contour image, and wherein a mouth of the distorted image is generated using audio data.

4 . The computer-implemented method of claim 1 , wherein the distortion function is a continuous monotonically non-decreasing function that distorts the ROI by generating at least one of a spline or radial expansion from an approximate center of the ROI.

5 . The computer-implemented method of claim 1 , wherein the reverse distortion function is generated by fitting a polynomial to sampled points of the distorted image.

6 . The computer-implemented method of claim 1 , wherein the control image includes one or more additional ROIs, the method further comprising:

sequentially distorting the control image using the distortion function to distend each of the one or more additional ROIs by re-aligning and augmenting pixels corresponding to each of the one or more additional ROIs; and

generating the undistorted image by applying the reverse distortion function to the ROI and the one or more additional ROIs in a reverse sequence of the distorting the ROI and the one or more additional ROIs.

7 . A computer-implemented method of training a generative artificial intelligence (AI) model, the method comprising:

identifying a region of interest (ROI) within a control image of a human;

generating a distorted control image by distorting the control image using a distortion function that distends the ROI by reconfiguring and augmenting pixels of the control image corresponding to the ROI to thereby expand the ROI relative to other regions of the control image; and

generating, using a generative AI model, a distorted image, wherein the generative AI model learns to generate the distorted image against the distorted control image as distorted by the distortion function.

8 . The computer-implemented method of claim 7 , wherein the generative AI model learns to generate the distorted image from data input to the generative AI model, wherein the data comprises a distorted contour image generated from landmarks of the control image in which the landmarks are distorted by the distortion function.

9 . The computer-implemented method of claim 7 , wherein the generative AI model learns to generate the distorted AI image using multimodal data input to the generative AI model.

10 . The computer-implemented method of claim 9 , wherein the multimodal data includes a distorted contour image, and wherein a mouth of the distorted image is generated using audio data.

11 . The computer-implemented method of claim 7 , wherein the distortion function is a monotonically non-decreasing function that distorts the ROI by generating at least one of a spline or radial expansion from an approximate center of the ROI.

12 . The computer-implemented method of claim 7 , wherein the generative AI model is a generative adversarial network.

13 . The computer-implemented method of claim 7 , wherein the control image includes one or more additional ROIs, the method further comprising:

sequentially distorting the control image using the distortion function to distend each of the one or more additional ROIs by reconfiguring and augmenting pixels corresponding to each of the one or more additional ROIs; and

generating an undistorted image by applying a reverse distortion function to the ROI and to the one or more additional ROIs in a reverse sequence of the distorting of the ROI and the one or more additional ROIs.

14 . A system, comprising:

one or more processors configured to execute operations including:

generating, with a generative AI model, a distorted image of a digital human by enhancing a region of interest (ROI) within the distorted image relative to other regions of the distorted image;

wherein the generative AI model is trained against a distorted control image generated using a distortion function to distort a control image used to guide image generation by the generative AI model, the distorted control image generated by reconfiguring pixels of the control image with the distortion function and augmenting the pixels with pixels generated by interpolation; and

generating an undistorted image of the digital human using a reverse distortion function to reverse distortion of the distorted image.

15 . The system of claim 14 , wherein the generative AI model learns to generate the distorted image from data input to the generative AI model, wherein the data comprises a distorted contour image generated from landmarks of the control image in which the landmarks are distorted by the distortion function.

16 . The system of claim 14 , wherein the generative AI model learns to generate the distorted image using multimodal data input to the generative AI model.

17 . The system of claim 16 , wherein the multimodal data includes a distorted contour image, and wherein a mouth of the distorted image is generated using audio data.

18 . The system of claim 14 , wherein the distortion function is a monotonically non-decreasing function that distorts the ROI by generating at least one of a spline or radial expansion from an approximate center of the ROI.

19 . The system of claim 14 , wherein the reverse distortion function is generated by fitting a polynomial to sampled points of the distorted image.

20 . The system of claim 14 , wherein the control image includes one or more additional ROIs, wherein the one or more processors are configured to execute operations further comprising:

sequentially distorting the control image of the human using the distortion function to distend each of the one or more additional ROIs by reconfiguring and augmenting pixels corresponding to each of the one or more additional ROIs; and

generating the undistorted image by applying the reverse distortion function to the ROI and to the one or more additional ROIs in a reverse sequence of the distorting of the ROI and the additional ROIs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2024
From: SADI, SAJID; MENON, VARUN; RAVICHANDRAN, SIDDARTH; WANG, CHUHUA; KANG, HYUN JAE; LOKESH, RAHUL; GOKUL, VIGNESH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066406/0573 →
Continuity (2)
Provisional Application 63468854 · May 25, 2023
Related Publication 20240394830A1 · Nov 28, 2024
References Cited (21)
US 11158102B2 · Liu et al. · 2021 [cited by applicant]
US 20210327404A1 · Savchenkov · 2021 [cited by examiner]
US 20230014604A1 · Kim et al. · 2023 [cited by applicant]
US 20230042654A1 · Zhang · 2023 [cited by examiner]
US 20240055015A1 · Chae · 2024 [cited by examiner]
CM 113554737A · 2021 [cited by applicant]
CN 107644228A · 2018 [cited by applicant]
CN 111652796A · 2020 [cited by applicant]
CN 113609255A · 2021 [cited by applicant]
WO 2020091891A1 · 2020 [cited by applicant]
WO 2022195305A1 · 2022 [cited by applicant]
WO 2022255529A1 · 2022 [cited by applicant]
WIPO Int'l. Appln. PCT/KR2024/004003, International Search Report, Jul. 16, 2024, 3 pg. [cited by applicant]
WIPO Int'l. Appln. PCT/KR2024/004003, Written Opinion, Jul. 16, 2024, 5 pg. [cited by applicant]
Ma, Y. et al., “Variable rate roi image compression optimized for visual quality,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2021, Sep. 1, 2021, pp. 1936-1940. [cited by applicant]
Siarohin, A. et al., “First order motion model for image animation,” Advances in Neural Information Processing Systems, vol. 32, 2019. [cited by applicant]
Prajwal, KR et al., “A lip sync expert is all you need for speech to lip generation in the wild,” In Proceedings of the 28th ACM International Conference on Multimedia, Oct. 12, 2020, pp. 484-492. [cited by applicant]
“Best AI Video Generator in 2024—Synthesia,” [online] Synthesia Limited © 2024 [retrieved Jan. 25, 2024], retrieved from the Internet: <https://www.synthesia.io/>, 9 pg. [cited by applicant]
EPO Appln. 24811245.0, Extended European Search Report, Dec. 11, 2025, 11 pg. [cited by applicant]
Chu, W. et al., “Learning to Caricature Via Semantic Shape Transform,” Int'l. J. of Computer Vision, vol. 129, No. 9, Jul. 9, 2021, pp. 2663-2679. [cited by applicant]
Ravichandran, S. et al., “Synthesizing Photorealistic Virtual Humans Through Cross-model Disentanglement,” In Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Mar. 24. 2023, arXiv:2209.01320v… [cited by applicant]