IP Library › Granted Patent US 12,361,750
Granted Patent B2
US 12,361,750 · App. 17/446,193 · Granted Jul 15, 2025

Synthetic emotion in continuously generated voice-to-video system

Inventors: Seth Jacob Rothschild (Littleton, MA); Alex Robbins (Cambridge, MA)
Assignee: EMC IP Holding Company LLC
G06V40/161G06V20/46G06V40/176G10L25/63G11B27/034G11B27/036G11B27/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,750
App. No.
17/446,193
Granted
Jul 15, 2025
Kind
B2
Abstract

One example method includes collecting an audio segment that includes audio data generated by a user, analyzing the audio data to identify an emotion expressed by the user, computing start and end indices of a video segment, selecting video data that shows the emotion expressed by the user, using the video data and the start and end indices of the video segment to modify a face of the user as the face appears in the video segment so as to generate modified face frames, and stitching the modified face frames into the video segment to create a modified video segment with the emotion expressed by the user, and the modified video segment includes the audio data generated by the user.

Claims (32)

1. A method, comprising:

collecting an audio segment that comprises audio data generated by a user;

analyzing the audio data to identify an emotion expressed by the user;

computing start and end indices for frames of a video segment, which includes a representation of a face of the user, and the representation comprises facial features;

selecting video data that shows the emotion expressed by the user from among a plurality of video data, of which each has been prerecorded by the user expressing a respective emotion;

using the video data and the start and end indices for the frames of the video segment to modify the representation appearing in the video segment so as to generate modified faces in the selected video data based on the audio segment so that the modified faces in the selected video data appear to be speaking words in the audio segment; and

stitching the modified faces from the selected video data over faces in the video segment, thereby swapping the faces in the video segment with the modified faces in the selected video data, to create a modified video segment with the emotion expressed by the user,

wherein the modified video segment includes the audio data generated by the user.

2. The method as recited in claim 1 , wherein the representation in the video segment is modified using a faceswap process that comprises replacing the representation in the video segment with an image of the face of the user taken from another video.

3. The method as recited in claim 2 , wherein selecting video data that shows the emotion expressed by the user comprises selecting a pre-recorded emotion video that shows the emotion expressed by the user.

4. The method as recited in claim 3 , further comprising computing start and end indices for the pre-recorded emotion video, and the faceswap process is performed based in part on the start and end indices of the pre-recorded emotion video.

5. The method as recited in claim 1 , comprising presenting the modified video segment so that although the audio data in the modified video segment does not comprise real time audio data, the modified video segment, as presented, still appears to be a real time video of the user.

6. The method as recited in claim 1 , wherein the representation in the video segment is modified using a facewarp process based on an input data set comprising labeled sequences of facial movements.

7. The method as recited in claim 6 , wherein the facewarp process is performed based in part on face movement rules.

8. The method as recited in claim 6 , wherein the audio segment of the modified video segment is presented to a viewer in near real time as the audio segment is being recorded.

9. A non-transitory computer readable storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:

collecting an audio segment that comprises audio data generated by a user;

analyzing the audio data to identify an emotion expressed by the user;

computing start and end indices for frames of a video segment, which includes a representation of a face of the user, and the representation comprises facial features;

selecting video data that shows the emotion expressed by the user from among a plurality of video data, of which each has been prerecorded by the user expressing a respective emotion;

using the video data and the start and end indices for the frames of the video segment to modify the representation appearing in the video segment so as to generate modified faces in the selected video data based on the audio segment so that the modified faces in the selected video data appear to be speaking words in the audio segment; and

stitching the modified faces from the selected video data over faces in the video segment, thereby swapping the faces in the video segment with the modified faces in the selected video data, to create a modified video segment with the emotion expressed by the user, and the modified video segment includes the audio data generated by the user.

10. The non-transitory computer readable storage medium as recited in claim 9 , wherein the representation in the video segment is modified using a faceswap process that comprises replacing the representation in the video segment with an image of the face of the user taken from another video.

11. The non-transitory computer readable storage medium as recited in claim 10 , wherein selecting video data that shows the emotion expressed by the user comprises selecting a pre-recorded emotion video that shows the emotion expressed by the user.

12. The non-transitory computer readable storage medium as recited in claim 11 , further comprising computing start and end indices for the pre-recorded emotion video, and the faceswap process is performed based in part on the start and end indices of the pre-recorded emotion video.

13. The non-transitory computer readable storage medium as recited in claim 9 , wherein the operations comprise presenting the modified video segment so that although the audio data in the modified video segment does not comprise real time audio data, the modified video segment, as presented, still appears to be a real time video of the user.

14. The non-transitory computer readable storage medium as recited in claim 9 , wherein the representation in the video segment is modified using a facewarp process based on an input data set comprising labeled sequences of facial movements.

15. The non-transitory computer readable storage medium as recited in claim 14 , wherein the facewarp process is performed based in part on face movement rules.

16. The non-transitory computer readable storage medium as recited in claim 14 , wherein the audio segment of the modified video segment is presented to a viewer in near real time as the audio segment is being recorded.

17. The method as recited in claim 1 , further comprising:

performing, in non-real time after the user has finished generating the audio data, a mouth prediction process in which a mouth of the facial features of the representation in the video segment is altered to appear to be speaking words which were spoken in the audio segment.

18. The non-transitory computer readable storage medium as recited in claim 9 , wherein the operations further comprises performing, in non-real time after the user has finished generating the audio data, a mouth prediction process in which a mouth of the facial features of the representation in the video segment is altered to appear to be speaking words which were spoken in the audio segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2022
From: ROTHSCHILD, SETH JACOB; ROBBINS, ALEX
To: EMC IP HOLDING COMPANY
Reel/Frame 059280/0783 →
Continuity (1)
Related Publication 20230061761A1 · Mar 2, 2023
References Cited (13)
US 6208373B1 · Fong et al. · 2001 [cited by applicant]
US 9204098B1 · Cunico et al. · 2015 [cited by applicant]
US 10440324B1 · Lichtenberg et al. · 2019 [cited by applicant]
US 10904488B1 · Weisz · 2021 [cited by examiner]
US 20100201780A1 · Bennett et al. · 2010 [cited by applicant]
US 20150381939A1 · Cunico · 2015 [cited by examiner]
US 20160148043A1 · Bathiche · 2016 [cited by examiner]
US 20160232941A1 · Cunico · 2016 [cited by examiner]
US 20180227339A1 · Rodriguez · 2018 [cited by examiner]
US 20230035306A1 · Liu · 2023 [cited by examiner]
US 20230121540A1 · Stratton · 2023 [cited by examiner]
Wikipedia: Deepfake, https://en.wikipedia.org/wiki/Deepfake, Accessed as early as Feb. 2021. [cited by applicant]
Faceswap website, https://faceswap.dev/, Accessed as early as Feb. 2021. [cited by applicant]