IP Library Granted Patent US 12,200,322
Granted Patent B2
US 12,200,322 · App. 18/389,764 · Granted Jan 14, 2025

Systems and methods for generating a video summary of a virtual event

Inventors: Subham Biswas (Thane, IN); Saurabh Tahiliani (Noida, IN)
Assignee: Verizon Patent and Licensing Inc.
H04N21/8549G10L15/02G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,200,322
App. No.
18/389,764
Granted
Jan 14, 2025
Kind
B2
Abstract

A video summary device may generate a textual summary of a transcription of a virtual event. The video summary device may generate a phonemic transcription of the textual summary and generate a text embedding based on the phonemic transcription. The video summary device may generate an audio embedding based on a target voice. The video summary device may generate an audio output of the phonemic transcription uttered by the target voice. The audio output may be generated based on the text embedding and the audio embedding. The video summary device may generate an image embedding based on video data of a target user. The image embedding may include information regarding images of facial movements of the target user. The video summary device may generate a video output of different facial movements of the target user uttering the phonemic transcription, based on the text embedding and the image embedding.

Claims (73)

1. A method performed by a video summary device, the method comprising:

generating an audio embedding based on a target voice, wherein the audio embedding includes information regarding audio classification of the target voice;

generating an audio output of a phonemic transcription, of an event, uttered by the target voice, wherein the audio output is generated based on a text embedding and the audio embedding, and wherein the text embedding includes information regarding text classification of the phonemic transcription;

generating an image embedding based on video data of a target user, wherein the image embedding includes information regarding images of facial movements of the target user;

generating a video output of different facial movements of the target user uttering the phonemic transcription, wherein the video output is generated based on the text embedding and the image embedding; and

generating a video summary of the event based on the audio output and the video output.

2. The method of claim 1 , further comprising:

generating the text embedding using a machine learning model trained for text classification.

3. The method of claim 2 , wherein generating the text embedding comprises:

providing the phonemic transcription as an input to the machine learning model,

wherein the machine learning model generates the text embedding as an output.

4. The method of claim 1 , further comprising:

identifying, as the target voice, a voice of a participant of the event; or

identifying, as the target voice, a voice of a user that was not a participant of the event.

5. The method of claim 1 , wherein generating the audio embedding comprises:

providing voice samples of the target voice as inputs to a machine learning model trained for audio classification,

wherein the machine learning model generates the audio embedding based on the voice samples.

6. The method of claim 1 , wherein generating the audio output comprises:

providing the text embedding and the audio embedding as inputs to a machine learning model,

wherein the machine learning model generates a spectrogram based on the text embedding and the audio embedding; and

generating the audio output based on the spectrogram.

7. The method of claim 6 , wherein generating the audio output comprises:

generating a waveform based on the spectrogram,

wherein the audio output includes the waveform.

8. A device, comprising:

one or more processors configured to:

generate an audio embedding based on a target voice, wherein the audio embedding includes information regarding audio classification of the target voice;

generate an audio output of a phonemic transcription, of an event, being uttered by the target voice, wherein the audio output is generated based on a text embedding and the audio embedding, and wherein the text embedding includes information regarding text classification of the phonemic transcription;

generate an image embedding based on video data of a target user, wherein the image embedding includes information regarding images of facial movements of the target user;

generate a video output of different facial movements of the target user uttering the phonemic transcription, wherein the video output is generated based on the text embedding and the image embedding; and

generate a video summary of the event based on the audio output and the video output.

9. The device of claim 8 , wherein the one or more processors, to generate the image embedding, are configured to:

identify the target user based on the target voice,

wherein the video data includes video data of facial movements of the target user as the target user utters at least one of one or more words or one or more phrases.

10. The device of claim 8 , wherein the one or more processors, to generate the image embedding, are configured to:

provide the video data as an input to a machine learning model trained for image classification,

wherein the machine learning model generates the image embedding based on the video data.

11. The device of claim 10 , wherein the one or more processors are further configured to:

process a transcription of the event to generate a processed input; and

provide the processed input as an input to a language model,

wherein the phonemic transcription is generated by the language model based on the processed input.

12. The device of claim 8 , wherein the one or more processors, to generate the video output, are configured to:

generate a first plurality of images depicting the target user uttering a first portion of a plurality of portions of the phonemic transcription; and

generate a second plurality of images depicting the target user uttering a second portion of the plurality of portions,

wherein the video output includes the first plurality of images and the second plurality of images.

13. The device of claim 12 , wherein the one or more processors, to generate the audio output, are configured to:

generate a first image of the first plurality of images; and

generate a second image of the first plurality of images,

wherein the second image is generated based on the first image, the text embedding, and the image embedding.

14. The device of claim 13 , wherein the one or more processors, to generate the second image, are configured to:

modify one or more pixel values of the first image to generate the second image.

15. A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

generate an audio output of a phonemic transcription being uttered by a target voice, wherein the audio output is generated based on a text embedding and an audio embedding, wherein the text embedding includes information regarding text classification of the phonemic transcription, and wherein the audio embedding includes information regarding audio classification of the target voice;

generate a video output of different facial movements of a target user uttering the phonemic transcription, wherein the video output is generated based on the text embedding and an image embedding generated based on video data of the target user, and wherein the image embedding includes information regarding images of facial movements of the target user;

generate a video summary based on the audio output and the video output; and

provide the video summary.

16. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions further cause the device to:

generate the text embedding using a machine learning model trained for text classification,

wherein the phonemic transcription is provided as an input to the machine learning model.

17. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions further cause the device to:

generate the audio embedding using a machine learning model trained for audio classification,

wherein voice samples, of the target voice, are provided as inputs to the machine learning model.

18. The non-transitory computer-readable medium of claim 17 , wherein the one or more instructions further cause the device to:

generate the image embedding using a machine learning model trained for image classification,

wherein the video data is provided as an input to the machine learning model.

19. The non-transitory computer-readable medium of claim 18 , wherein the one or more instructions, that cause the device to generate the image embedding, cause the device to:

identify the target user based on the target voice,

wherein the video data includes video data of facial movements of the target user as the target user utters at least one of one or more words or one or more phrases.

20. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to generate the video output, cause the device to:

generate a first plurality of images depicting the target user uttering a first portion of a plurality of portions of the phonemic transcription; and

generate a second plurality of images depicting the target user uttering a second portion of the plurality of portions,

wherein the video output includes the first plurality of images and the second plurality of images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2023
From: BISWAS, SUBHAM; TAHILIANI, SAURABH
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 066088/0021 →
Continuity (2)
Continuation 17811732 · Jul 11, 2022
Related Publication 20240121487A1 · Apr 11, 2024
References Cited (23)
US 8209623B2 · Barletta et al. · 2012 [cited by applicant]
US 8296797B2 · Olstad · 2012 [cited by examiner]
US 9635307B1 · Mysore Vijaya Kumar · 2017 [cited by examiner]
US 10096033B2 · Heath · 2018 [cited by applicant]
US 10582149B1 · Mysore Vijaya Kumar · 2020 [cited by examiner]
US 11334618B1 · Adlersberg et al. · 2022 [cited by applicant]
US 11605402B2 · Xu · 2023 [cited by examiner]
US 11616658B2 · Hu · 2023 [cited by examiner]
US 20120284276A1 · Fernando et al. · 2012 [cited by applicant]
US 20160014482A1 · Chen et al. · 2016 [cited by applicant]
US 20210081056A1 · Divakaran et al. · 2021 [cited by applicant]
US 20220067385A1 · Kaushik · 2022 [cited by examiner]
US 20230140369A1 · Aminian · 2023 [cited by examiner]
US 20230169990A1 · Biswas et al. · 2023 [cited by applicant]
US 20230283851A1 · Stathacopoulos · 2023 [cited by applicant]
US 20240020337A1 · Maharana · 2024 [cited by examiner]
US 20240037824A1 · Biswas · 2024 [cited by examiner]
US 20240095449A1 · Ranganathan · 2024 [cited by examiner]
US 20240223872A1 · Sharma · 2024 [cited by examiner]
“ToPhonetics”, Accessed on Jul. 8, 2022, 1 page. [Retrieved from https://tophonetics.com/]. [cited by applicant]
Bernard, et al., “Phonemizer”, GitHub, bootphon, 2021, 3 pages. [Retrieved from https://github.com/bootphon/phonemizer]. [cited by applicant]
Jurafsky, et al., “Switchboard SWBD-DAMSL Shallow-Discourse-Function Annotation Coders Manual, Draft 13”, University of Colorado at Boulder & +SRI International, Aug. 1, 1997, 32 pages. Retrieved from https://web.stanfo… [cited by applicant]
Mortensen, “epitran 1.22”, Python Package Index (PyPI), Jun. 11, 2022, 24 pages. [Retrieved from https://pypi.org/project/epitran/]. [cited by applicant]