IP Library › Granted Patent US 12,032,922
Granted Patent B2
US 12,032,922 · App. 17/318,170 · Granted Jul 9, 2024

Automated script generation and audio-visual presentations

Inventors: Ji Li (San Jose, CA); Konstantin Seleskerov (Palo Alto, CA); Huey-Ru Tsai (Los Altos, CA); Muin Barkatali Momin (Santa Clara, CA); Ramya Tridandapani (Sunnyvale, CA); Sindhu Vigasini Jambunathan (San Jose, CA); Amit Srivastava (San Jose, CA); Derek Martin Johnson (Sunnyvale, CA); Gencheng Wu (Campbell, CA); Sheng Zhao (Beijing, CN); Xinfeng Chen (Beijing, CN); Bohan Li (Beijing, CN)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F40/58G06F3/0481G06F16/24578G06F40/205G10L13/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,032,922
App. No.
17/318,170
Filed
May 12, 2021
Granted
Jul 9, 2024
Kind
B2
Art Unit
2655
USPC
704/3
Abstract

Automatic generation of intelligent content is created using a system of computers including a user device and a cloud-based component that processes the user information. The system performs a process that includes receiving an input document and parsing the input document to generate inputs for a natural language generation model using a text analysis model. The natural language generation model generates one or more candidate presentation scripts based on the inputs. A presentation script is selected from the candidate presentation scripts and displayed. A text-to-speech model may be used to generate a synthesized audio presentation of the presentation script. A final presentation may be generated that includes a visual display of the input document and the corresponding audio presentation in sync with the visual display.

Claims (65)

1. A computer-implemented method for automatically generating a presentation script, the method comprising:

receiving a presentation slide deck comprising a visual presentation;

determining a topic of the presentation slide deck;

generating a natural language generation model prompt for a first natural language generation model of a plurality of natural language generation models using an input design model, the generating comprising:

providing the presentation slide deck, the topic, and user information as first inputs to the input design model trained to:

select the first natural language generation model from the plurality of natural language generation models based on the first inputs, and

generate the natural language generation model prompt for the first natural language generation model based on the first inputs, wherein the natural language generation model prompt is designed to request a candidate presentation script for a speech on the topic, and

receiving the natural language generation model prompt as output from the input design model in response to the first inputs;

providing the natural language generation model prompt as an input prompt to the first natural language generation model trained to generate a natural language response to the input prompt;

receiving, in response to the input prompt, the candidate presentation script for the speech on the topic from the first natural language generation model;

generating an audio presentation by inputting the candidate presentation script into a text-to-speech model;

generating a final presentation comprising a visual display of the presentation slide deck synchronized with the audio presentation; and

outputting the final presentation to an audience.

2. The computer-implemented method of claim 1 , further comprising:

displaying one or more candidate presentation scripts comprising the candidate presentation script; and

receiving a selection of the candidate presentation script from the displayed one or more candidate presentation scripts.

3. The computer-implemented method of claim 1 , further comprising:

parsing the presentation slide deck to determine the topic.

4. The computer-implemented method of claim 1 , further comprising:

generating one or more candidate presentation scripts comprising the candidate presentation script;

ranking the one or more candidate presentation scripts with a ranking model; and

displaying the one or more candidate presentation scripts in ranked order.

5. The computer-implemented method of claim 4 , wherein each of the plurality of natural language generation models generates at least one of the one or more candidate presentation scripts.

6. The computer-implemented method of claim 1 , further comprising:

inputting a user voice into the text-to-speech model, wherein the audio presentation is generated using the user voice.

7. The computer-implemented method of claim 1 , further comprising:

receiving a request to modify an output language of the audio presentation in the final presentation to a requested language; and

translating the output language to the requested language in the final presentation.

8. The computer-implemented method of claim 1 , further comprising:

receiving feedback from the audience after outputting the final presentation to the audience; and

adjusting parameters of the input design model, the natural language generation model, or a combination of both based on the feedback.

9. A system comprising:

one or more processors; and

a memory having stored thereon instructions that, upon execution by the one or more processors, cause the one or more processors to:

receive a presentation slide deck comprising a visual presentation;

determine a topic of the presentation slide deck;

generate a natural language generation model prompt for a first natural language generation model of a plurality of natural language generation models using an input design model, the generating comprising:

provide the presentation slide deck, the topic, and user information as first inputs to the input design model trained to:

select the first natural language generation model from the plurality of natural language generation models based on the first inputs, and

generate the natural language generation model prompt for the first natural language generation model based on the first inputs,

wherein the natural language generation model prompt is designed to request a candidate presentation script for a speech on the topic, and

receive the natural language generation model prompt as output from the input design model in response to the first inputs;

provide the natural language generation model prompt as an input prompt to the first natural language generation model trained to generate a natural language response to the input prompt;

receive, in response to the input prompt, the candidate presentation script for the speech on the topic from the first natural language generation model;

generate an audio presentation by inputting the candidate presentation script into a text-to-speech model;

generate a final presentation comprising a visual display of the presentation slide deck synchronized with the audio presentation; and

output the final presentation to an audience.

10. The system of claim 9 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

display one or more candidate presentation scripts comprising the candidate presentation script; and

receive a selection of the candidate presentation script from the displayed one or more candidate presentation scripts.

11. The system of claim 9 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

parse the presentation slide deck to determine the topic.

12. The system of claim 9 , the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

generate one or more candidate presentation scripts comprising the candidate presentation script;

rank the one or more candidate presentation scripts with a ranking model; and

display the one or more candidate presentation scripts in ranked order.

13. The system of claim 12 , wherein each of the plurality of natural language generation models generates at least one of the one or more candidate presentation scripts.

14. The system of claim 9 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

input a user voice into the text-to-speech model, wherein the audio presentation is generated using the user voice.

15. The system of claim 9 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

receive a request to modify an output language of the audio presentation in the final presentation to a requested language; and

translate the output language to the requested language in the final presentation.

16. The system of claim 9 , wherein the instructions comprise further instructions that, upon execution by the one or more processors, cause the one or more processors to:

receive feedback from the audience after outputting the final presentation to the audience; and

adjust parameters of the input design model, the natural language generation model, or a combination of both based on the feedback.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2021
From: LI, JI; SELESKEROV, KONSTANTIN; TSAI, HUEY-RU; MOMIN, MUIN BARKATALI; TRIDANDAPANI, RAMYA; JAMBUNATHAN, SINDHU VIGASINI; SRIVASTAVA, AMIT; JOHNSON, DEREK MARTIN; WU, GENCHENG; ZHAO, SHENG; CHEN, XINFENG; LI, BOHAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056230/0319 →
Continuity (1)
Related Publication 20220366153A1 · Nov 17, 2022
Cited By (6)
US 12,493,740 US 12,541,638 US 12,566,916 US 12,632,669 US 12,718,033 US 12,724,978