IP Library › Granted Patent US 12,361,936
Granted Patent B2
US 12,361,936 · App. 17/410,136 · Granted Jul 15, 2025

Method and system of automated question generation for speech assistance

Inventors: Robert Fernand Gordan (San Francisco, CA); Amit Srivastava (San Jose, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/22G06F40/40G10L15/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,936
App. No.
17/410,136
Granted
Jul 15, 2025
Kind
B2
Abstract

A method and system for generating one or more questions relating to a presentation session includes receiving audio data from the presentation session, retrieving a transcript for the audio data, receiving other data relating to the presentation session, providing at least one of the transcript and the other data to a machine-learning (ML) model as input for automatically generating the one or more questions relating to the presentation session, receiving from the ML model the one or more questions, and providing the one or more questions for display on a user interface associated with the presentation session.

Claims (55)

1. A data processing system comprising:

a processor; and

a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor or another element, cause the data processing system to perform functions of:

receiving audio data from a practice presentation session;

retrieving a transcript for the audio data;

receiving other data relating to the practice presentation session, the other data being extracted from content of presentation materials used to conduct the practice presentation session;

generating, at a questions engine, a prompt to be provided as an input to a natural language generation (NLG) model, by utilizing a prompt script that includes blank spaces for inserting data relating to at least one of the transcript and the other data, wherein the prompt script allows the NLG model to perform zero-shot adaptation in real time, and the NLG model includes a model that receives the prompt and generates one or more questions that are likely to be asked during a live presentation session based on the transcript of the practice presentation session and the content of the presentation materials;

providing the prompt to the NLG model as input for automatically generating the one or more questions;

receiving from the NLG model the one or more questions;

providing the one or more questions for display on a user interface associated with the practice presentation session; and

training the NLG model using a training data set of textual data.

2. The data processing system of claim 1 , wherein the other data includes as least one of content of a presentation document, multimodal content, content of a document shared during the presentation session or one or more questions from a previous presentation session.

3. The data processing system of claim 1 , wherein the practice presentation session is at least one of a speech rehearsal session or a virtual meeting.

4. The data processing system of claim 1 , wherein the audio data is received in real time while the practice presentation session is occurring.

5. The data processing system of claim 1 , wherein the one or more questions are displayed in real time during the practice presentation session.

6. The data processing system of claim 1 , wherein the user interface includes a UI element for receiving a response to the one or more questions from a user.

7. The data processing system of claim 6 , wherein the audio data includes data relating to the response and the executable instructions, when executed by the processor, further cause the data processing system to perform functions of:

providing at least one of the transcript and the other data to a second ML model as input for evaluating the response;

receiving from the second ML model one or more evaluation results for the response; and

providing the one or more evaluation results for display on the user interface associated with the practice presentation session.

8. The data processing system of claim 7 , wherein the evaluation results include at least one of a first indication of clarity of the response, and a second indication of a length of the response.

9. A method for generating one or more questions relating to a presentation session comprising:

receiving audio data from a practice presentation session;

retrieving a transcript for the audio data;

receiving other data relating to the practice presentation session, the other data being extracted from content of presentation materials used to conduct the practice presentation session;

generating, at a questions engine, a prompt to be provided as an input to a natural language generation (NLG) model, by utilizing a prompt script that includes blank spaces for inserting data relating to at least one of the transcript and the other data wherein the prompt script allows the NLG model to perform zero-shot adaptation in real time, and the NLG model includes a model that receives the prompt and generates one or more questions that are likely to be asked during a live presentation session based on the transcript of the practice presentation session and the content of the presentation materials;

providing the prompt to the NLG model as input for automatically generating the one or more questions;

receiving from the NLG model the one or more questions;

providing the one or more questions for display on a user interface associated with the practice presentation session; and

training the NLG model using a training data set of textual data.

10. The method of claim 9 , wherein the other data includes as least one of content of a presentation document, multimodal content, content of a document shared during the presentation session or one or more questions from a previous presentation session.

11. The method of claim 9 , wherein the practice presentation session is at least one of a speech rehearsal session or a virtual meeting.

12. The method of claim 9 , wherein the audio data is received in real time while the practice presentation session is occurring.

13. The method of claim 9 , wherein the one or more questions are displayed in real time during the practice presentation session.

14. The method of claim 9 , wherein the one or more questions are displayed after the practice presentation session is completed.

15. The method of claim 9 , wherein the user interface includes a UI element for receiving a response to the one or more questions from a user.

16. The method of claim 15 , wherein the audio data includes data relating to the response and the method further comprises:

providing at least one of the transcript and the other data to a second ML model as input for evaluating the response;

receiving from the second ML model one or more evaluation results for the response; and

providing the one or more evaluation results for display on the user interface associated with the practice presentation session.

17. A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to:

receive audio data from a practice presentation session;

retrieve a transcript for the audio data;

receiving other data relating to the practice presentation session, the other data being extracted from content of presentation materials used to conduct the practice presentation session;

generating, at a questions engine, a prompt to be provided as an input to a natural language generation (NLG) model, by utilizing a prompt script that includes blank spaces for inserting data relating to at least one of the transcript and the other data wherein the prompt script allows the NLG model to perform zero-shot adaptation in real time, and the NLG model includes a model that receives the prompt and generates one or more questions that are likely to be asked during a live presentation session based on the transcript of the practice presentation session and the content of the presentation materials;

providing the prompt to the NLG model as input for automatically generating the one or more questions;

receive from the NLG model the one or more questions;

provide the one or more questions for display on a user interface associated with the practice presentation session; and

training the NLG model using a training data set of textual data.

18. The non-transitory computer readable medium of claim 17 , wherein the other data includes as least one of content of a presentation document, multimodal content, content of a document shared during the presentation session or one or more questions from a previous presentation session.

19. The non-transitory computer readable medium of claim 17 , wherein the user interface includes a UI element for receiving a response to the one or more questions from a user.

20. The non-transitory computer readable medium of claim 19 , wherein the audio data includes data relating to the response and the instructions when executed, further cause a programmable device to:

provide at least one of the transcript and the other data to a second ML model as input for evaluating the response;

receive from the second ML model one or more evaluation results for the response; and

provide the one or more evaluation results for display on the user interface associated with the practice presentation session.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2021
From: GORDAN, ROBERT FERNAND; SRIVASTAVA, AMIT
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057270/0417 →
Continuity (1)
Related Publication 20230061210A1 · Mar 2, 2023
References Cited (67)
US 6411932B1 · Molnar · 2002 [cited by applicant]
US 7778831B2 · Chen · 2010 [cited by applicant]
US 8050922B2 · Chen · 2011 [cited by applicant]
US 8457967B2 · Audhkhasi et al. · 2013 [cited by applicant]
US 9747897B2 · Peng et al. · 2017 [cited by applicant]
US 9792908B1 · Bassemir et al. · 2017 [cited by applicant]
US 10069971B1 · Shaw et al. · 2018 [cited by applicant]
US 10560492B1 · Ledet · 2020 [cited by applicant]
US 11341331B2 · Liao et al. · 2022 [cited by applicant]
US 20020120447A1 · Charlesworth · 2002 [cited by applicant]
US 20090089062A1 · Lu · 2009 [cited by applicant]
US 20120322035A1 · Julia et al. · 2012 [cited by applicant]
US 20130231930A1 · Sanso · 2013 [cited by applicant]
US 20140356822A1 · Hoque · 2014 [cited by examiner]
US 20150310852A1 · Spizzo et al. · 2015 [cited by applicant]
US 20160049094A1 · Gupta et al. · 2016 [cited by applicant]
US 20160077719A1 · Threewits · 2016 [cited by applicant]
US 20160133155A1 · Lee et al. · 2016 [cited by applicant]
US 20160253999A1 · Kang et al. · 2016 [cited by applicant]
US 20180075145A1 · Zhao · 2018 [cited by examiner]
US 20180315420A1 · Ash et al. · 2018 [cited by applicant]
US 20190361842A1 · Wood et al. · 2019 [cited by applicant]
US 20190385480A1 · Suzuki · 2019 [cited by applicant]
US 20200135050A1 · Nunez · 2020 [cited by applicant]
US 20200184958A1 · Norouzi · 2020 [cited by applicant]
US 20200296457A1 · Church et al. · 2020 [cited by applicant]
US 20210065582A1 · Liao et al. · 2021 [cited by applicant]
US 20210099317A1 · Hilleli et al. · 2021 [cited by applicant]
US 20210103635A1 · Liao · 2021 [cited by examiner]
US 20210103851A1 · Spotanski · 2021 [cited by examiner]
US 20210118426A1 · Li · 2021 [cited by examiner]
US 20210151036A1 · Diment · 2021 [cited by applicant]
US 20210312399A1 · Asokan · 2021 [cited by examiner]
US 20210319786A1 · Kain et al. · 2021 [cited by applicant]
US 20220223066A1 · Chen · 2022 [cited by applicant]
US 20220230082A1 · Poddar · 2022 [cited by examiner]
US 20240013790A1 · Li · 2024 [cited by applicant]
CN 101551947A · 2009 [cited by applicant]
CN 107945788A · 2018 [cited by applicant]
CN 111081229A · 2020 [cited by applicant]
EP 0953970A2 · 1999 [cited by applicant]
JP 2016191739A · 2016 [cited by applicant]
KR 101672484B1 · 2016 [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 16/560,783”, Mailed Date: Feb. 17, 2023, 24 Pages. [cited by applicant]
“Final Office Action Issued in U.S. Appl. No. 16/560,783”, Mailed Date: Sep. 30, 2022, 20 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/037092”, Mailed Date: Nov. 11, 2022, 9 Pages. [cited by applicant]
“(How to) Pronounce”, Retrieved from: https://web.archive.org/web/20200320093418/apps.apple.com/us/app/how-to-pronounce/id717945069, Mar. 20, 2020, 3 Pages. [cited by applicant]
“ELSA: Learn And Speak English”, Retrieved from: https://apps.apple.com/us/app/elsa-learn-english-speech/id1083804886, Retrieved On: Apr. 18, 2020, 3 Pages. [cited by applicant]
“English Pronunciation IPA”, Retrieved from: https://apps.apple.com/us/app/english-pronunciation-ipa/id939357791, Retrieved On: Mar. 16, 2022, 3 Pages. [cited by applicant]
“Look Up: Pronunciation Checker & Dictionary”, Retrieved from: https://apps.apple.com/us/app/look-up-pronunciation-checker-dictionary/id1217022803, Retrieved On: Mar. 16, 2022, 3 Pages. [cited by applicant]
“Say It: English Pronunciation”, Retrieved from: https://web.archive.org/web/20190828012815/https://apps.apple.com/us/app/say-it-english-pronunciation/id919978521, Aug. 28, 2019, 3 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/CN21/096621”, Mailed Date: Feb. 28, 2022, 10 Pages. [cited by applicant]
Pearce, James, “English Vocabulary: How to Speak with Fluency Like a Native”, Retrieved from: https://web.archive.org/web/20210501023209/https://www.fluentu.com/blog/english/, May 1, 2021, 6 pages. [cited by applicant]
Audhkhasi, et al., “Formant-Based Technique for Automatic Filled-Pause Detection in Spontaneous Spoken English”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2009,… [cited by applicant]
“Applicant Initiated Interview Summary Issued in U.S. Appl. No. 16/560,783”, Mailed Date: Jul. 26, 2022, 3 Pages. [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 16/560,783”, Mailed Date: May 5, 2022, 22 Pages. [cited by applicant]
Kaushik, et al., “Laughter and filler detection in naturalistic audio”, In Proceedings of 16th Annual Conference of the International Speech Communication Association, Sep. 6, 2015, pp. 2509-2513. [cited by applicant]
Kurihara, et al., “Presentation sensei: a presentation training system using speech and image processing”, In Proceedings of the 9th international conference on Multi modal interfaces, Nov. 12, 2007, pp. 358-365. [cited by applicant]
Shangavi, et al., “Self-Speech Evaluation with Speech Recognition and Gesture Analysis”, In Proceedings of National Information Technology Conference, Oct. 2, 2018, 7 Pages. [cited by applicant]
“GPT-3”, Retrieved from: https://en.wikipedia.org/wiki/GPT-3, Jun. 28, 2021, 6 Pages. [cited by applicant]
U.S. Appl. No. 17/415,675, filed May 28, 2021. [cited by applicant]
Communication under Rule 70(2) and 70a(2) Received for European Application No. 21942358.9, mailed on Feb. 11, 2025, 01 pages. [cited by applicant]
Final Office Action Issued in U.S. Appl. No. 16/593,724, Mailed Date : Oct. 21, 2021 9 Pages. [cited by applicant]
Non Final Office Action Issued in U.S. Appl. No. 16/593,724, Mailed Date: Feb. 3, 2021 23 Pages. [cited by applicant]
Extended European Search Report Received in European Patent Application No. 21942358.9, mailed on Jan. 22, 2025, 09 pages. [cited by applicant]
Non-Final Office Action mailed on Mar. 18, 2025, in U.S. Appl. No. 17/415,675, 31 Pages. [cited by applicant]
US-2021-0065582-A1, Mar. 4, 2021. [cited by applicant]