IP Library › Granted Patent US 12,512,102
Granted Patent B2
US 12,512,102 · App. 17/926,291 · Granted Dec 30, 2025

Providing prompts in speech recognition results in real time

Inventor: Kun Wu (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L17/22G10L17/14G10L17/18G06F3/0237
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,512,102
App. No.
17/926,291
Granted
Dec 30, 2025
Kind
B2
Abstract

The present disclosure provides methods and apparatuses for providing prompts in speech recognition results in real time. A current speech input in an audio stream for a target event may be obtained. A current utterance text corresponding to the current speech input may be identified. A prompt may be generated based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence. A speech recognition result for the current speech input may be provided, the speech recognition result comprising the current utterance text and the prompt.

Claims (60)

1 . A method for providing prompts in speech recognition results in real time, comprising:

obtaining a current speech input in an audio stream for a target event;

identifying a current utterance text corresponding to the current speech input;

evaluating a corpus of documents associated with the target event using text analysis to generate a relevance matrix for the corpus of documents;

training a prompt generation recurrent neural network model using training data comprising a plurality of word sequences extracted from a plurality of documents associated with a plurality of target events to generate a prompt based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence comprising words selected from the relevance matrix based on respective absolute values among the words, wherein an order of the words of the at least one predicted subsequent utterance text sequence is determined using raw values among the words; and

providing a speech recognition result for the current speech input, the speech recognition result comprising the current utterance text and the prompt.

2 . The method of claim 1 , wherein

the current utterance text comprises one or more words, and

each predicted subsequent utterance text sequence comprises one or more predicted subsequent utterance texts, and each predicted subsequent utterance text comprises one or more words.

3 . The method of claim 1 , wherein

the prompt is generated further based on at least one previous utterance text identified before the current utterance text.

4 . The method of claim 1 , further comprising:

obtaining an event identity (ID) of the target event and/or a speaker ID of a speaker of the current speech input,

wherein the prompt is generated further based on the event ID and/or the speaker ID.

5 . The method of claim 1 , wherein the generating a prompt comprises:

predicting, through a previously-established prompt generator, the at least one predicted subsequent utterance text sequence based at least on the current utterance text.

6 . The method of claim 5 , wherein

the prompt generator is based on a relevance matrix.

7 . The method of claim 6 , wherein

the relevance matrix comprises a plurality of text items and relevance values among the plurality of text items, and

the predicting the at least one predicted subsequent utterance text sequence comprises:

selecting, based at least on the current utterance text, at least one text item sequence highest ranked by relevance from the plurality of text items.

8 . The method of claim 5 , further comprising:

obtaining at least one document associated with the target event; and

establishing the prompt generator based at least on the at least one document.

9 . The method of claim 8 , further comprising:

obtaining event information associated with the target event, the event information comprising at least one of: an event identity (ID) of the target event, time information of the target event, agenda of the target event, participant IDs of participants of the target event, and a speaker ID of a speaker associated with the at least one document.

10 . The method of claim 5 , further comprising:

obtaining a plurality of documents associated with a plurality of events and/or a plurality of speakers; and

establishing the prompt generator based at least on the plurality of documents.

11 . The method of claim 10 , further comprising:

obtaining event information associated with the plurality of events respectively and/or a plurality of speaker identities (IDs) corresponding to the plurality of speakers respectively,

wherein the prompt generator is established further based on the event information and/or the plurality of speaker IDs.

12 . The method of claim 1 , further comprising:

obtaining a second speech input subsequent to the current speech input in the audio stream;

identifying a second utterance text corresponding to the second speech input;

comparing the second utterance text with the at least one predicted subsequent utterance text sequence;

generating a second prompt based at least on the second utterance text and/or a result of the comparison; and

providing a speech recognition result for the second speech input, the speech recognition result comprising the second utterance text and the second prompt.

13 . An apparatus for providing prompts in speech recognition results in real time, comprising:

at least one processor; and

a memory storing computer-executable instructions that, when executed, cause the at least one processor to:

obtain a current speech input in an audio stream for a target event,

identify a current utterance text corresponding to the current speech input,

evaluating a corpus of documents associated with the target event using text analysis to generate a relevance matrix for the corpus of documents,

train a prompt generation recurrent neural network model using training data comprising a plurality of word sequences extracted from a plurality of documents associated with a plurality of target events to generate a prompt based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence comprising words selected from the relevance matrix based on respective absolute values among the words, wherein an order of the words of the at least one predicted subsequent utterance text sequence is determined using raw values among the words, and

provide a speech recognition result for the current speech input, the speech recognition result comprising the current utterance text and the prompt.

14 . The apparatus of claim 13 , the instructions to generate a prompt further comprising instructions to:

predict, through a previously-established prompt generator, the at least one predicted subsequent utterance text sequence based at least on the current utterance text.

15 . The apparatus of claim 14 , wherein the prompt generator is based on a relevance matrix.

16 . The apparatus of claim 15 , wherein the relevance matrix comprises a plurality of text items and relevance values among the plurality of text items, and the instructions to predict the at least one predicted subsequent utterance text sequence further comprising instructions to:

select, based at least on the current utterance text, at least one text item sequence highest ranked by relevance from the plurality of text items.

17 . At least one non-transitory machine-readable medium comprising instructions for providing prompts in speech recognition results in real time that, when executed by at least one processor, cause the at least one processor to:

obtain a current speech input in an audio stream for a target event;

identify a current utterance text corresponding to the current speech input;

evaluating a corpus of documents associated with the target event using text analysis to generate a relevance matrix for the corpus of documents;

train a prompt generation recurrent neural network model using training data comprising a plurality of word sequences extracted from a plurality of documents associated with a plurality of target events to generate a prompt based at least on the current utterance text, the prompt comprising at least one predicted subsequent utterance text sequence comprising words selected from the relevance matrix based on respective absolute values among the words, wherein an order of the words of the at least one predicted subsequent utterance text sequence is determined using raw values among the words; and

provide a speech recognition result for the current speech input, the speech recognition result comprising the current utterance text and the prompt.

18 . The at least one non-transitory machine-readable medium of claim 17 , the instructions to generate a prompt further comprising instructions to:

predict, through a previously-established prompt generator, the at least one predicted subsequent utterance text sequence based at least on the current utterance text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2022
From: WU, KUN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061975/0777 →
Priority Claims (1)
CN 202010517639.2 · Jun 4, 2020 · national
Continuity (1)
Related Publication 20230215441A1 · Jul 6, 2023
References Cited (36)
US 7562288B2 · Dittrich · 2009 [cited by applicant]
US 9626955B2 · Fleizach et al. · 2017 [cited by applicant]
US 10366149B2 · Wolfram et al. · 2019 [cited by applicant]
US 11341331B2 · Liao et al. · 2022 [cited by applicant]
US 20080072143A1 · Assadollahi · 2008 [cited by examiner]
US 20100185447A1 · Krumel · 2010 [cited by applicant]
US 20110197128A1 · Assadollahi · 2011 [cited by examiner]
US 20120173222A1 · Wang · 2012 [cited by examiner]
US 20140297267A1 · Spencer · 2014 [cited by examiner]
US 20150089435A1 · Kuzmin · 2015 [cited by examiner]
US 20160275070A1 · Corston · 2016 [cited by examiner]
US 20160328377A1 · Spencer · 2016 [cited by examiner]
US 20160364087A1 · Thompson et al. · 2016 [cited by applicant]
US 20180150143A1 · Orr · 2018 [cited by examiner]
US 20200019308A1 · Medlock · 2020 [cited by examiner]
US 20200098366A1 · Chakraborty · 2020 [cited by applicant]
US 20210089860A1 · Heere · 2021 [cited by examiner]
CN 112395412A · 2021 [cited by examiner]
EP 1160767A2 · 2001 [cited by applicant]
EP 3657501A1 · 2020 [cited by applicant]
WO WO2008064137A3 · 2007 [cited by examiner]
WO 2017112813A1 · 2017 [cited by applicant]
James, et al., “Text Input for Mobile Devices: Comparing Model Prediction to Actual Performance,” CHI 2001. (Year: 2001). [cited by examiner]
James, et al., “Text Input for Mobile Devices: Comparing Model Prediction to Actual Performance,” CHI 2001. (Year: 2001)—see attached reference in the previous Office action. (Year: 2001). [cited by examiner]
James, et al., “Text Input for Mobile Devices: Comparing Model Prediction to Actual Performance,” CHI 2001—see attached reference in the 1st Office action. (Year: 2001). [cited by examiner]
“Ali Linger”, Retrieved From: https://ai.aliyun.com/, Mar. 18, 2020, 3 Pages. [cited by applicant]
Yoshioka, et al., “Advances in Online Audio-Visual Meeting Transcription”, In Repository of arXiv:1912.04979v1, Dec. 10, 2019, 8 Pages. [cited by applicant]
“Speech Service Documentation”, Retrieved From: https://learn.microsoft.com/en-us/azure/cognitive-services/Speech-Service/, May 16, 2018, 3 Pages. [cited by applicant]
Asadi, et al., “IntelliPrompter: Speech-based Dynamic Note Display Interface for Oral Presentations”, In Proceedings of the 19th ACM International Conference on Multimodal Interaction, Nov. 13, 2017, pp. 172-180. [cited by applicant]
Fasbinder, Fia, “5 Public Speaking Apps to Perfect Your Next Presentation”, Retrieved From: https://www.inc.com/fia-fasbinder/need-to-nail-your-next-speech-theres-an-app-for-th.html, Jul. 24, 2017, 6 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US2021/028277”, Mailed Date: Aug. 4, 2021, 13 Pages. [cited by applicant]
Office Action Received for Chinese Application No. 202010517639.2, mailed on Feb. 27, 2024, 19 pages (English Translation Provided). [cited by applicant]
Communication under Rule 71(3) received in European Application No. 21724481.3 mailed on Aug. 30, 2024, 8 pages. [cited by applicant]
Decision to Grant pursuant to Article 97(1) received in European Application No. 21724481.3, mailed on Jan. 8, 2025, 2 pages. [cited by applicant]
Notice of Allowance Received for Chinese Application No. 202010517639.2, mailed on Mar. 25, 2025, 08 pages (English Translation Provided). [cited by applicant]
Second Office Action Received for Chinese Application No. 202010517639.2, mailed on Dec. 18, 2024, 23 pages (English Translation Provided). [cited by applicant]