IP Library › Granted Patent US 12,321,383
Granted Patent B2
US 12,321,383 · App. 18/408,792 · Granted Jun 3, 2025

Suggested queries for transcript search

Inventors: Adi Miller (Ramat Hasharon, IL); Haim Somech (Herzliya, IL); Michael Sterenberg (Herzliya, IL)
Assignee: Microsoft Technology Licensing, LLC
G06F16/685G06F3/0481G06F3/0484G06F16/3329G06F40/295G06F40/30G10L17/22G10L25/57H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,383
App. No.
18/408,792
Filed
Jan 10, 2024
Granted
Jun 3, 2025
Kind
B2
Art Unit
2692
USPC
704/235
Abstract

Systems and methods for surfacing natural language queries from one or more transcripts. An example method may include converting received audio to text, through automated speech recognition, to form a transcript of the audio, wherein the transcript includes text of the audio and identifications of speakers associated with portions of the text corresponding to utterances from the respective speakers; generating input signals based on at least the transcript; executing at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate at least one of a suggested natural language query for searching the transcript or a key moment within the received audio; and causing at least one of the suggested natural language query or the key moment to be surfaced on one or more remote devices.

Claims (52)

1. A system for surfacing natural language queries, the system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

generate input signals for content of a virtual meeting;

execute at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate a suggested natural language query for searching the content of the virtual meeting;

cause the suggested natural language query to be surfaced on one or more remote devices;

receive a selection of the suggested natural language query;

based on receiving the selection of the suggested natural language query, execute the suggested natural language query against the contents of the meeting; and

generate query results.

2. The system of claim 1 , wherein the input signals include at least one of: durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting.

3. The system of claim 1 , wherein the input signals are generated for a transcript of the virtual meeting.

4. The system of claim 3 , wherein the input signals include at least one of the text of the transcript, audio corresponding to the text of the transcript, speaker identification for utterances in the transcript.

5. The system of claim 1 , wherein the operations further comprise:

analyze the input signals identify a topic in the virtual meeting; and

wherein the suggested natural language query includes the identified topic.

6. The system of claim 1 , wherein:

a chat entry is received during the virtual meeting;

the chat entry receives a count of reactions greater than a reaction-count threshold; and

the suggested natural language query includes question regarding key moments of the virtual meeting.

7. The system of claim 1 , wherein the operations further comprise:

performing a semantic analysis of content to identify an uttered name; and

the suggested natural language query is based on the uttered name.

8. The system of claim 1 , wherein the suggested natural language query is surfaced during the virtual meeting.

9. A computer-implemented method for surfacing suggested natural language queries, the method comprising:

generate input signals for content of a virtual meeting;

performing a semantic analysis of one or more of the input signals to identify an utterance, by a first participant, of a name of a second participant;

executing at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate at least one of a suggested natural language query for searching the transcript or a key moment within the received audio; and

causing at least one of the suggested natural language query or the key moment to be surfaced on one or more remote devices, wherein at least one of the suggested natural language query or the key moment is based on the uttered name of the second participant and the at least one of the suggested natural language query or the key moment is caused to be surfaced on the remote device of the second participant.

10. The method of claim 9 , wherein the input signals include at least one of: text of a generated transcript, audio of the virtual meeting, speaker identification for utterances in the virtual meeting, durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting.

11. The method of claim 9 , wherein the natural language query and the key moment are surfaced during the virtual meeting.

12. The method of claim 11 , wherein the natural language query and the key moment are concurrently surfaced in a videoconference interface during a videoconference.

13. The method of claim 12 , wherein the videoconference interface includes a plurality of video segments, a transcript-search input field, and at least a portion of the transcript.

14. The method of claim 13 , further comprising:

receiving a selection of the surfaced key moment; and

in response to receiving the selection, navigating the video segments to a time point at which the key moment occurred.

15. A computer-implemented method for surfacing natural language queries, the method comprising:

generating input signals for content of a virtual meeting;

executing at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate a suggested natural language query for searching the content of the virtual meeting;

causing the suggested natural language query to be surfaced on one or more remote devices;

receiving a selection of the suggested natural language query;

based on receiving the selection of the suggested natural language query, executing the suggested natural language query against the contents of the meeting; and

generating query results.

16. The method of claim 15 , wherein the input signals include at least one of: durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting.

17. The method of claim 15 , wherein the input signals are generated for a transcript of the virtual meeting.

18. The method of claim 17 , wherein the input signals include at least one of the text of the transcript, audio corresponding to the text of the transcript, speaker identification for utterances in the transcript.

19. The method of claim 15 , further comprising:

analyzing the input signals identify a topic in the virtual meeting; and

wherein the suggested natural language query includes the identified topic.

20. The method of claim 15 , wherein:

a chat entry is received during the virtual meeting;

the chat entry receives a count of reactions greater than a reaction-count threshold; and

the suggested natural language query includes question regarding key moments of the virtual meeting.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2025
From: MILLER, ADI; SOMECH, HAIM; STERENBERG, MICHAEL
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070135/0901 →
Continuity (2)
Continuation 17498296 · Oct 11, 2021
Related Publication 20240273139A1 · Aug 15, 2024
References Cited (31)
US 9064006B2 · Hakkani-Tur · 2015 [cited by examiner]
US 9742912B2 · Srivastava · 2017 [cited by examiner]
US 10546001B1 · Nguyen · 2020 [cited by examiner]
US 11500865B1 · Wang · 2022 [cited by examiner]
US 11604794B1 · Nallapati · 2023 [cited by examiner]
US 11914644B2 · Miller · 2024 [cited by examiner]
US 12197477B2 · Schwartz · 2025 [cited by examiner]
US 20090012778A1 · Feng · 2009 [cited by examiner]
US 20140059030A1 · Hakkani-Tur · 2014 [cited by examiner]
US 20170083633A1 · Bosko · 2017 [cited by applicant]
US 20170212895A1 · Ahmed · 2017 [cited by examiner]
US 20180121500A1 · Reschke · 2018 [cited by examiner]
US 20180129704A1 · Chandrasekaran · 2018 [cited by examiner]
US 20180129733A1 · Chandrasekaran · 2018 [cited by examiner]
US 20180144064A1 · Krasadakis · 2018 [cited by examiner]
US 20180203924A1 · Agrawal · 2018 [cited by examiner]
US 20190130026A1 · Ventura · 2019 [cited by examiner]
US 20190317994A1 · Singh · 2019 [cited by examiner]
US 20190340172A1 · McElvain · 2019 [cited by examiner]
US 20200125602A1 · Sezgin · 2020 [cited by examiner]
US 20200387551A1 · Hays · 2020 [cited by examiner]
US 20210149901A1 · Fonseca De Lima · 2021 [cited by examiner]
US 20210192134A1 · Yue · 2021 [cited by examiner]
US 20210240750A1 · Schwartz · 2021 [cited by examiner]
US 20220004571A1 · Ganapathy · 2022 [cited by examiner]
US 20220121685A1 · Zheng · 2022 [cited by examiner]
US 20220180056A1 · Hong · 2022 [cited by examiner]
US 20220229991A1 · Duong · 2022 [cited by examiner]
US 20230115098A1 · Miller · 2023 [cited by examiner]
US 20240273139A1 · Miller · 2024 [cited by examiner]
Communication pursuant to Article 94(3) Received in European Patent Application No. 22758363.0, mailed on Apr. 7, 2025, 11 pages. [cited by applicant]
Cited By (1)
US 12,717,849