IP Library › Granted Patent US 11,914,644
Granted Patent B2
US 11,914,644 · App. 17/498,296 · Granted Feb 27, 2024

Suggested queries for transcript search

Inventors: Adi Miller (Ramat Hasharon, IL); Haim Somech (Herzliya, IL); Michael Sterenberg (Herzliya, IL)
Assignee: Microsoft Technology Licensing, LLC
G06F16/685G06F3/0481G06F3/0484G06F16/3329G06F40/295G06F40/30G10L17/22G10L25/57H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,644
App. No.
17/498,296
Filed
Oct 11, 2021
Granted
Feb 27, 2024
Kind
B2
Art Unit
2692
USPC
704/235
Abstract

Systems and methods for surfacing natural language queries from one or more transcripts. An example method may include converting received audio to text, through automated speech recognition, to form a transcript of the audio, wherein the transcript includes text of the audio and identifications of speakers associated with portions of the text corresponding to utterances from the respective speakers; generating input signals based on at least the transcript; executing at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate at least one of a suggested natural language query for searching the transcript or a key moment within the received audio; and causing at least one of the suggested natural language query or the key moment to be surfaced on one or more remote devices.

Claims (65)

1. A system for surfacing natural language queries from one or more transcripts, the system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

receive audio associated with different participants of a virtual meeting;

convert, through automated speech recognition, the received audio to text to form a transcript of the audio, wherein the transcript includes text of the audio and identifications of speakers associated with portions of the text corresponding to utterances from the respective speakers;

generate input signals based on at least the transcript;

execute at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate a suggested natural language query for searching the transcript;

cause the suggested natural language query to be surfaced on one or more remote devices;

receive a selection of the suggested natural language query;

based on receiving the selection of the suggested natural language query, execute the suggested natural language query against the transcript; and

generate query results.

2. The system of claim 1 , wherein the input signals include at least one of: the text of the transcript, audio corresponding to the text of the transcript, speaker identification for utterances in the transcript, durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting.

3. The system of claim 1 , wherein:

an utterance in the transcript has a duration greater than an utterance-duration threshold;

a speaker associated with the utterance has a hierarchy level greater than a hierarchy threshold; and

the suggested natural language query includes a query regarding what the speaker said in the utterance.

4. The system of claim 1 , wherein the operations further comprise:

analyze the text of transcript to identify a topic in the transcript; and

wherein the suggested natural language query includes the identified topic.

5. The system of claim 1 , wherein:

a chat entry is received during the virtual meeting;

the chat entry receives a count of reactions greater than a reaction-count threshold; and

the suggested natural language query includes question regarding key moments of the virtual meeting.

6. The system of claim 1 , wherein the operations further comprise:

performing a semantic analysis of the text of the transcript to identify an uttered name; and

the suggested natural language query is based on the uttered name.

7. The system of claim 1 , wherein the suggested natural language query is surfaced during the virtual meeting.

8. A computer-implemented method for surfacing suggested natural language queries from one or more transcripts, the method comprising:

converting received audio to text, through automated speech recognition, to form a transcript of the audio, wherein the transcript includes text of the audio and identifications of speakers associated with portions of the text corresponding to utterances from the respective speakers;

generating input signals based on at least the transcript;

performing a semantic analysis of the text of the transcript to identify an utterance, by a first participant, of a name of a second participant;

executing at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate at least one of a suggested natural language query for searching the transcript or a key moment within the received audio; and

causing at least one of the suggested natural language query or the key moment to be surfaced on one or more remote devices, wherein at least one of the suggested natural language query or the key moment is based on the uttered name of the second participant and the at least one of the suggested natural language query or the key moment is caused to be surfaced on the remote device of the second participant.

9. The method of claim 8 , wherein the input signals include at least one of: the text of the transcript, audio corresponding to the text of the transcript, speaker identification for utterances in the transcript, durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting.

10. The method of claim 8 , wherein the natural language query and the key moment are surfaced.

11. The method of claim 10 , wherein the natural language query and the key moment are concurrently surfaced in a videoconference interface during a videoconference.

12. The method of claim 11 , wherein the videoconference interface includes a plurality of video segments, a transcript-search input field, and at least a portion of the transcript.

13. The method of claim 12 , further comprising:

receiving a selection of the surfaced key moment; and

in response to receiving the selection, navigating the video segments to a time point at which the key moment occurred.

14. The method of claim 8 , further comprising:

receiving a selection of the suggested natural language query;

based on receiving the selection of the suggested natural language query, execute the suggested natural language query against the transcript; and

generating query results.

15. A computer-implemented system for surfacing suggested natural language queries from one or more transcripts, the system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

access a transcript generated from audio of multiple participants of a virtual meeting, wherein the transcript includes text of the audio and identifications of the speakers associated with the respective portions of the text;

generate input signals based on at least the transcript, wherein the input signals include at least two of: the text of the transcript, audio corresponding to the text of the transcript, speaker identification for utterances in the transcript, durations of utterances, sentiment of an utterance, a title of a speaker, hierarchy level of a speaker, relationship of a speaker to a user, a relationship of a speaker to other speakers in the virtual meeting, identification of participants of the virtual meeting, a title of the virtual meeting, a duration of the virtual meeting, a chat entry for the virtual meeting, live reactions received during the meeting, reactions to a chat entry, or content presented during the virtual meeting;

execute at least one of one or more heuristics or a trained machine-learning (ML) model, using the generated input signals as an input, to generate a suggested natural language query for searching the transcript;

cause the suggested natural language query to be surfaced on one or more remote devices;

receive a selection of the suggested natural language query;

based on receiving the selection of the suggested natural language query, execute the suggested natural language query against the transcript; and

generate query results.

16. The system of claim 15 , wherein the operations further comprise:

cause at least one of key topics or key speakers for the virtual meeting to be surfaced.

17. The system of claim 16 , wherein the natural language query and the at least one of the key topics or key speakers are concurrently surfaced within a videoconference interface.

18. The system of claim 15 , wherein generating the input signals comprises:

performing a semantic analysis on the transcript to identify one or more topics of the transcript; and

filtering the transcript based on the identified topics.

19. The system of claim 18 , wherein filtering the transcript causes removal of non-informative content from the transcript for generation of input signals.

20. The system of claim 15 , wherein:

an utterance in the transcript has a duration greater than an utterance-duration threshold;

a speaker associated with the utterance has a hierarchy level greater than a hierarchy threshold; and

the suggested natural language query includes a query regarding what the speaker said in the utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2021
From: MILLER, ADI; STERENBERG, MICHAEL; SOMECH, HAIM
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057755/0047 →
Continuity (1)
Related Publication 20230115098A1 · Apr 13, 2023
Cited By (2)
US 12,321,383 US 12,717,849