IP Library Granted Patent US 11,288,034
Granted Patent B2
US 11,288,034 · App. 16/849,052 · Granted Mar 29, 2022

Hierarchical topic extraction and visualization for audio streams

Inventors: Tom Neckermann (Seattle, WA); Romain Gabriel Paul Rey (Vancouver, CA); Alexander James Wilson (Seattle, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F3/165G06F3/0482G06F3/04855G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,288,034
App. No.
16/849,052
Granted
Mar 29, 2022
Kind
B2
Abstract

An audio stream is subjected to speech-to-text processing in order to obtain a textual representation of the audio stream. Hierarchical topic extraction is performed on the textual representation to obtain a multi-level hierarchical topic representation of the textual representation. A user interface actuator is generated, which allows a user to search through the audio stream. Different levels of the multi-level hierarchical topic representation are displayed to the user, based upon the speed of actuation of the user interface actuator.

Claims (70)

1. A computer implemented method, comprising:

generating a multi-level hierarchical topic representation of an audio stream, the multi-level hierarchical topic representation comprising textual topic elements, each textual topic element extracted from a corresponding position in the audio stream, wherein

a first level of the multi-level hierarchical topic representation includes more general textual topic elements, and

a second level of the multi-level hierarchical topic representation includes more detailed textual topic elements;

assigning a speed stamp to each level in the multi-level hierarchical topic representation;

generating a representation of a display of a seek/scroll actuator corresponding to the audio stream;

detecting an operator input on the seek/scroll actuator;

detecting an actuator speed at which the operator is moving through the audio stream based on the operator input on the seek/scroll actuator;

identifying a position in the audio stream based on a position of the operator input on the seek/scroll actuator;

identifying which speed stamp most closely corresponds to the detected actuator speed; and

selecting a level in the multi-level hierarchical topic representation that has the identified speed stamp that most closely corresponds to the detected actuator speed;

identifying, from the selected level, a particular textual topic element extracted from the identified position in the audio stream; and

generating a representation of a display of the particular textual topic element along with the seek/scroll actuator.

2. The computer implemented method of claim 1 , and further comprising:

obtaining the audio stream; and

performing speech-to-text processing on the audio stream to generate the textual representation of the audio stream.

3. The computer implemented method of claim 1 , wherein generating the multi-level hierarchical topic representation comprises:

for each level in the multi-level hierarchical topic representation, dividing the audio stream into different extraction windows, each extraction window corresponding to a different window of time in the audio stream; and

for each extraction window, generating a set of textual topic display elements indicative of topics extracted from the corresponding extraction window.

4. The computer implemented method of claim 3 wherein generating a set of textual topic display elements for each extraction window comprises:

generating more general topic display elements for each window, on the first level of the multi-level hierarchical topic representation; and

generating more detailed topic display elements for each window, on the second level of the multi-level hierarchical topic representation.

5. The computer implemented method of claim 1 , wherein the display of the particular textual topic element comprises an extracted topic display that is adjacent the seek/scroll actuator.

6. The computer implemented method of claim 1 , wherein the multi-level hierarchical topic representation comprises:

a first level topic representation indicative of a first level of topic detail extracted from the identified position in the audio stream, and

a second level topic representation indicative of a second level of topic detail extracted from the identified position in the audio stream, the second level of topic detail being more detailed than the first level of topic detail; and

identifying the particular textual topic element comprises selecting a topic representation from the first or second level topic representations based on the scroll speed.

7. The computer implemented method of claim 3 wherein generating a set of textual topic display elements indicative of topics extracted from the corresponding extraction window comprises at least one of:

extracting, as the set of textual topic display elements, words used in the extraction window of the audio stream,

obtaining, as the set of textual topic display elements, a summary of words used in the extraction window of the audio stream, or

for extraction windows on a lowest level of the multi-level hierarchical topic representation, extracting a representative text fragment of a portion of the audio stream in the extraction window as the set of textual topic display elements.

8. A computing system, comprising:

at least one processor; and

memory storing instructions executable by the at least one processor, wherein the instructions, when executed, cause the computing system to:

generate a multi-level hierarchical topic representation of an audio stream, the multi-level hierarchical topic representation comprising textual topic elements, each textual topic element extracted from a corresponding position in the audio stream, wherein

a first level of the multi-level hierarchical topic representation includes more general topic elements, and

a second level of the multi-level hierarchical topic representation includes more detailed topic elements;

assign a speed stamp to each level in the multi-level hierarchical topic representation;

generate a representation of a display of a seek/scroll actuator corresponding to the audio stream;

detect an operator input on the seek/scroll actuator;

detect an actuator speed at which the operator is moving through the audio stream based on the operator input on the seek/scroll actuator;

identify a position in the audio stream based on a position of the operator input on the seek/scroll actuator;

identify which speed stamp most closely corresponds to the detected actuator speed; and

select a level in the multi-level hierarchical topic representation that has the identified speed stamp that most closely corresponds to the detected actuator speed;

identify, from the selected level, a particular textual topic element extracted from the identified position in the audio stream; and

generate a representation of a display of the particular textual topic element along with the seek/scroll actuator.

9. The computing system of claim 8 , wherein the instructions cause the computing system to:

perform speech-to-text processing on the audio stream; and

generate the textual representation of the audio stream based on the speech-to-text processing.

10. The computing system of claim 8 , wherein the instructions cause the computing system to:

for each level in the multi-level hierarchical topic representation, divide the audio stream into different extraction windows, each extraction window corresponding to a different window of time in the audio stream, and

for each extraction window, generate a set of textual topic display elements indicative of topics extracted from the corresponding extraction window.

11. The computing system of claim 10 , wherein the instructions cause the computing system to:

generate more general topic display elements for each window, on the first level of the multi-level hierarchical topic display representation, and

generate more detailed topic display elements for each window, on the second level of the multi-level hierarchical topic display representation.

12. The computing system of claim 8 , wherein the display of the particular textual topic element comprises an extracted topic display that is adjacent the seek/scroll actuator.

13. The computing system of claim 8 , wherein the instructions cause the computing system to:

extract, as a set of textual topic elements in the second level, words used in the extraction window of the audio stream.

14. The computing system of claim 8 , wherein the instructions cause the computing system to:

obtain, as a set of textual topic elements in the first level, a summary of words used in the extraction window of the audio stream.

15. A computer implemented method, comprising:

generating a multi-level hierarchical topic representation of a textual document, the multi-level hierarchical topic representation comprising:

a first textual topic element indicative of a first level of topic detail extracted from a particular position in the textual document, and

a second textual topic element indicative of a second level of topic detail extracted from the particular position in the textual document, the second level of topic detail being more detailed than the first level of topic detail;

assigning a speed stamp to each level in the multi-level hierarchical topic representation;

generating a representation of a positioning actuator display element that is actuated to navigate to the particular position in the textual document;

detecting a speed of operator actuation of the positioning actuator display element;

identifying which speed stamp most closely corresponds to the detected speed of operator actuation of the positioning actuator display element;

selecting a textual topic element, from the first and second textual topic elements in the multi-level hierarchical topic representation, that has the identified speed stamp that most closely corresponds to the detected speed of operator actuation of the positioning actuator display element; and

generating a display of the selected textual topic element from the multi-level hierarchical topic representation along with the positioning actuator display element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2020
From: NECKERMANN, TOM; REY, ROMAIN GABRIEL PAUL; WILSON, ALEXANDER JAMES
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 052402/0728 →
Continuity (1)
Related Publication 20210326098A1 · Oct 21, 2021