IP Library Granted Patent US 10,942,953
Granted Patent B2
US 10,942,953 · App. 16/007,890 · Granted Mar 9, 2021

Generating summaries and insights from meeting recordings

Inventor: Mohamed Gamal Mohamed Mahmoud (Santa Clara, CA)
Assignee: CISCO TECHNOLOGY, INC.
G06F16/313G06F16/345G06F16/685G10L15/08G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,942,953
App. No.
16/007,890
Granted
Mar 9, 2021
Kind
B2
Abstract

One embodiment of the present invention sets forth a technique for generating a summary of a recording. The technique includes generating an index associated with the recording, wherein the index identifies a set of terms included in the recording and, for each term in the set of terms, a corresponding location of the term in the recording. The technique also includes determining categories of predefined terms to be identified in the index and identifying a first subset of the terms in the index that match a first portion of the predefined terms in the categories. The technique further includes outputting a summary of the recording comprising the locations of the first subset of terms in the recording and listings of the first subset of terms under one or more corresponding categories.

Claims (59)

1. A method for generating a summary of a recording, comprising:

generating an index associated with the recording, wherein the index identifies a set of terms included in the recording and, for each term in the set of terms, a corresponding location of the term in the recording;

determining categories of predefined terms to be identified in the index;

identifying a first subset of the terms in the index that match a first portion of the predefined terms in the categories;

outputting a summary of the recording comprising locations of the first subset of terms in the recording and listings of the first subset of terms under one or more corresponding categories, wherein outputting the summary comprising the locations of the first subset of terms comprises outputting the summary comprising temporal locations of the first subset of terms in the recording, and wherein outputting the summary comprising the listings of the first subset of terms under the one or more corresponding categories comprises outputting the summary comprising the first subset of terms under a highest precedence category of the one or more corresponding categories under which the first subset of terms were found;

generating semantic expansions of the predefined terms;

identifying a second subset of the terms in the index that match the semantic expansions; and

including the second subset of the terms in the summary.

2. The method of claim 1 , further comprising:

analyzing the summary and the index for insights related to the recording; and

outputting the insights with the summary.

3. The method of claim 2 , wherein the insights comprise at least one of an inquisitiveness, a quantitativeness, a sentiment, a topic, a theme, and an entity.

4. The method of claim 1 , wherein the semantic expansions of the predefined terms comprise at least one of:

additional terms that are within a semantic distance of the predefined terms; and

user-defined paraphrases of the predefined terms.

5. The method of claim 1 , wherein generating the index comprises:

generating, by a set of automatic speech recognition (ASR) engines, a set of transcript lattices comprising the set of terms, locations of the set of terms in the recording, and confidences representing predictive accuracies for the set of terms;

calculating contemporary word counts for the set of terms from the set of transcript lattices; and

creating the index by storing, for each term in the set of terms, the term in association with one or more locations of the term in the recording, one or more confidences associated with the one or more locations of the term, and one or more contemporary word counts associated with the one or more locations of the term.

6. The method of claim 5 , wherein generating the index further comprises:

for each term in the set of terms, storing the term in association with one or more related terms and one or more ASRs used to produce the term.

7. The method of claim 6 , wherein generating the index further comprises:

filtering the set of terms in the index by a set of criteria.

8. The method of claim 7 , wherein the set of criteria comprises at least one of a maximum contemporary word count, a confidence threshold, a minimum number of ASR engines used to generate each term, a stop word list, a blacklist, and a part-of-speech (POS) tag.

9. The method of claim 1 , wherein determining the categories associated with the recording comprises:

obtaining selections of the categories from curated categories, dynamic categories, and user-defined categories.

10. The method of claim 9 , wherein the curated categories comprise at least one of a business jargon category, a numbers category, a dates category, an actions category, a locations category, a commitments category, a questions category, a points of contention category, an idioms and sayings category, and a requests category.

11. The method of claim 9 , wherein the dynamic categories comprise at least one of an attendees category, an agenda category, an introductions category, a title category, a highlights category, and a related transcripts category.

12. A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of:

generating an index associated with the recording, wherein the index identifies a set of terms included in the recording and, for each term in the set of terms, a corresponding location of the term in the recording;

determining categories of predefined terms to be identified in the index;

identifying a first subset of the terms in the index that match a first portion of the predefined terms in the categories;

outputting a summary of the recording comprising locations of the first subset of terms in the recording and listings of the first subset of terms under one or more corresponding categories, wherein outputting the summary comprising the locations of the first subset of terms comprises outputting the summary comprising temporal locations of the first subset of terms in the recording, and wherein outputting the summary comprising the listings of the first subset of terms under the one or more corresponding categories comprises outputting the summary comprising the first subset of terms under a highest precedence category of the one or more corresponding categories under which the first subset of terms were found;

generating semantic expansions of the predefined terms;

identifying a second subset of the terms in the index that match the semantic expansions; and

including the second subset of the terms in the summary.

13. The non-transitory computer readable medium of claim 12 , wherein the steps further comprise:

identifying a second subset of the terms in the index that match a second portion of the predefined terms in the categories; and

excluding the second subset of the terms from the summary.

14. The non-transitory computer readable medium of claim 13 , wherein the first portion of the predefined terms comprises whitelisted terms and the second portion of the predefined terms comprises blacklisted terms.

15. The non-transitory computer-readable medium of claim 13 , wherein the steps further comprise:

analyzing the summary and the index for insights related to the recording; and

outputting the insights with the summary.

16. The non-transitory computer-readable medium of claim 13 , wherein generating the index comprises:

generating, by a set of automatic speech recognition (ASR) engines, a set of transcript lattices comprising the set of terms, locations of the set of terms in the recording, and confidences representing predictive accuracies for the set of terms;

calculating contemporary word counts for the set of terms from the set of transcript lattices;

creating the index by storing, for each term in the set of terms, the term in association with one or more locations of the term in the recording, one or more confidences associated with the one or more locations of the term, one or more contemporary word counts associated with the one or more locations of the term, one or more related terms, and one or more ASR engines used to produce the term; and

filtering the set of terms in the index by a set of criteria.

17. The non-transitory computer-readable medium of claim 16 , wherein the set of criteria comprises at least one of a maximum contemporary word count, a confidence threshold, a minimum number of ASR engines used to generate each term, a stop word list, a blacklist, and a part-of-speech (POS) tag.

18. A system, comprising:

a memory that stores instructions, and

a processor that is coupled to the memory and, when executing the instructions, is configured to:

generate an index associated with the recording, wherein the index identifies a set of terms included in the recording and, for each term in the set of terms, a corresponding location of the term in the recording;

determine categories of predefined terms to be identified in the index;

identify a first subset of the terms in the index that match a first portion of the predefined terms in the categories;

output a summary of the recording comprising locations of the first subset of terms in the recording and listings of the first subset of terms under one or more corresponding categories, wherein the processor being operative to output the summary comprising the locations of the first subset of terms comprises the processor being operative to output the summary comprising temporal locations of the first subset of terms in the recording, and wherein the processor being operative to output the summary comprising the listings of the first subset of terms under the one or more corresponding categories comprises the processor being operative to output the summary comprising the first subset of terms under a highest precedence category of the one or more corresponding categories under which the first subset of terms were found;

generate semantic expansions of the predefined terms;

identify a second subset of the terms in the index that match the semantic expansions; and

include the second subset of the terms in the summary.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2020
From: RIZIO LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 053195/0862 →
CHANGE OF NAME Recorded May 26, 2020
From: RIZIO, INC.
To: RIZIO LLC
Reel/Frame 052751/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2019
From: MAHMOUD, MOHAMED GAMAL MOHAMED
To: RIZIO, INC.
Reel/Frame 048604/0430 →
Continuity (1)
Related Publication 20190384854A1 · Dec 19, 2019
Cited By (1)
US 12,725,215