IP Library › Granted Patent US 10,185,711
Granted Patent B1
US 10,185,711 · App. 15/202,039 · Granted Jan 22, 2019

Speech recognition and summarization

Inventors: Glen Shires (Danville, CA); Sterling Swigart (Mountain View, CA); Jonathan Zolla (Belmont, CA); Jason J. Gauci (Mountain View, CA)
Assignee: Google LLC
G06F17/2765G10L15/26G10L21/10H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,185,711
App. No.
15/202,039
Granted
Jan 22, 2019
Kind
B1
Abstract

The subject matter of this specification can be embodied in, among other things, a method that includes receiving two or more data sets each representing speech of a corresponding individual attending an internet-based social networking video conference session, decoding the received data sets to produce corresponding text for each individual attending the internet-based social networking video conference, and detecting characteristics of the session from a coalesced transcript produced from the decoded text of the attending individuals for providing context to the internet-based social networking video conference session.

Claims (69)

1. A computer-implemented method comprising:

receiving, by a videoconference system that includes an automated speech recognizer and a context builder and from a first computing device that includes a microphone, speech data representing an utterance spoken by a particular participant of a video conference and captured by the microphone of the first computing device;

transcribing, by the automated speech recognizer of the video-conference system, the speech data representing the utterance spoken by the particular participant of the video conference into text in real-time;

determining, by the automated speech recognizer of the videoconference system, a topic of the video conference by analyzing one or more words and/or phrases in the text of the speech data;

annotating, by the automated speech recognizer of the videoconference system, the text of the speech data by:

determining one or more relevant terms in the text of the speech data as being potentially relevant to the determined topic; and

identifying, using the one or more relevant terms in the text, one or more resources associated with the determined topic of the video conference, each identified resource comprising at least one of advertising content, a search result, an event, or a location; and

for each identified resource:

generating, using the context builder of the video conference system, a user interface component for the identified resource; and

outputting, by the context builder of the video conference system, the corresponding user interface component for the identified resource to a second computing device in real-time, the corresponding user interface component when received by the second computing device causing the second computing device to display the corresponding user interface component on a videoconference graphical user interface executing on the second computing device.

2. The computer-implemented method of claim 1 , wherein identifying one or more resources associated with the determined topic of the video conference comprises obtaining, from an advertising module, the advertising content that corresponds to the determined topic of the video conference.

3. The computer-implemented method of claim 1 , wherein identifying one or more resources associated with the determined topic of the video conference comprises:

obtaining, from a search engine, one or more search results that are identified as a result of performing a query using at least one of the one or more relevant terms in the text of the speech data representing the utterance spoken by the particular participant of the video conference; and

selecting a particular search result from among the one or more search results identified as the result of performing the query.

4. The computer-implemented method of claim 1 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the event associated with the determined topic of the video conference; and

generating an invitation for the identified event, the invitation comprising at least one of a calendar date, time, location, or guest list for the identified event.

5. The computer-implemented method of claim 1 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating a map image associated with the identified location.

6. The computer-implemented method of claim 1 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating a hyperlink to a map associated with the identified location.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a videoconference system that includes an automated speech recognizer and a context builder and from a first computing device that includes a microphone, speech data representing an utterance spoken by a particular participant of a video conference and captured by the microphone of the first computing device;

transcribing, by the automated speech recognizer of the video-conference system, the speech data representing the utterance spoken by the particular participant of the video conference into text in real-time;

determining, by the automated speech recognizer of the videoconference system, a topic of the video conference by analyzing one or more words and/or phrases in the text of the speech data;

annotating, by the automated speech recognizer of the videoconference system, the text of the speech data by:

determining one or more relevant terms in the text of the speech data as being potentially relevant to the determined topic; and

identifying, using the one or more relevant terms in the text, one or more resources associated with the determined topic of the video conference, each identified resource comprising at least one of advertising content, a search result, an event, or a location; and

for each identified resource:

generating, using the context builder of the video conference system, a corresponding user interface component for the identified resource; and

outputting, by the context builder of the video conference system, the corresponding user interface component for the identified resource to a second computing device in real-time, the corresponding user interface component when received by the second computing device causing the second computing device to display the corresponding user interface component on a videoconference graphical user interface executing on the second computing device.

8. The system of claim 7 , wherein identifying one or more resources associated with the determined topic of the video conference comprises obtaining, from an advertising module, the advertising content that corresponds to the determined topic of the video conference.

9. The system of claim 7 , wherein identifying one or more resources associated with the determined topic of the video conference comprises:

obtaining, from a search engine, one or more search results that are identified as a result of performing a query using at least one of the one or more relevant terms in the text of the speech data representing the utterance spoken by the particular participant of the video conference; and

selecting a particular search result from among the one or more search results identified as the result of performing the query.

10. The system of claim 7 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the event associated with the determined topic of the video conference; and

generating an invitation for the identified event, the invitation comprising at least one of a calendar date, time, location, or guest list for the identified event.

11. The system of claim 7 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating a map image associated with the identified location.

12. The system of claim 7 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating a hyperlink to a map associated with the identified location.

13. A computer-readable storage device storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a videoconference system that includes an automated speech recognizer and a context builder and from a first computing device that includes a microphone, speech data representing an utterance spoken by a particular participant of a video conference and captured by the microphone of the first computing device;

transcribing, by the automated speech recognizer of the video-conference system, the speech data representing the utterance spoken by the particular participant of the video conference into text in real-time;

determining, by the automated speech recognizer of the videoconference system, a topic of the video conference by analyzing one or more words and/or phrases in the text of the speech data;

annotating, by the automated speech recognizer of the videoconference system, the text of the speech data by:

determining one or more relevant terms in the text of the speech data as being potentially relevant to the determined topic; and

identifying, using the one or more relevant terms in the text, one or more resources associated with the determined topic of the video conference, each identified resource comprising at least one of advertising content, a search result, an event, or a location; and

for each identified resource:

generating, using the context builder of the video conference system, a corresponding user interface component for the identified resource; and

outputting, by the context builder of the video conference system, the corresponding user interface component for the identified resource to a second computing device in real-time, the corresponding user interface component when received by the second computing device causing the second computing device to display the corresponding user interface component on a videoconference graphical user interface executing on the second computing device.

14. The computer-readable storage device of claim 13 , wherein identifying one or more resources associated with the determined topic of the video conference comprises:

obtaining, from a search engine, one or more search results that are identified as a result of performing a query using at least one of the one or more relevant terms in the text of the speech data representing the utterance spoken by the particular participant of the video conference; and

selecting a particular search result from among the one or more search results identified as the result of performing the query.

15. The computer-readable storage device of claim 13 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the event associated with the determined topic of the video conference; and

generating an invitation for the identified event, the invitation comprising at least one of a calendar date, time, location, or guest list for the identified event.

16. The computer-readable storage device of claim 13 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating map image associated with the identified location.

17. The computer-readable storage device of claim 13 , wherein generating the corresponding user interface component for the identified resource comprises:

identifying the location associated with the determined topic of the video conference; and

generating a hyperlink to a map associated with the identified location.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2016
From: SHIRES, GLEN; SWIGART, STERLING; ZOLLA, JONATHAN; GAUCI, JASON J.
To: GOOGLE INC.
Reel/Frame 039110/0916 →
Continuity (3)
Continuation 14078800 · Nov 13, 2013
Continuation 13743838 · Jan 17, 2013
Provisional Application 61699072 · Sep 10, 2012
Cited By (1)
US 12,651,118