IP Library Granted Patent US 10,019,989
Granted Patent B2
US 10,019,989 · App. 15/262,284 · Granted Jul 10, 2018

Text transcript generation from a communication session

Inventor: Jason John Gauci (Mountain View, CA)
Assignee: Google LLC
G10L15/08G06F17/2235G06F17/241G06F17/28G06Q30/0277G10L15/26G10L25/78G10L2015/088H04L65/403H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,989
App. No.
15/262,284
Granted
Jul 10, 2018
Kind
B2
Abstract

Techniques, systems, and devices for managing streaming media among end user devices in a video conferencing system are described. For example, a transcript may be automatically generated for a video conference. In one example, a method may include receiving a combined media stream comprising a plurality of media sub-streams each associated with one of a plurality of end user devices, wherein each of the plurality of media sub-streams comprises a respective video component and a respective audio component. The method may also include, for each of the media-sub-streams, separating the audio component from the respective video component, for each audio component of the respective media sub-streams, transcribing speech from the audio component to text for the respective media sub-stream, and combining the text for each of the respective media sub-streams into a combined transcription. In some examples, the combined transcription may also be translated into a user selected language.

Claims (61)

1. A method for transcribing speech in a communication session comprising:

transmitting a virtual communication session in substantially real-time to a plurality of end user devices;

receiving, by one or more processors, a combined media stream comprising a plurality of media sub-streams each associated with one of the plurality of end user devices, wherein each of the plurality of media sub-streams in the combined media stream comprises a respective video component and a respective audio component;

for each of the plurality of media sub-streams, separating, by the one or more processors, the respective audio component from the respective video component;

for each separate audio component, transcribing, by the one or more processors, at least a portion of speech from the audio component to text;

providing a transcription in substantially real-time; and

annotating the text for the audio component of each respective media sub-stream to include additional content, wherein annotating the text comprises:

determining one or more keywords of the text;

selecting, based on the one or more keywords, one or more advertisements or a link; and

updating the transcription with the one or more advertisements or the link in association with at least a portion of the text.

2. The method of claim 1 , wherein the one or more advertisements are provided within the text.

3. The method of claim 1 , wherein the one or more advertisements are provided at least one of in a border and next to a field containing the text.

4. The method of claim 1 , wherein:

annotating the text for the audio component of each respective media sub-stream to include additional content further comprises determining that the one or more keywords include a street address of a property; and

the link is a map to the street address of the property included in the one or more keywords.

5. The method of claim 1 , wherein:

determining the one or more keywords of the text includes determining that the one or more keywords include a phone number; and

the link is for dialing the phone number.

6. The method of claim 1 , wherein determining the one or more keywords of the text is based on at least one of a context of the text and a frequency with which at least one of a word and a phrase is used in the text.

7. The method of claim 1 , wherein:

the communication session is a real-time communication session; and

the text and one or more advertisements in association with the text is provided during the real-time communication session.

8. A server device operable to transcribe speech in a communication session comprising:

a memory; and

one or more processors coupled to the memory and operable to execute instructions stored in the memory, the one or more processors configured to:

transmit a virtual communication session in substantially real-time to the plurality of end user devices;

receive a media stream associated with a plurality of end user devices, wherein the media stream comprises a video component and an audio component;

separate the audio component from the video component;

transcribe at least a portion of speech from the audio component to text;

provide a transcription in substantially real-time; and

annotate the text for the audio component to include additional content by:

determining one or more keywords of the text;

searching for one or more of an image, a video, music, and an article that correspond to the one or more keywords of the text;

selecting, based on the one or more keywords, one or more advertisements or a link that correspond to the one or more of the image, the video, the music, and the article; and

updating the transcription with the one or more advertisements or the link in association with at least a portion of the text to a user.

9. The server device of claim 8 , wherein the one or more advertisements are provided at least one of in a border and next to a field containing the text.

10. The server device of claim 8 , wherein annotating the text for the audio component to include additional content further comprises:

selecting, based on the one or more keywords, one or more hyperlinks; and

inserting at least one of the one or more hyperlinks into the text.

11. The server device of claim 10 , wherein the one or more hyperlinks include at least one of a map of an address based on the one or more keywords including the address, an option to dial a phone number based on the one or more keywords including the phone number, an image, a video, music, and an article.

12. The server device of claim 8 , wherein determining the one or more keywords of the text is based on at least one of a context of the text and a frequency with which at least one of a word and a phrase is used in the text.

13. The server device of claim 8 , wherein:

the communication session is a real-time communication session; and

the text and one or more advertisements in association with the text is provided during the real-time communication session.

14. A non-transitory computer storage medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

transmitting a virtual communication session in substantially real-time to the plurality of end user devices;

receiving, by one or more processors, a combined media stream comprising a plurality of media sub-streams each associated with one of the plurality of end user devices, wherein each of the plurality of media sub-streams in the combined media stream comprises a respective video component and a respective audio component;

for each of the plurality of media sub-streams, separating, by the one or more processors, the respective audio component from the respective video component;

for each separate audio component, transcribing, by the one or more processors, at least a portion of speech from the audio component to text;

providing a transcription in substantially real-time; and

annotating the text for the audio component of each respective media sub-stream to include additional content, wherein annotating the text comprises:

determining one or more keywords of the text;

selecting, based on the one or more keywords, one or more advertisements or a link; and

updating the transcription with the one or more advertisements or the link in association with at least a portion of the text.

15. The computer storage medium of claim 14 , wherein the one or more advertisements are provided within the text.

16. The computer storage medium of claim 14 , wherein the one or more advertisements are provided at least one of in a border and next to a field containing the text.

17. The computer storage medium of claim 14 , wherein annotating the text for the audio component of each respective media sub-stream to include additional content further comprises:

selecting, based on the one or more keywords, one or more hyperlinks; and

inserting at least one of the one or more hyperlinks into the text.

18. The computer storage medium of claim 17 , wherein the one or more hyperlinks include at least one of a map of an address based on the one or more keywords including the address, an option to dial a phone number based on the one or more keywords including the phone number, an image, a video, music, and an article.

19. The computer storage medium of claim 14 , wherein determining the one or more keywords of the text is based on at least one of a context of the text and a frequency with which at least one of a word and a phrase is used in the text.

Assignments (2)
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2017
From: GAUCI, JASON JOHN
To: GOOGLE INC.
Reel/Frame 043707/0378 →
Continuity (3)
Continuation 13599908 · Aug 30, 2012
Provisional Application 61529607 · Aug 31, 2011
Related Publication 20170011740A1 · Jan 12, 2017
Cited By (1)
US 12,632,321