IP Library Granted Patent US 8,120,638
Granted Patent B2
US 8,120,638 · App. 11/625,068 · Granted Feb 21, 2012

Speech to text conversion in a videoconference

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,120,638
App. No.
11/625,068
Granted
Feb 21, 2012
Kind
B2
Abstract

Various embodiments of a method for automatically converting audio speech in a videoconference into text information are described. According to one embodiment of the method, a videoconferencing device at a first endpoint in the videoconference may receive a stream of video information and audio information from a videoconferencing device at a second endpoint in the videoconference. The audio information includes speech of a participant at the second endpoint. The videoconferencing device at the first endpoint may automatically convert the speech into text information.

Claims (37)

1. A method for processing audio information in a videoconference, comprising:

a first videoconferencing device at a first endpoint in the videoconference receiving video information and audio information from a second videoconferencing device at a second endpoint in the videoconference, wherein the audio information includes speech of a videoconference participant;

the first videoconferencing device automatically converting the speech into text information;

the first videoconferencing device storing the text information in a memory;

the first videoconferencing device creating a composite image including the video information and the text information; and

the first videoconferencing device displaying the composite image on a display device at the first endpoint.

2. The method of claim 1 , wherein said automatically converting the speech into text information comprises dynamically converting the speech into text information as the audio information is received by the first videoconferencing device.

3. The method of claim 1 , further comprising:

the first videoconferencing device creating one or more files; and

the first videoconferencing device storing the text information in the one or more files, wherein the one or more files provide a transcript of the speech of at least one participant in the videoconference.

4. The method of claim 1 , further comprising:

the first videoconferencing device performing voice recognition to identify a participant based on speech from the participant; and

the first videoconferencing device associating the text information with the participant in response to said identifying the participant based on the speech of the participant.

5. The method of claim 4 , wherein said associating the text information with the participant comprises including a name of the participant in the text information.

6. The method of claim 4 , wherein said associating the text information with the participant comprises displaying the text information proximally to an image of the participant.

7. The method of claim 1 ,

wherein the speech of the participant is in a first language;

wherein said converting the speech into text information comprises converting the speech into first text information in the first language;

wherein the method further comprises the first videoconferencing device automatically translating the first text information into second text information, wherein the second text information is in a second language.

8. The method of claim 7 , further comprising the first videoconferencing device displaying the second text information in the second language on the display device.

9. A videoconferencing device, comprising:

an input port configured to receive, at first endpoint in a videoconference, video information and audio information transmitted from a second endpoint in the videoconference, wherein the audio information includes speech of a participant at the second endpoint;

a memory; and

one or more computational elements configured to:

automatically convert the speech into text information;

store the text information in the memory

create a composite image including the video information and the text information; and

display the composite image on a display device at the first endpoint via an output port.

10. The videoconferencing device of claim 9 , wherein the one or more computational elements include at least one processor.

11. The videoconferencing device of claim 9 , wherein said automatically converting the speech into text information comprises dynamically converting the speech into text information as the audio information is received.

12. The videoconferencing device of claim 9 , wherein the one or more computational elements are further configured to store the text information as a transcript of the speech of the participant in the videoconference.

13. The videoconferencing device of claim 9 ,

wherein the speech of the participant is in a first language;

wherein said converting the speech into text information comprises converting the speech into first text information in the first language; and

wherein the one or more computational elements are further configured to automatically translate the first text information into second text information, wherein the second text information is in a second language.

14. The videoconferencing device of claim 13 ,

wherein the one or more computational elements are further configured to send the second text information in the second language for display on the display device via the output port.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2025
From: SL MIDCO 2, LLC; SL MIDCO 1, LLC; LO PLATFORM MIDCO, INC.; SERENOVA, LLC; LIFESIZE, INC.; TELESTRAT LLC; LIGHT BLUE OPTICS INC.; SERENOVA WFM, INC.
To: ENGHOUSE INTERACTIVE INC.
Reel/Frame 070932/0415 →
SECURITY INTEREST Recorded Mar 16, 2020
From: LIFESIZE, INC.
To: WESTRIVER INNOVATION LENDING FUND VIII, L.P.
Reel/Frame 052179/0063 →
SECURITY INTEREST Recorded Mar 2, 2020
From: SERENOVA, LLC; LIFESIZE, INC.; LO PLATFORM MIDCO, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 052066/0126 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2016
From: LIFESIZE COMMUNICATIONS, INC.
To: LIFESIZE, INC.
Reel/Frame 037900/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2007
From: KENOYER, MICHAEL L.
To: LIFESIZE COMMUNICATIONS, INC.
Reel/Frame 018795/0136 →