IP Library Granted Patent US 11,227,602
Granted Patent B2
US 11,227,602 · App. 16/689,662 · Granted Jan 18, 2022

Speech transcription using multiple data sources

Inventors: Vincent Charles Cheung (San Carlos, CA); Chengxuan Bai (San Mateo, CA); Yating Sasha Sheng (San Francisco, CA)
Assignee: Facebook Technologies, LLC
G10L17/00G06F3/011G06K9/00228G06T19/006G10L25/63H04R1/406H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,602
App. No.
16/689,662
Granted
Jan 18, 2022
Kind
B2
Abstract

This disclosure describes transcribing speech using audio, image, and other data. A system is described that includes an audio capture system configured to capture audio data associated with a plurality of speakers, an image capture system configured to capture images of one or more of the plurality of speakers, and a speech processing engine. The speech processing engine may be configured to recognize a plurality of speech segments in the audio data, identify, for each speech segment of the plurality of speech segments and based on the images, a speaker associated with the speech segment, transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of the speaker associated with the speech segment, and analyze the transcription to produce additional data derived from the transcription.

Claims (40)

1. A system comprising:

an audio capture system configured to capture audio data associated with a plurality of speakers;

an image capture system configured to capture images of one or more of the plurality of speakers; and

a speech processing engine configured to:

recognize a plurality of speech segments in the audio data,

access external data;

identify, for each speech segment of the plurality of speech segments and based on the images and the external data, a speaker associated with the speech segment, wherein the external data includes location information associated with the speaker,

transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of which speaker is associated with the speech segment, and

analyze the transcription to produce additional data derived from the transcription.

2. The system of claim 1 , wherein to recognize the plurality of speech segments, the speech processing engine is further configured to recognize, based on the images, the plurality of speech segments.

3. The system of claim 2 , wherein to identify the speaker, the speech processing engine is further configured to detect one or more faces in the images.

4. The system of claim 2 , wherein the speech processing engine is further configured to choose, based on the identity of the speaker associated with each speech segment, one or more speech recognition models.

5. The system of claim 4 , wherein to identify, for each speech segment of the plurality of speech segments, the speaker, the speech processing engine is further configured to detect one or more faces in the images with moving lips.

6. The system of claim 1 , wherein the external data further includes one or more of calendar information and information about meeting invitees.

7. The system of claim 4 , further comprising a head-mounted display (HMD) capable of being worn by a user, and wherein the one or more speech recognition models comprises a voice recognition model for the user.

8. The system of claim 4 , further comprising a head-mounted display (HMD) capable of being worn by a user, wherein the speech processing engine is further configured to identify the user of the HIVID as the speaker of the plurality of speech segments based on attributes of the plurality of speech segments.

9. The system of claim 1 , wherein the audio capturing system comprises a microphone array.

10. The system of claim 7 , wherein the HMD is configured to output artificial reality content, and wherein the artificial reality content comprises a virtual conferencing application including a video stream and an audio stream.

11. The system of claim 1 , wherein the additional data comprises one or more of a calendar invitation for a meeting or event described in the transcription, information related to topics identified in the transcription, or a task list including tasks identified in the transcription.

12. The system of claim 1 , wherein the additional data comprises at least one of: statistics about the transcription including number of words spoken by the speaker, tone of the speaker, information about filler words used by the speaker, percent of time the speaker spoke, information about profanity used, information about the length of words used, a summary of the transcription, or sentiment of the speaker.

13. The system of claim 1 , wherein the additional data includes an audio stream including a modified version of the speech segments associated with at least one of the plurality of speakers.

14. A method comprising:

capturing audio data associated with a plurality of speakers;

capturing images of one or more of the plurality of speakers;

recognizing a plurality of speech segments in the audio data;

accessing external data;

identifying, for each speech segment of the plurality of speech segments and based on the images and the external data, a speaker associated with the speech segment, wherein the external data includes location information associated with the speaker;

transcribing each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of which speaker is associated with the speech segment; and

analyzing the transcription to produce additional data derived from the transcription.

15. The method of claim 14 , wherein the external data further includes one or more of calendar information and information about meeting invitees.

16. The method of claim 14 , wherein the additional data comprises one or more of a calendar invitation for a meeting or event described in the transcription, information related to topics identified in the transcription, or a task list including tasks identified in the transcription.

17. The method of claim 14 , wherein the additional data comprises at least one of: statistics about the transcription including number of words spoken by the speaker, tone of the speaker, information about filler words used by the speaker, percent of time the speaker spoke, information about profanity used, information about the length of words used, a summary of the transcription, or sentiment of the speaker.

18. A non-transitory computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system to:

capture audio data associated with a plurality of speakers;

capture images of one or more of the plurality of speakers;

recognize a plurality of speech segments in the audio data;

access external data;

identify, for each speech segment of the plurality of speech segments and based on the images and the external data, a speaker associated with the speech segment, wherein the external data includes location information associated with the speaker;

transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of which speaker is associated with the speech segment; and

analyze the transcription to produce additional data derived from the transcription.

Assignments (2)
CHANGE OF NAME Recorded Jul 21, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060802/0799 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2020
From: CHEUNG, VINCENT CHARLES; BAI, CHENGXUAN; SHENG, YATING SASHA
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 051673/0862 →
Continuity (1)
Related Publication 20210151058A1 · May 20, 2021