IP Library Granted Patent US 11,315,570
Granted Patent B2
US 11,315,570 · App. 16/373,503 · Granted Apr 26, 2022

Machine learning-based speech-to-text transcription cloud intermediary

Inventor: Shamir Allibhai (San Francisco, CA)
Assignee: Facebook Technologies, LLC
G10L15/30G10L15/22G10L21/028G10L25/84G10L25/90G10L15/26G10L15/32G10L17/02G10L21/0208
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,315,570
App. No.
16/373,503
Granted
Apr 26, 2022
Kind
B2
Abstract

The technology disclosed relates to a machine learning based speech-to-text transcription intermediary which, from among multiple speech recognition engines, selects a speech recognition engine for accurately transcribing an audio channel based on sound and speech characteristics of the audio channel.

Claims (70)

1. A computer-implemented method of cloud-based speech recognition from audio channels without prior training to adapt to speaker(s) in the audio channels, the method including:

accessing a machine-learning engine;

submitting hundreds of speech samples to multiple speech recognition engines, and analyzing how transcription error rates of the speech recognition engines vary with sound and speech characteristics of the speech samples;

receiving an audio channel and qualifying the speech recognition engines as capable of transcribing the audio channel and/or its parts, taking into account at least a recording codec of the audio channel, available transcoding from the recording codec to a speech recognition engine supported codec, a length of the audio channel, and a language of the audio channel;

applying an audio channel analyzer to the audio channel to characterize audio fidelity, background noise, concurrent speech by multiple speakers, timbre, pitch, and audio distortion of the audio channel; and

using the machine learning engine, selecting a speech recognition engine that is qualified as capable of transcribing the audio channel and/or its parts or a transcoded version of the audio channel and/or its parts, based on a combination of least three of the following characteristics:

audio fidelity;

background noise;

concurrent speech by multiple users;

timbre;

pitch; and

audio distortion of the audio channel.

2. The computer-implemented method of claim 1 , further including:

using the multiple speech recognition engines on the audio channel, including using the speech recognition engines sequentially when a first speech recognition engine reports a low confidence score on some or all of its transcription.

3. The computer-implemented method of claim 1 , further including:

using the multiple speech recognition engines on all or separate parts of the audio channel, including using the speech recognition engines when voting on transcription results is used, when different speakers on different tracks of the audio channel, and when different speakers take turns during segments of the audio channel.

4. The computer-implemented method of claim 1 , further including applying the method to separation/identification (diarization) engines.

5. The computer-implemented method of claim 1 , further including applying the method to auto-punctuation engines.

6. The computer-implemented method of claim 1 , further including applying a silence analyzer to the speech samples and the audio channel prior to submission to parse out silent parts of speech.

7. The computer-implemented method of claim 1 , further including performing testing periodically, including daily, weekly, or monthly, wherein the testing comprises the submitting of hundreds of speech samples to multiple speech recognition engines, and the analyzing of how transcription error rates of the speech recognition engines vary with sound and speech characteristics of the speech samples.

8. A computer-implemented method of cloud-based speech recognition from audio channels without prior training to adapt to speaker(s) in the audio channels, the method including:

accessing a machine-learning engine;

submitting thousands of speech samples to multiple speech recognition engines, and analyzing how transcription error rates of the speech recognition engines vary with sound and speech characteristics of the speech samples;

receiving an audio channel and qualifying the speech recognition engines as capable of transcribing the audio channel and/or its parts, taking into account at least a recording codec of the audio channel, available transcoding from the recording codec to a speech recognition engine supported codec, a length of the audio channel, and a language of the audio channel;

applying an audio channel analyzer to the audio channel to characterize audio fidelity, background noise, concurrent speech by multiple speakers, timbre, pitch, and audio distortion of the audio channel; and

using the machine learning engine, selecting a speech recognition engine that is qualified as capable of transcribing the audio channel and/or its parts or a transcoded version of the audio channel and/or its parts, based on a combination of at least three of the following characteristics:

audio fidelity;

background noise;

concurrent speech by multiple users;

timbre;

pitch; and

audio distortion of the audio channel.

9. A computer-implemented method of cloud-based speech recognition from audio channels without prior training to adapt to speaker(s) in the audio channels, the method including:

accessing a machine-learning engine;

submitting dozens of speech samples to multiple speech recognition engines, and analyzing how transcription error rates of the speech recognition engines vary with sound and speech characteristics of the speech samples;

receiving an audio channel and qualifying the speech recognition engines as capable of transcribing the audio channel and/or its parts, taking into account at least a recording codec of the audio channel, available transcoding from the recording codec to a speech recognition engine supported codec, a length of the audio channel, and a language of the audio channel;

applying an audio channel analyzer to the audio channel to characterize audio fidelity, background noise, concurrent speech by multiple speakers, timbre, pitch, and audio distortion of the audio channel; and

using the machine learning engine, selecting a speech recognition engine that is qualified as capable of transcribing the audio channel and/or its parts or a transcoded version of the audio channel and/or its parts, based on a combination of at least three of the following characteristics:

audio fidelity;

background noise;

concurrent speech by multiple users;

timbre;

pitch; and

audio distortion of the audio channel.

10. A system for cloud-based speech recognition from audio channels without prior training to adapt to speaker(s) in the audio channels, comprising:

a machine-learning engine trained to process speech samples;

an analyzer for submitting and analyzing submitting hundreds of speech samples to multiple speech recognition engines, the analyzer determining error rates of the speech recognition engines based on sound and speech characteristics of the speech samples;

the analyzer including receiving an audio channel and qualifying the speech recognition engines as capable of transcribing the audio channel and/or its parts, taking into account at least a recording codec of the audio channel, available transcoding from the recording codec to a speech recognition engine supported codec, a length of the audio channel, and a language of the audio channel;

applying an audio channel analyzer to the audio channel to characterize audio fidelity, background noise, concurrent speech by multiple speakers, timbre, pitch, and audio distortion of the audio channel; and

using the machine learning engine to select a speech recognition engine that is qualified as capable of transcribing the audio channel and/or its parts or a transcoded version of the audio channel and/or its parts based a combination of at least three of the following characteristics:

audio fidelity;

background noise;

concurrent speech by multiple users;

timbre;

pitch; and

audio distortion of the audio channel.

11. The computer-implemented method of claim 1 , wherein the selecting of a qualified speech recognition engine is based on a combination of least four of the following characteristics:

audio fidelity;

background noise;

concurrent speech by multiple users;

timbre;

pitch; and

audio distortion of the audio channel.

12. The computer-implemented method of claim 11 , further including:

using the multiple speech recognition engines on the audio channel, including using the speech recognition engines sequentially when a first speech recognition engine reports a low confidence score on some or all of its transcription.

13. The computer-implemented method of claim 11 , further including:

using the multiple speech recognition engines on all or separate parts of the audio channel, including using the speech recognition engines when voting on transcription results is used, when different speakers on different tracks of the audio channel, and when different speakers take turns during segments of the audio channel.

14. The computer-implemented method of claim 11 , further including applying the method to auto-punctuation engines.

15. The computer-implemented method of claim 11 , further including applying a silence analyzer to the speech samples and the audio channel prior to submission to parse out silent parts of speech.

16. The computer-implemented method of claim 11 , further including performing testing periodically, including daily, weekly, or monthly, wherein the testing comprises the submitting of hundreds of speech samples to multiple speech recognition engines, and the analyzing of how transcription error rates of the speech recognition engines vary with sound and speech characteristics of the speech samples.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2022
From: SIMON SAYS, INC.
To: FACEBOOK, INC.
Reel/Frame 059494/0913 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2022
From: FACEBOOK, INC.
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 059392/0564 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2019
From: ALLIBHAI, SHAMIR
To: SIMON SAYS, INC.
Reel/Frame 048931/0676 →
Cited By (2)
US 12,273,573 US 12,367,859