IP Library Granted Patent US 10,971,153
Granted Patent B2
US 10,971,153 · App. 16/749,970 · Granted Apr 6, 2021

Transcription generation from multiple speech recognition systems

Inventors: David Thomson (North Salt Lake, UT); Jadie Adams (Salt Lake City, UT); Jonathan Skaggs (Provo, UT); Joshua McClellan (Silver Spring, MD); Shane Roylance (Farmington, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/22G10L15/187G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,971,153
App. No.
16/749,970
Granted
Apr 6, 2021
Kind
B2
Abstract

A method may include obtaining first audio data originating at a first device during a communication session between the first device and a second device. The method may also include obtaining a first text string that is a transcription of the first audio data, where the first text string may be generated using automatic speech recognition technology using the first audio data. The method may also include obtaining a second text string that is a transcription of second audio data, where the second audio data may include a revoicing of the first audio data by a captioning assistant and the second text string may be generated by the automatic speech recognition technology using the second audio data. The method may further include generating an output text string from the first text string and the second text string and using the output text string as a transcription of the speech.

Claims (42)

1. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device;

obtaining a first text string that is a transcription of the first audio data, the first text string generated by a first automatic speech recognition system using a first speech recognition model;

obtaining a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data and the second text string generated by a second automatic speech recognition system using a second speech recognition model;

before any part of the first text string and any part of the second text string is provided to the second device, generating an output text string from the first text string and the second text string, the output text string includes one or more first words from the first text string and one or more second words from the second text string such that the output text string does not include an entirety of the first text string or an entirety of the second text string; and

providing the output text string, without providing the first text string and the second text string, to the second device for presentation during the communication session.

2. The method of claim 1 , wherein the first speech recognition model includes one or more of the following: a feature model, a transform model, an acoustic model, a language model, and a pronunciation model.

3. The method of claim 1 , wherein generating the output text string further includes:

de-normalizing the first text string and the second text string;

aligning the first text string and the second text string; and

comparing the aligned and de-normalized first and second text strings.

4. The method of claim 1 , wherein generating the output text string includes:

selecting the one or more second words based on the first text string and the second text string both including the one or more second words; and

selecting the one or more first words from the first text string based on the second text string not including the one or more first words.

5. The method of claim 1 , wherein generating the output text string includes:

selecting the one or more first words based on the first text string and the second text string both including the one or more first words; and

selecting the one or more second words from the second text string based on the first text string not including the one or more second words.

6. The method of claim 1 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on a third text string generated by the first automatic speech recognition system using the first audio data.

7. The method of claim 6 , wherein the first text string and the third text string are both hypotheses generated by the first automatic speech recognition system for the substantially same portion of the first audio data.

8. The method of claim 1 , further comprising obtaining a third text string that is a transcription of the first audio data or the second audio data, the third text string generated using a third speech recognition model, wherein the output text string is generated from the first text string, the second text string, and the third text string.

9. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 1 .

10. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device;

obtaining a first text string that is a transcription of the first audio data, the first text string generated using a first automatic speech recognition technology;

obtaining a second text string that is a transcription of the first audio data, the second text string generated using a second automatic speech recognition technology; and

generating an output text string using the first text string in response to the first text string being generated by the first automatic speech recognition technology and without regard to a quality measure of the first text string, wherein in response to the second text string having a quality measure satisfying a quality threshold, the output text string is generated using the first text string and the second text string instead of generating the output text string using only the first text string.

11. The method of claim 10 , wherein the first automatic speech recognition technology includes a first speech recognition model and the second automatic speech recognition technology includes a second speech recognition model that is different from the first speech recognition model.

12. The method of claim 11 , further comprising obtaining a third text string that is a transcription of the first audio data, the third text string generated using a third speech recognition model that is different from the first speech recognition model and the second speech recognition model, wherein the output text string is generated from the first text string, the second text string, and the third text string.

13. The method of claim 10 , wherein generating the output text string includes:

selecting one or more second words from the second text string for inclusion in the output text string based on the first text string and the second text string both including the one or more second words; and

selecting one or more first words from the first text string for inclusion in the output text string based on the second text string not including the one or more first words.

14. The method of claim 10 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on a third text string generated by the first automatic speech recognition technology using the first audio data.

15. The method of claim 14 , wherein the first text string and the third text string are both hypotheses generated by the first automatic speech recognition technology for the substantially same portion of the first audio data.

16. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 10 .

17. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device;

obtaining a first text string that is a transcription of the first audio data, the first text string generated by a first automatic speech recognition system using a first speech recognition model;

obtaining a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data and the second text string generated by a second automatic speech recognition system using a second speech recognition model; and

in response to the second text string having a quality measure satisfying a quality threshold, generating an output text string using the first text string and the second text string instead of generating the output text string using only the first text string in response to the quality measure not satisfying the quality threshold.

18. The method of claim 17 , wherein generating the output text string using the first text string and the second text string results in the output text string including one or more first words from the first text string and one or more second words from the second text string such that the output text string does not include an entirety of the first text string or an entirety of the second text string.

19. The method of claim 17 , wherein the first speech recognition model is trained for a plurality of individuals and the second speech recognition model is trained for a captioning assistant that revoices the first audio data.

20. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 17 .

Assignments (6)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 051597/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: HOLM, MICHAEL; BLACK, DAVID; BAROCIO, JESSE; THOMSON, DAVID; BOEKWEG, SCOTT; ROYLANCE, SHANE; CLEMENTS, KIERSTEN; BOEHME, KENNETH; ADAMS, JADIE; SKAGGS, JONATHAN; ORZECHOWSKI, GRZEGORZ; MCCLELLAN, JOSHUA
To: CAPTIONCALL, LLC
Reel/Frame 051597/0020 →