IP Library Granted Patent US 10,573,312
Granted Patent B1
US 10,573,312 · App. 16/209,623 · Granted Feb 25, 2020

Transcription generation from multiple speech recognition systems

Inventors: David Thomson (North Salt Lake, UT); Jadie Adams (Salt Lake City, UT); Jonathan Skaggs (Provo, UT); Joshua McClellan (Silver Spring, MD); Shane Roylance (Farmington, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/22G10L15/187G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,573,312
App. No.
16/209,623
Granted
Feb 25, 2020
Kind
B1
Abstract

A method may include obtaining first audio data originating at a first device during a communication session between the first device and a second device. The method may also include obtaining a first text string that is a transcription of the first audio data, where the first text string may be generated using automatic speech recognition technology using the first audio data. The method may also include obtaining a second text string that is a transcription of second audio data, where the second audio data may include a revoicing of the first audio data by a captioning assistant and the second text string may be generated by the automatic speech recognition technology using the second audio data. The method may further include generating an output text string from the first text string and the second text string and using the output text string as a transcription of the speech.

Claims (57)

1. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the first audio data includes speech;

obtaining a first text string that is a transcription of the first audio data, the first text string generated by a first automatic speech recognition system using the first audio data and using a first model trained for a plurality of individuals;

obtaining a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data by a captioning assistant and the second text string generated by a second automatic speech recognition system using the second audio data and using a second model trained for the captioning assistant;

obtaining a third text string that is a transcription of the first audio data or the second audio data, the third text string generated by a third automatic speech recognition system using a third model;

generating an output text string from the first text string, the second text string, and the third text string; and

providing the output text string as a transcription of the speech to the second device for presentation during the communication session concurrently with the presentation of the first audio data by the second device.

2. The method of claim 1 , wherein the first model includes one or more of the following: a feature model, a transform model, an acoustic model, a language model, and a pronunciation model.

3. The method of claim 1 , wherein generating the output text string further includes:

de-normalizing the first text string and the second text string;

aligning the first text string and the second text string; and

comparing the aligned and de-normalized first and second text strings.

4. The method of claim 1 , wherein generating the output text string further includes:

selecting one or more second words from the second text string for the output text string based on the first text string and the second text string both including the one or more second words; and

selecting one or more first words from the first text string for the output text string based on the second text string not including the one or more first words.

5. The method of claim 1 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on input obtained from a device associated with the captioning assistant.

6. The method of claim 5 , wherein the input obtained from the device is based on a fourth text string generated by the first automatic speech recognition system using the first audio data.

7. The method of claim 6 , wherein the first text string and the fourth text string are both hypothesis generated by the first automatic speech recognition system for the substantially same portion of the first audio data.

8. The method of claim 1 , wherein the third text string is a transcription of the first audio data, the method further comprising obtaining a fourth text string that is a transcription of the second audio data, the fourth text string generated by a fourth automatic speech recognition system using the second audio data and using a fourth model, wherein the output text string is generated from the first text string, the second text string, the third text string, and the fourth text string.

9. The method of claim 1 , further comprising:

obtaining fourth audio data that includes speech and that originates at the first device during the communication session;

obtaining a third text string that is a transcription of the fourth audio data, the fourth text string generated by the first automatic speech recognition system using the fourth audio data and using the first model; and

in response to either no revoicing of the fourth audio data or a fifth transcription generated using the second automatic speech recognition system having a quality measure satisfying a quality threshold, generating a second output text string using only the fourth text string.

10. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 1 .

11. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the first audio data includes speech;

obtaining a first text string that is a transcription of the first audio data, the first text string generated using automatic speech recognition technology using the first audio data;

obtaining a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data by a captioning assistant and the second text string generated by the automatic speech recognition technology using the second audio data;

obtaining a third text string that is a transcription of the first audio data or the second audio data, the third text string generated by the automatic speech recognition technology;

generating an output text string from the first text string, the second text string, and the third text string; and

using the output text string as a transcription of the speech.

12. The method of claim 11 , wherein the automatic speech recognition technology used to generate the first text string is a first automatic speech recognition system that includes a first model trained for a plurality of individuals and the automatic speech recognition technology used to generate the second text string is a second automatic speech recognition system that includes a second model adapted to the captioning assistant.

13. The method of claim 11 , wherein the output text string includes one or more first words from the first text string and one or more second words from the second text string.

14. The method of claim 11 , further comprising correcting at least one word in one or more of: the output text string, the first text string, and the second text string based on input obtained from a device associated with the captioning assistant.

15. The method of claim 14 , wherein the input obtained from the device is based on a fourth text string generated by the automatic speech recognition technology using the first audio data.

16. The method of claim 15 , wherein the first text string and the fourth text string are both hypothesis generated by the automatic speech recognition technology for the substantially same portion of the first audio data.

17. The method of claim 11 , further comprising:

obtaining third audio data that includes speech and that originates at the first device during the communication session;

obtaining a fourth text string that is a transcription of the third audio data, the fourth text string generated by the automatic speech recognition technology using the third audio data; and

in response to either no revoicing of the third audio data or a fourth transcription, generated using the automatic speech recognition technology and revoicing of the third audio data, having a quality measure satisfying a quality threshold, generating a second output text string using only the fourth text string.

18. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 11 .

19. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the first audio data includes speech;

obtaining a first text string that is a transcription of the first audio data, the first text string generated using automatic speech recognition technology using the first audio data;

obtaining a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data and the second text string generated by the automatic speech recognition technology using the second audio data;

obtaining a third text string that is a transcription of the first audio data or the second audio data, the third text string generated by the automatic speech recognition technology;

generating an output text string from the first text string, the second text string, and the third text string, the output text string including one or more words based on at least two of the first text string, the second text string, and the third text string including the one or more words; and

providing the output text string as a transcription of the speech to the second device for presentation during the communication session by the second device.

20. A system comprising:

one or more processors; and

at least one non-transitory computer-readable media coupled to the one or more processors, the at least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by the one or more processors cause the system to perform operations, the operations comprising:

obtain first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication such that the first audio data includes speech;

obtain a first text string that is a transcription of the first audio data, the first text string generated using automatic speech recognition technology using the first audio data;

obtain a second text string that is a transcription of second audio data, the second audio data including a revoicing of the first audio data and the second text string generated by the automatic speech recognition technology using the second audio data;

obtain a third text string that is a transcription of the first audio data or the second audio data, the third text string generated by the automatic speech recognition technology;

generate an output text string from the first text string, the second text string, and the third text string; and

provide the output text string as a transcription of the speech.

Assignments (11)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2021
From: CORTLAND CAPITAL MARKET SERVICES LLC
To: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 058533/0467 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
LIEN Recorded Feb 11, 2020
From: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 051894/0665 →
RELEASE OF SECURITY INTEREST Recorded May 8, 2019
From: U.S. BANK NATIONAL ASSOCIATION
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049115/0468 →
RELEASE OF SECURITY INTEREST Recorded May 7, 2019
From: JPMORGAN CHASE BANK, N.A.
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049109/0752 →
PATENT SECURITY AGREEMENT Recorded Apr 29, 2019
From: SORENSEN COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 050084/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2019
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 047896/0013 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2018
From: HOLM, MICHAEL; BLACK, DAVID; BAROCIO, JESSE; THOMSON, DAVID; BOEKWEG, SCOTT; ROYLANCE, SHANE; CLEMENTS, KIERSTEN; BOEHME, KENNETH; ADAMS, JADIE; SKAGGS, JONATHAN; ORZECHOWSKI, GRZEGORZ; MCCLELLAN, JOSHUA
To: CAPTIONCALL, LLC
Reel/Frame 047698/0893 →
Cited By (82)
US 50,388 US 50,834 US 12,190,870 US 12,211,490 US 12,212,945 US 12,217,748 US 12,217,758 US 12,217,765 US 12,230,291 US 12,236,932 US 12,236,955 US 12,248,748 US 12,249,313 US 12,271,695 US 12,278,925 US 12,283,269 US 12,284,059 US 12,288,558 US 12,306,809 US 12,308,030 US 12,314,263 US 12,314,633 US 12,314,672 US 12,321,853 US 12,322,390 US 12,323,554 US 12,327,549 US 12,327,556 US 12,334,054 US 12,341,621 US 12,360,734 US 12,373,027 US 12,375,052 US 12,380,896 US 12,387,716 US 12,400,661 US 12,406,672 US 12,406,684 US 12,412,583 US 12,424,220 US 12,430,569 US 12,436,994 US 12,438,977 US 12,456,465 US 12,462,802 US 12,462,808 US 12,468,883 US 12,469,491 US 12,475,883 US 12,481,841 US 12,488,785 US 12,494,929 US 12,498,899 US 12,505,832 US 12,511,098 US 12,511,498 US 12,513,479 US 12,518,748 US 12,518,755 US 12,518,756 US 12,525,242 US 12,530,111 US 12,530,112 US 12,555,581 US 12,561,536 US 12,562,152 US 12,567,416 US 12,567,418 US 12,579,978 US 12,579,982 US 12,608,542 US 12,619,818 US 12,626,717 US 12,645,874 US 12,652,260 US 12,658,185 US 12,670,911 US 12,675,484 US 12,682,187 US 12,699,543 US 12,706,092 US 12,711,962