IP Library Granted Patent US 11,017,778
Granted Patent B1
US 11,017,778 · App. 16/209,594 · Granted May 25, 2021

Switching between speech recognition systems

Inventors: David Thomson (North Salt Lake, UT); David Black (West Valley City, UT); Jonathan Skaggs (Provo, UT); Kenneth Boehme (South Jordan, UT); Shane Roylance (Farmington, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/32G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,778
App. No.
16/209,594
Granted
May 25, 2021
Kind
B1
Abstract

A method may include obtaining first audio data originating at a first device during a communication session between the first device and a second device. The method may also include obtaining an availability of revoiced transcription units in a transcription system and in response to establishment of the communication session, selecting, based on the availability of revoiced transcription units, a revoiced transcription unit instead of a non-revoiced transcription unit to generate a transcript of the first audio data. The method may also include obtaining revoiced audio generated by a revoicing of the first audio data by a captioning assistant and generating a transcription of the revoiced audio using an automatic speech recognition system. The method may further include in response to selecting the revoiced transcription unit, directing the transcription of the revoiced audio to the second device as the transcript of the first audio data.

Claims (49)

1. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication;

obtaining an availability of revoiced transcription units in a transcription system;

in response to establishment of the communication session, selecting, based on the availability of revoiced transcription units, a revoiced transcription unit instead of a non-revoiced transcription unit to generate a transcript of the first audio data to direct to the second device such that only the transcript of the first audio data generated by the selected revoiced transcription unit is directed to the second device during the communication session and the non-selected non-revoiced transcription unit does not generate a transcript of the first audio data of the communication session;

obtaining, by the revoiced transcription unit, revoiced audio generated by a revoicing of the first audio data by a captioning assistant;

generating, by the revoiced transcription unit, a transcription of the revoiced audio using an automatic speech recognition system;

in response to selecting the revoiced transcription unit, directing the transcription of the revoiced audio to the second device as the transcript of the first audio data;

obtaining second audio data originating at a third device during a second communication session between the third device and the second device;

obtaining a second availability of the revoiced transcription units in the transcription system;

in response to establishment of the second communication session, selecting, based on the second availability of the revoiced transcription units, the non-revoiced transcription unit instead of one of the revoiced transcription units to generate a transcript of the second audio data to direct to the second device such that none of the revoiced transcription units generate a transcript of the second audio data; and

generating, by the selected non-revoiced transcription unit, a transcription of the second audio data.

2. The method of claim 1 , wherein the availability of revoiced transcription units is based on one or more of: a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

3. The method of claim 1 , wherein the availability of revoiced transcription units is based on four or more of: a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

4. The method of claim 1 , wherein the automatic speech recognition system is trained specifically for speech of the captioning assistant.

5. The method of claim 1 , wherein the automatic speech recognition system is adapted for speech of the captioning assistant and a second automatic speech recognition system used by the non-revoiced transcription unit is trained for a plurality of speakers.

6. The method of claim 1 , wherein the availability of revoiced transcription units is based on a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

7. The method of claim 1 , wherein the transcription system includes the non-revoiced transcription unit and one or more additional non-revoiced transcription units.

8. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 1 .

9. A system comprising:

at least one computing system;

at least one computer-readable media coupled to the at least one computing system, the at least one computer-readable media configured to store one or more instructions that in response to being executed by the at least one computing system cause performance of operations, the operations comprising:

obtain first audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication;

obtain an availability of revoiced transcription units in a transcription system;

in response to establishment of the communication session, select, based on the availability of revoiced transcription units in the transcription system, a revoiced transcription unit instead of a non-revoiced transcription unit to generate a transcript of the first audio data that is directed to the second device and no non-selected non-revoiced transcription unit generates a transcript of the first audio data of the communication session;

in response to selecting the revoiced transcription unit, obtain a transcription of revoiced audio generated using an automatic speech recognition system of the revoiced transcription unit, the revoiced audio generated by a revoicing of the first audio data by a captioning assistant associated with the revoiced transcription unit;

direct the transcription of the revoiced audio to the second device as the transcript of the first audio data;

obtain second audio data originating at a third device during a second communication session between the third device and the second device;

obtain a second availability of the revoiced transcription units in the transcription system;

in response to establishment of the second communication session, select, based on the second availability of the revoiced transcription units, the non-revoiced transcription unit instead of one of the revoiced transcription units to generate a transcript of the second audio data to direct to the second device such that none of the revoiced transcription units generate a transcript of the second audio data; and

obtain a transcription of the second audio data generated by the selected non- revoiced transcription unit.

10. The system of claim 9 , wherein the availability of revoiced transcription units is based on one or more of: a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

11. The system of claim 9 , wherein the availability of revoiced transcription units is based on four or more of: a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

12. The system of claim 9 , wherein the automatic speech recognition system is adapted for speech of the captioning assistant.

13. The system of claim 9 , wherein the automatic speech recognition system is trained specifically for speech of the captioning assistant and a second automatic speech recognition system used by the non-revoiced transcription unit is trained for a plurality of speakers.

14. The system of claim 9 , wherein the availability of revoiced transcription units is based on a current peak number of transcriptions being generated, a current average number of transcriptions being generated, a projected peak number of transcriptions to be generated, a projected average number of transcriptions to be generated, a projected number of revoiced transcription units, and a number of available revoiced transcription units.

15. The system of claim 9 , wherein the system is included in the transcription system.

16. The system of claim 9 , wherein the transcription system includes the non-revoiced transcription unit and one or more additional non-revoiced transcription units.

17. A method comprising:

obtaining first audio data originating at a first device during a communication session between the first device and a second device;

obtaining an availability of revoiced transcription units in a transcription system;

in response to establishment of the communication session, selecting, based on the availability of revoiced transcription units in the transcription system, a revoiced transcription unit instead of a non-revoiced transcription unit to generate a transcript of the first audio data that is directed to the second device and the non-selected non-revoiced transcription unit does not generate a transcript of the first audio data of the communication session;

in response to selecting the revoiced transcription unit, obtaining a transcription of revoiced audio generated using an automatic speech recognition system of the revoiced transcription unit, the revoiced audio generated by a revoicing of the first audio data;

directing the transcription of the revoiced audio to the second device as the transcript of the first audio data;

obtaining second audio data originating at a third device during a second communication session between the third device and the second device;

obtaining a second availability of the revoiced transcription units in the transcription system;

in response to establishment of the second communication session, selecting, based on the second availability of the revoiced transcription units, the non-revoiced transcription unit instead of one of the revoiced transcription units to generate a transcript of the second audio data to direct to the second device such that none of the revoiced transcription units generate a transcript of the second audio data; and

generating, by the selected non-revoiced transcription unit, a transcription of the second audio data.

18. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 17 .

19. The method of claim 17 , wherein the transcription system includes the non-revoiced transcription unit and one or more additional non-revoiced transcription units.

Assignments (11)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2021
From: CORTLAND CAPITAL MARKET SERVICES LLC
To: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 058533/0467 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
LIEN Recorded Feb 11, 2020
From: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 051894/0665 →
RELEASE OF SECURITY INTEREST Recorded May 8, 2019
From: U.S. BANK NATIONAL ASSOCIATION
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049115/0468 →
RELEASE OF SECURITY INTEREST Recorded May 7, 2019
From: JPMORGAN CHASE BANK, N.A.
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049109/0752 →
PATENT SECURITY AGREEMENT Recorded Apr 29, 2019
From: SORENSEN COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 050084/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2019
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 047896/0017 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2018
From: HOLM, MICHAEL; BLACK, DAVID; BAROCIO, JESSE; THOMSON, DAVID; BOEKWEG, SCOTT; ROYLANCE, SHANE; CLEMENTS, KIERSTEN; BOEHME, KENNETH; ADAMS, JADIE; SKAGGS, JONATHAN; ORZECHOWSKI, GRZEGORZ; MCCLELLAN, JOSHUA
To: CAPTIONCALL, LLC
Reel/Frame 047698/0839 →
Cited By (27)
US 12,190,870 US 12,243,534 US 12,284,059 US 12,354,599 US 12,373,647 US 12,400,660 US 12,400,661 US 12,406,668 US 12,406,672 US 12,406,684 US 12,407,777 US 12,437,763 US 12,456,465 US 12,462,808 US 12,475,887 US 12,482,458 US 12,488,799 US 12,494,929 US 12,512,097 US 12,518,748 US 12,555,581 US 12,555,593 US 12,567,416 US 12,567,418 US 12,578,922 US 12,579,982 US 12,658,185