IP Library Granted Patent US 10,325,597
Granted Patent B1
US 10,325,597 · App. 16/154,553 · Granted Jun 18, 2019

Transcription of communications

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,325,597
App. No.
16/154,553
Granted
Jun 18, 2019
Kind
B1
Abstract

A method to transcribe communications may include obtaining audio data originating at a first device during a communication session between the first device and a second device and providing the audio data to an automated speech recognition system configured to transcribe the audio data. The method may further include obtaining multiple hypothesis transcriptions generated by the automated speech recognition system. Each of the multiple hypothesis transcriptions may include one or more words determined by the automated speech recognition system to be a transcription of a portion of the audio data. The method may further include determining one or more consistent words that are included in two or more of the multiple hypothesis transcriptions and in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device.

Claims (57)

1. A method to transcribe communications, the method comprising:

obtaining audio data originating at a first device during a communication session between the first device and a second device, the communication session configured for verbal communication;

providing the audio data to an automated speech recognition system configured to transcribe the audio data;

before a final transcription of the audio data is determined by the automated speech recognition system, the method including:

obtaining a first hypothesis transcription generated by the automated speech recognition system, the first hypothesis transcription including one or more first words determined by the automated speech recognition system to be a transcription of at least a first portion of the audio data;

obtaining a second hypothesis transcription generated by the automated speech recognition system, the second hypothesis transcription including a plurality of second words determined by the automated speech recognition system to be a transcription of at least a second portion of the audio data that includes the first portion of the audio data, a number of the plurality of second words being greater than a number of the one or more first words;

determining one or more consistent words that are included in both the one or more first words of the first hypothesis transcription and the plurality of second words of the second hypothesis transcription; and

in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before the final transcription of the audio data is provided to the second device.

2. The method of claim 1 , wherein the first hypothesis transcription and the second hypothesis transcription are not provided to the second device.

3. The method of claim 1 , wherein the audio data is a portion of total audio data originating at the first device during the communication session.

4. The method of claim 1 , further comprising:

obtaining a third hypothesis transcription generated by the automated speech recognition system, the third hypothesis transcription including a plurality of third words determined by the automated speech recognition system to be a transcription of at least a third portion of the audio data that includes the first portion and the second portion of the audio data, a number of the plurality of third words being greater than a number of the plurality of second words;

determining one or more second consistent words that are included in both the plurality of second words of the second hypothesis transcription and the plurality of third words of the third hypothesis transcription; and

in response to determining the one or more second consistent words, providing the one or more second consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before the final transcription of the audio data is provided to the second device.

5. The method of claim 1 , further comprising:

obtaining the final transcription of the audio data, the final transcription including a plurality of third words that the automated speech recognition system outputs together as the finalized transcription of the audio data;

determining when a portion of the final transcription that corresponds to the one or more consistent words includes an update word that is different from any of the one or more consistent words; and

in response to determining the update word, providing an indication of the update word to the second device such that the second device changes the presentation of the one or more consistent words to include the update word.

6. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 1 .

7. A system comprising:

at least one processor; and

at least one non-transitory computer-readable media communicatively coupled to the at least one processor and configured to store one or more instructions that when executed by the at least one processor cause or direct the system to perform operations comprising:

obtain audio data originating at a first device during a communication session between the first device and a second device;

provide the audio data to an automated speech recognition system configured to transcribe the audio data;

obtain a plurality of hypothesis transcriptions generated by the automated speech recognition system, each of the plurality of hypothesis transcriptions including one or more words determined by the automated speech recognition system to be a transcription of a portion of the audio data;

determine one or more consistent words that are included in two or more of the plurality of hypothesis transcriptions; and

in response to determining the one or more consistent words, provide the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before a final transcription of the audio data is provided to the second device.

8. The system of claim 7 , wherein the plurality of hypothesis transcriptions are not provided to the second device.

9. The system of claim 7 , wherein the automated speech recognition system is included in the system.

10. The system of claim 7 , wherein the plurality of hypothesis transcriptions are obtained sequentially over time and a first portion of the audio data associated with a first one of the plurality of hypothesis transcriptions includes all of the audio data associated with all of the plurality of hypothesis transcriptions obtained previous to obtaining the first one of the plurality of hypothesis transcriptions.

11. The system of claim 7 , wherein the plurality of hypothesis transcriptions are obtained sequentially over time and the operation to determine the one or more consistent words includes compare a first hypothesis transcription of the plurality of hypothesis transcriptions with a second hypothesis transcription of the plurality of hypothesis transcriptions, the second hypothesis transcription directly following the first hypothesis transcription among the plurality of hypothesis transcriptions.

12. The system of claim 7 , wherein the operations further comprise:

obtain a second hypothesis transcription generated by the automated speech recognition system after generation of the plurality of hypothesis transcriptions, the second hypothesis transcription including a plurality of second words determined by the automated speech recognition system to be a transcription of at least a second portion of the audio data;

determine one or more second consistent words that are included in both the plurality of second words of the second hypothesis transcription and the one or more words of one of the plurality of hypothesis transcriptions; and

in response to determining the one or more second consistent words, provide the one or more second consistent words to the second device for presentation of the one or more second consistent words by the second device.

13. The system of claim 7 , wherein the operations further comprise:

obtain the final transcription of the audio data, the final transcription including a plurality of words that the automated speech recognition system outputs together as the final transcription of the audio data;

determine when a portion of the final transcription that corresponds to the one or more consistent words includes an update word that is different from any of the one or more consistent words; and

in response to determining the update word, provide an indication of the update word to the second device such that the second device changes the presentation of the one or more consistent words to include the update word.

14. A method to transcribe communications, the method comprising:

obtaining audio data originating at a first device during a communication session between the first device and a second device;

providing the audio data to an automated speech recognition system configured to transcribe the audio data;

obtaining a plurality of hypothesis transcriptions generated by the automated speech recognition system, each of the plurality of hypothesis transcriptions including one or more words determined by the automated speech recognition system to be a transcription of a portion of the audio data;

determining one or more consistent words that are included in two or more of the plurality of hypothesis transcriptions; and

in response to determining the one or more consistent words, providing the one or more consistent words to the second device for presentation of the one or more consistent words by the second device, the presentation of the one or more consistent words configured to occur before a final transcription of the audio data is provided to the second device.

15. The method of claim 14 , wherein the plurality of hypothesis transcriptions are not provided to the second device.

16. The method of claim 14 , wherein the plurality of hypothesis transcriptions are obtained sequentially over time and a first portion of the audio data associated with a first one of the plurality of hypothesis transcriptions includes all of the audio data associated with all of the plurality of hypothesis transcriptions obtained previous to obtaining the first one of the plurality of hypothesis transcriptions.

17. The method of claim 14 , wherein the plurality of hypothesis transcriptions are obtained sequentially over time and the determining the one or more consistent words includes comparing a first hypothesis transcription of the plurality of hypothesis transcriptions with a second hypothesis transcription of the plurality of hypothesis transcriptions, the second hypothesis transcription directly following the first hypothesis transcription among the plurality of hypothesis transcriptions.

18. The method of claim 14 , further comprising:

obtaining a second hypothesis transcription generated by the automated speech recognition system after generation of the plurality of hypothesis transcriptions, the second hypothesis transcription including a plurality of second words determined by the automated speech recognition system to be a transcription of at least a second portion of the audio data;

determining one or more second consistent words that are included in both the plurality of second words of the second hypothesis transcription and the one or more words of one of the plurality of hypothesis transcriptions; and

in response to determining the one or more second consistent words, providing the one or more second consistent words to the second device for presentation of the one or more second consistent words by the second device.

19. The method of claim 14 , further comprising:

obtaining the final transcription of the audio data, the final transcription including a plurality of words that the automated speech recognition system outputs together as the final transcription of the audio data;

determining when a portion of the final transcription that corresponds to the one or more consistent words includes an update word that is different from any of the one or more consistent words; and

in response to determining the update word, providing an indication of the update word to the second device such that the second device changes the presentation of the one or more consistent words to include the update word.

20. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 14 .

Assignments (11)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2021
From: CORTLAND CAPITAL MARKET SERVICES LLC
To: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 058533/0467 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
LIEN Recorded Feb 11, 2020
From: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 051894/0665 →
RELEASE OF SECURITY INTEREST Recorded May 8, 2019
From: U.S. BANK NATIONAL ASSOCIATION
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049115/0468 →
RELEASE OF SECURITY INTEREST Recorded May 7, 2019
From: JPMORGAN CHASE BANK, N.A.
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049109/0752 →
PATENT SECURITY AGREEMENT Recorded Apr 29, 2019
From: SORENSEN COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 050084/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2018
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 047367/0451 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2018
From: CHEVRIER, BRIAN; ROYLANCE, SHANE; BOEHME, KENNETH
To: CAPTIONCALL, LLC
Reel/Frame 047119/0136 →