IP Library Granted Patent US 10,580,410
Granted Patent B2
US 10,580,410 · App. 15/964,661 · Granted Mar 3, 2020

Transcription of communications

Inventor: Craig Nelson (Midvale, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/26G06K9/00288H04M3/2218
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,580,410
App. No.
15/964,661
Granted
Mar 3, 2020
Kind
B2
Abstract

A system may include a camera configured to obtain an image of a user, at least one processor, and at least one non-transitory computer-readable media communicatively coupled to the at least one processor. The non-transitory computer-readable media configured to store one or more instructions that when executed cause or direct the system to perform operations. The operations may include establish a communication session between the system and a device. The communication session may be configured such that the device provides audio for the system. The operations may further include compare the image to a particular user image associated with the system and select a first method of transcription generation from among two or more methods of transcription generation based on the comparison of the image to the particular user image. The operations may also include present, a transcription of the audio generated using the selected first method of transcription generation.

Claims (70)

1. A method to transcribe communications, the method comprising:

obtaining a first image of a first user associated with a first device;

obtaining audio during a communication session between the first device and a second device, the audio originating at the second device and being based on speech of a second user of the second device;

comparing the first image of the first user associated with the first device to a particular user image associated with the first device;

in response to the first image matching the particular user image, presenting, to the first user associated with the first device, a first transcription of a first portion of the audio originating at the second device;

after presenting the first transcription of the audio and during the communication session, obtaining, by the first device, a second image;

comparing the second image to the particular user image;

in response to the second image not matching the particular user image:

ceasing the presentation of the first transcription of the first portion of the audio originating at the second device; and

obtaining a request to present a second transcription of a second portion of the audio originating at the second device; and

in response to the request to present the second transcription, presenting the second transcription.

2. The method of claim 1 , further comprising in response to the first image matching the particular user image:

providing the first portion of the audio to a transcription system configured to generate the first transcription; and

obtaining the first transcription from the transcription system.

3. The method of claim 2 , further comprising:

in response to the second image not matching the particular user image, ceasing to provide the first portion of the audio to the transcription system; and

in response to the request to present the second transcription:

providing the second portion of the audio to the transcription system; and

obtaining the second transcription from the transcription system.

4. The method of claim 1 , wherein the first transcription is generated using a first speech recognition system and the second transcription is generated using a second speech recognition system different than the first speech recognition system.

5. The method of claim 4 , wherein the first speech recognition system and the second speech recognition system are each configured to use a different speech recognition engine to automatically recognize speech in audio independent of human interaction.

6. The method of claim 4 , wherein the first speech recognition system is a re-voicing speech recognition system and the second speech recognition system automatically recognizes speech independent of human interaction.

7. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 1 .

8. A method to transcribe communications, the method comprising:

obtaining first audio of a first communication session between a first device and a second device, the first communication session configured such that the first audio originates at the second device and the second device provides the first audio for the first device;

obtaining a first image of a first user associated with the first device;

comparing the first image to a particular user image associated with the first device;

in response to the first image matching the particular user image, presenting, to the first user, a first transcription of the first audio, the first transcription generated using a first speech recognition system that includes a first speech engine trained to automatically recognize speech in audio;

after presenting the first transcription, obtaining second audio of a second communication session between the first device and a third device, the second communication session configured such that the second audio originates at the third device and the third device provides the second audio for the first device;

after presenting the first transcription, obtaining a second image of a second user associated with the first device;

comparing the second image to the particular user image; and

in response to the second image not matching the particular user image, presenting, to the second user, a second transcription of the second audio, the second transcription generated using a second speech recognition system that includes a second speech engine trained to automatically recognize speech in audio, the second speech recognition system being different than the first speech recognition system, including the first speech engine being different than the second speech engine.

9. The method of claim 8 , further comprising:

after presenting the first transcription and during the first communication session, obtaining a third image;

comparing the third image to the particular user image;

in response to the third image not matching the particular user image:

ceasing the presentation of the first transcription; and

presenting a third transcription of the first audio, the third transcription generated using the second speech recognition system.

10. The method of claim 8 , further comprising:

after presenting the first transcription and during the first communication session, obtaining a third image;

comparing the third image to the particular user image;

in response to the third image not matching the particular user image:

ceasing the presentation of the first transcription; and

obtaining, at the first device, a request to present a third transcription of the first audio; and

in response to the request to present the third transcription, presenting the third transcription.

11. The method of claim 10 , wherein the third transcription is generated by the second speech recognition system.

12. The method of claim 8 , wherein the first and second speech recognition systems both automatically recognize speech independent of human interaction.

13. The method of claim 8 , wherein the first speech recognition system is configured to:

broadcasting the second audio; and

obtaining third audio based on a human re-voicing of the broadcasted second audio, wherein the first transcription is generated by the first speech engine based on the third audio.

14. The method of claim 8 , wherein the first speech recognition system is a re-voicing speech recognition system and the second speech recognition system automatically recognizes speech independent of human interaction.

15. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 8 .

16. A system comprising:

a camera configured to obtain an image of a user associated with the system;

at least one processor coupled to the camera and configured to receive the image from the camera; and

at least one non-transitory computer-readable media communicatively coupled to the at least one processor and configured to store one or more instructions that when executed by the at least one processor cause or direct the system to perform operations comprising:

establish a communication session between the system and a device, the communication session configured such that the device provides audio for the system, wherein the image of the user is obtained after the communication session is established;

compare the image to a particular user image associated with the system;

select a first speech recognition system from among two or more speech recognition systems based on the comparison of the image to the particular user image, each of the two or more speech recognition systems including a different speech engine trained to automatically recognize speech in audio; and

present, to the user, a transcription of the audio, the transcription of the audio generated using the selected first speech recognition system.

17. The system of claim 16 , wherein the transcription of the audio is a first transcription of a first portion of the audio, the operations further comprising:

after presenting the first transcription of the audio and during the communication session, obtain a second image;

compare the second image to the particular user image;

select a second speech recognition system from among the two or more speech recognition systems based on the comparison of the second image to the particular user image; and

present, to the user, a second transcription of a second portion of the audio, the second transcription generated using the selected second speech recognition system.

18. The system of claim 16 , wherein the first speech recognition system is selected from among two or more speech recognition systems based on the image not matching the particular user image and the first speech recognition system automatically recognizes speech independent of human interaction.

19. The system of claim 16 , wherein the first speech recognition system is selected from among two or more speech recognition systems based on the image matching the particular user image and the first speech recognition system is configured to:

broadcasting the audio; and

obtaining second audio based on a human re-voicing of the broadcasted audio, wherein the transcription is generated based on the second audio.

20. The system of claim 16 , wherein the two or more speech recognition systems include at least: a speech recognition system that automatically recognizes speech independent of human interaction and a re-voicing speech recognition system.

Assignments (10)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2021
From: CORTLAND CAPITAL MARKET SERVICES LLC
To: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 058533/0467 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
LIEN Recorded Feb 11, 2020
From: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CORTLAND CAPITAL MARKET SERVICES LLC
Reel/Frame 051894/0665 →
RELEASE OF SECURITY INTEREST Recorded May 7, 2019
From: JPMORGAN CHASE BANK, N.A.
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049109/0752 →
PATENT SECURITY AGREEMENT Recorded Apr 29, 2019
From: SORENSEN COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 050084/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2018
From: NELSON, CRAIG
To: CAPTIONCALL, LLC
Reel/Frame 045865/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2018
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 045865/0239 →