IP Library Granted Patent US 11,170,761
Granted Patent B2
US 11,170,761 · App. 16/209,524 · Granted Nov 9, 2021

Training of speech recognition systems

Inventors: David Thomson (North Salt Lake, UT); Jadie Adams (Salt Lake City, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/063G10L15/26G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,761
App. No.
16/209,524
Granted
Nov 9, 2021
Kind
B2
Abstract

A method may include obtaining first audio data of a first communication session between a first and second device and during the first communication session, obtaining a first text string that is a transcription of the first audio data and training a model of an automatic speech recognition system using the first text string and the first audio data. The method may further include in response to completion of the training, deleting the first audio data and the first text string and after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third and fourth device and during the second communication session obtaining a second text string that is a transcription of the second audio data and further training the model of the automatic speech recognition system using the second text string and the second audio data.

Claims (37)

1. A method comprising:

obtaining first audio data of a first communication session between a first device of a first user and a second device of a second user, the first communication session configured for verbal communication;

obtaining, during the first communication session, a first text string that is a transcription of the first audio data;

training, during the first communication session, a model of an automatic speech recognition system using the first text string and the first audio data;

in response to completion of the training of the model using the first text string and the first audio data, deleting the first audio data and the first text string;

after training the model using the first text string and the first audio data, obtaining second audio data of a second communication session between a third device of a third user and a fourth device of a fourth user, wherein the third user and the fourth user are both separate and distinct from the first user and the second user;

generating, during the second communication session, a transcription of the second audio data by applying the model trained using the first text string and the first audio data; and

providing the transcription of the second audio data to the fourth device for presentation during the second communication session.

2. The method of claim 1 , wherein the model is an acoustic model, a language model, a confidence model, or classification model of the automatic speech recognition system.

3. The method of claim 1 , wherein the first text string is generated using automatic speech recognition technology.

4. The method of claim 3 , wherein the automatic speech recognition technology generates the first text string using a revoicing of the first audio data.

5. The method of claim 1 , wherein the first text string is generated from one or more words of a second text string and one or more words of a third text string, the second text string and the third text string generated by automatic speech recognition technology.

6. The method of claim 1 , wherein the training of the model of the automatic speech recognition system using the first text string and the first audio data completes after the first communication session.

7. The method of claim 1 , wherein the first audio data and the first text string are deleted during the first communication session.

8. The method of claim 1 , further comprising providing the transcription of the first audio data to the second device for presentation by the second device during the first communication session.

9. The method of claim 1 , further comprising:

training, during the second communication session using a second text string of the transcription of the second audio data and the second audio data, a second model used by automatic speech recognition technology; and

in response to completion of the training of the second model using the second text string and the second audio data, deleting the second audio data and the second text string.

10. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 1 .

11. A method comprising:

obtaining first audio data of a first communication session between a first device of a first user and a second device of a second user, the first communication session configured for verbal communication;

obtaining, during the first communication session, a first text string that is a transcription of the first audio data; training, during the first communication session, a model of an automatic speech recognition system using the first text string and the first audio data;

in response to completion of the training of the model using the first text string and the first audio data, deleting the first audio data and the first text string;

after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third device of a third user and a fourth device of a fourth user, wherein the third user and the fourth user are both separate and distinct from the first user and the second user;

obtaining, during the second communication session, a second text string that is a transcription of the second audio data; and

further training, during the second communication session, the model of the automatic speech recognition system using the second text string and the second audio data.

12. The method of claim 11 , wherein the model is an acoustic model, a language model, a confidence model, or classification model of the automatic speech recognition system.

13. The method of claim 11 , wherein the first text string is generated using automatic speech recognition technology.

14. The method of claim 13 , wherein the automatic speech recognition technology generates the first text string using a revoicing of the first audio data.

15. The method of claim 11 , wherein the first text string is generated from one or more words of a second text string and one or more words of a third text string, the second text string and the third text string generated by automatic speech recognition technology.

16. The method of claim 11 , wherein the training of the model of the automatic speech recognition system using the first text string and the first audio data completes after the first communication session.

17. The method of claim 11 , wherein the first audio data and the first text string are deleted during the first communication session.

18. The method of claim 11 , further comprising:

providing the transcription of the first audio data to the second device for presentation by the second device during the first communication session; and

providing the transcription of the second audio data to the fourth device for presentation by the fourth device during the second communication session.

19. The method of claim 11 , wherein the first audio data originates from the first device and is based on captured verbal communications of the first user during the first communication session.

20. At least one non-transitory computer-readable media configured to store one or more instructions that in response to being executed by at least one computing system cause performance of the method of claim 11 .

Assignments (10)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
RELEASE OF SECURITY INTEREST Recorded Dec 16, 2021
From: CORTLAND CAPITAL MARKET SERVICES LLC
To: SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 058533/0467 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
RELEASE OF SECURITY INTEREST Recorded May 8, 2019
From: U.S. BANK NATIONAL ASSOCIATION
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049115/0468 →
RELEASE OF SECURITY INTEREST Recorded May 7, 2019
From: JPMORGAN CHASE BANK, N.A.
To: SORENSON COMMUNICATIONS, LLC; SORENSON IP HOLDINGS, LLC; CAPTIONCALL, LLC; INTERACTIVECARE, LLC
Reel/Frame 049109/0752 →
PATENT SECURITY AGREEMENT Recorded Apr 29, 2019
From: SORENSEN COMMUNICATIONS, LLC; CAPTIONCALL, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 050084/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2019
From: THOMSON, DAVID; ADAMS, JADIE; HOLM, MICHAEL; BLACK, DAVID; BAROCIO, JESSE; BOEKWEG, SCOTT; ROYLANCE, SHANE; CLEMENTS, KIERSTEN; BOEHME, KENNETH; SKAGGS, JONATHAN; ORZECHOWSKI, GRZEGORZ; MCCLELLAN, JOSHUA
To: CAPTIONCALL, LLC
Reel/Frame 047905/0753 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2019
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 047905/0793 →
Cited By (14)
US 12,190,292 US 12,190,870 US 12,229,726 US 12,284,059 US 12,306,809 US 12,373,027 US 12,373,647 US 12,431,154 US 12,526,593 US 12,567,416 US 12,567,418 US 12,579,982 US 12,585,821 US 12,658,185