IP Library Granted Patent US 11,562,731
Granted Patent B2
US 11,562,731 · App. 16/997,846 · Granted Jan 24, 2023

Word replacement in transcriptions

Inventors: David Thomson (Bountiful, UT); Cody Barton (Holladay, UT)
Assignee: Sorenson IP Holdings, LLC
G10L15/02G10L15/08G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,731
App. No.
16/997,846
Granted
Jan 24, 2023
Kind
B2
Abstract

A method may include obtaining first audio data of a communication session between a first device and a second device and obtaining, during the communication session, a first text string that is a transcription of the first audio data. The method may further include directing the first text string to the first device for presentation of the first text string during the communication session and obtaining, during the communication session, a second text string that is a transcription of the first audio data. The method may further include comparing a first accuracy score of the first word to a second accuracy score of the second word and in response to a difference between the first accuracy score and the second accuracy score satisfying a threshold, directing the second word to the first device to replace the first word in the first location as displayed by the first device.

Claims (50)

1. A method comprising:

obtaining first audio data of a communication session between a first device and a second device;

obtaining, during the communication session, a first text string that is a first transcription of the first audio data, the first text string including a first word in a first location of the first transcription;

directing the first text string to the first device for presentation of the first text string during the communication session;

obtaining, during the communication session, a second text string that is a second transcription of the first audio data, the second text string including a second word in a second location of the second transcription that is different from the first word, the second location corresponding to the first location;

comparing a first accuracy score of the first word to a second accuracy score of the second word; and

in response to a difference between the first accuracy score and the second accuracy score satisfying a threshold, directing the second word to the first device to replace the first word in the first location as displayed by the first device.

2. The method of claim 1 , wherein the first text string is obtained from a first automatic transcription system and the second text string is obtained from a second automatic transcription system that is different than the first automatic transcription system.

3. The method of claim 1 , wherein both the first text string and the second text string are partial text strings that are not finalized text strings as generated by automatic transcription systems.

4. The method of claim 1 , wherein in response to the difference between the first accuracy score and the second accuracy score not satisfying the threshold, one or more words of the first text string are not replaced by one or more words of the second text string.

5. The method of claim 1 , further comprising obtaining an indication of a time lapse from when a second previous word is directed to the first device to replace a first previous word, wherein the second word is directed to the first device to replace the first word in the first location in further response to the time lapse satisfying a time threshold.

6. The method of claim 1 , wherein the threshold is adjusted in response to the second word being generated by a second automatic transcription system that is different than a first automatic transcription system that generates the first word.

7. The method of claim 1 , further comprising:

obtaining, during the communication session, a third text string that is a third transcription of the first audio data, the third text string including a third word in a third location of the third transcription;

directing the third text string to the first device for presentation of the third text string during the communication session;

obtaining, during the communication session, a fourth text string that is a fourth transcription of the first audio data, the fourth text string including a fourth word in a fourth location of the fourth transcription that is different from the third word, the third location corresponding to the fourth location;

comparing a third accuracy score of the third word to a fourth accuracy score of the fourth word; and

in response to the fourth accuracy score being greater than the third accuracy score and a difference between the third accuracy score and the fourth accuracy score not satisfying the threshold, determining to maintain the third word in the second location as displayed by the first device instead of directing the fourth word to the first device to replace the third word in the second location as displayed by the first device in response to the fourth accuracy score being greater than the third accuracy score and a difference between the third accuracy score and the fourth accuracy score satisfying the threshold.

8. The method of claim 1 , further comprising:

obtaining a first content score of the first word, the first content score indicating an effect of the first word on a meaning of the first transcription; and

obtaining a second content score of the second word, the second content score indicating an effect of the second word on a meaning of the second transcription,

wherein the second word is directed to the first device to replace the first word in the first location in further response to a sum of the first content score and the second content score satisfying a content threshold.

9. The method of claim 1 , further comprising in response to the difference between the first accuracy score and the second accuracy score satisfying the threshold, directing a third word to the first device to replace a fourth word in a third location in the first transcription as displayed by the first device.

10. The method of claim 9 , wherein a difference between a fourth accuracy score of the fourth word and a third accuracy score of the third word does not satisfy the threshold.

11. The method of claim 9 , wherein the second third location is before the first location in the first transcription.

12. A non-transitory computer-readable medium configured to store instructions that when executed by a computer system perform the method of claim 1 .

13. A method comprising:

obtaining first audio data of a communication session between a first device and a second device;

obtaining, during the communication session, a first text string that is a transcription of the first audio data, the first text string including a plurality of words;

directing the first text string to the first device for presentation of the first text string during the communication session;

determining, during the communication session, a plurality of replacement words to replace a subset of the plurality of words displayed by the first device;

determining a number of the plurality of replacement words; and

in response to the number of the plurality of replacement words satisfying a threshold, directing the plurality of replacement words to the first device to replace the subset of the plurality of words as displayed by the first device.

14. The method of claim 13 , further comprising obtaining an indication of a time lapse, wherein the plurality of replacement words are directed to the first device to replace the subset of the plurality of words as displayed by the first device in further response to the time lapse satisfying a time threshold.

15. The method of claim 13 , wherein a first accuracy score of one of the plurality of replacement words is greater than a second accuracy score of one of the subset of the plurality of words that corresponds to the one of the plurality of replacement words.

16. A method comprising:

obtaining first audio data of a communication session between a first device and a second device;

obtaining, during the communication session, a first text string that is a first transcription of the first audio data, the first text string including a first word in a first location of the first transcription;

directing the first text string to the first device for presentation of the first text string during the communication session;

obtaining, during the communication session, a second text string that is a second transcription of the first audio data, the second text string including a second word in a second location of the second transcription that is different from the first word, the second location corresponding to the first location;

obtaining a score of the second word, the score indicating an effect of the second word on a meaning of the second transcription; and

in response to the score satisfying a threshold, directing the second word to the first device to replace the first word in the first location as displayed by the first device.

17. The method of claim 16 , further comprising:

obtaining a first accuracy score of the first word; and

obtaining a second accuracy score of the second word,

wherein the second word is directed to the first device to replace the first word in the first location in further response to a difference between the first accuracy score and the second accuracy score satisfying an accuracy threshold.

18. The method of claim 16 , further comprising obtaining an indication of a time lapse from when a second previous word is directed to the first device to replace a first previous word,

wherein the second word is directed to the first device to replace the first word in the first location in further response to the time lapse satisfying a time threshold.

19. The method of claim 16 , wherein the first text string is obtained from a first automatic transcription system and the second text string is obtained from a second automatic transcription system that is different than the first automatic transcription system.

20. The method of claim 16 , further comprising in response to the score satisfying the threshold, directing a third word to the first device to replace a fourth word in a second location in the first transcription as displayed by the first device, wherein a score of the fourth word, which indicates an effect of the fourth word on a meaning of the first transcription, does not satisfy the threshold.

Assignments (6)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY DATA THE NAME OF THE LAST RECEIVING PARTY SHOULD BE CAPTIONCALL, LLC PREVIOUSLY RECORDED ON REEL 67190 FRAME 517. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded May 31, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONCALL, LLC
Reel/Frame 067591/0675 →
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2024
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: SORENSON IP HOLDINGS, LLC; SORENSON COMMUNICATIONS, LLC; CAPTIONALCALL, LLC
Reel/Frame 067190/0517 →
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
JOINDER NO. 1 TO THE FIRST LIEN PATENT SECURITY AGREEMENT Recorded Apr 22, 2021
From: SORENSON IP HOLDINGS, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056019/0204 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 055374/0753 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: MONTERO, ADAM; MANAM, AKHIL BABU; HARRIS, BEN; CHEVRIER, BRIAN; PETERSON, BRUCE; BARTON, CODY; BLACK, DAVID; THOMSON, DAVID; ADAMS, JADIE; DUNN, JASON P; ALLISON, JOSH; BOEHME, KENNETH; ADAMS, MARK; HOLM, MICHAEL; RAKOV, RACHEL; BOEKWEG, SCOTT; ROYLANCE, SHANE; HAO, SHUJIAN
To: CAPTIONCALL, LLC
Reel/Frame 055374/0776 →