IP Library Granted Patent US 11,699,043
Granted Patent B2
US 11,699,043 · App. 17/449,377 · Granted Jul 11, 2023

Determination of transcription accuracy

Inventor: Scott Boekweg (South Jordan, UT)
Assignee: Sorenson IP Holdings, LLC
G06F40/30G10L15/1815G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,699,043
App. No.
17/449,377
Granted
Jul 11, 2023
Kind
B2
Abstract

A method may include obtaining audio of a communication session between a first device of a first user and a second device of a second user. The method may further include obtaining a transcription of second speech of the second user. The method may also include identifying one or more first sound characteristics of first speech of the first user. The method may also include identifying one or more first words indicating a lack of understanding in the first speech. The method may further include determining an experienced emotion of the first user based on the one or more first sound characteristics. The method may also include determining an accuracy of the transcription of the second speech based on the experienced emotion and the one or more first words.

Claims (54)

1. A method comprising:

obtaining audio of a communication session between a first device of a first user and a second device of a second user, the communication session configured for verbal communication such that the audio includes first audio captured by the first device and second audio captured by the second device;

identifying one or more first characteristics of the first audio captured by the first device;

obtaining a transcription of the second audio captured by the second device;

identifying one or more first words in the transcription indicating the first user lacks understanding of the second audio; and

based on the one or more first characteristics of the first audio and the one or more first words, determining an accuracy of the transcription of the second audio.

2. The method of claim 1 , wherein the first audio includes first speech of the first user and the first characteristics of the first audio include: a tone, a volume, a pitch, an inflection, a timbre, or a speed of the first speech.

3. The method of claim 1 , further comprising:

obtaining a second transcription of the first audio; and

identifying one or more second words in the second transcription indicating the first user lacks understanding of the second audio, wherein the accuracy of the transcription of the second audio is determined further based on the second words.

4. The method of claim 1 , further comprising identifying one or more second characteristics of the second audio, wherein determining the accuracy of the transcription of the second audio is further based on the one or more second characteristics of the second audio.

5. The method of claim 1 , wherein the one or more first characteristics of the first audio are identified at a first time, the method further comprising:

identifying one or more second characteristics of the first audio at a second time different than the first time; and

determining a difference between the one or more first characteristics and the one or more second characteristics,

wherein the accuracy of the transcription of the second audio is determined based on the determined difference.

6. The method of claim 1 , further comprising obtaining one or more characteristics of sound in an environment of the first user, wherein the determining the accuracy of the transcription of the second audio is further based on the one or more characteristics of the sound in the environment.

7. The method of claim 1 , further comprising determining a topic of the communication session, wherein the determining the accuracy of the transcription of the second audio is further based on the topic.

8. A method comprising:

obtaining audio of a communication session between a first device of a first user and a second device of a second user, the communication session configured for verbal communication such that the audio includes first audio captured by the first device and second audio captured by the second device;

identifying one or more first characteristics of the first audio captured by the first device;

obtaining a transcription of the second audio captured by the second device; and

based on the one or more first characteristics of the first audio, determining an accuracy of the transcription of the second audio.

9. The method of claim 8 , wherein the first audio includes first speech of the first user and the first characteristics of the first audio include: a tone, a volume, a pitch, an inflection, a timbre, or a speed of the first speech.

10. The method of claim 8 , further comprising:

obtaining a second transcription of the first audio; and

identifying one or more words in the second transcription indicating the first user lacks understanding of the second audio, wherein the accuracy of the transcription of the second audio is determined further based on the one or more words.

11. The method of claim 8 , wherein the one or more first characteristics of the first audio are identified at a first time, the method further comprising:

identifying one or more second characteristics of the first audio at a second time different than the first time; and

determining a difference between the one or more first characteristics and the one or more second characteristics,

wherein the accuracy of the transcription of the second audio is determined based on the determined difference.

12. The method of claim 8 , further comprising obtaining one or more characteristics of sound in an environment of the first user, wherein the determining the accuracy of the transcription of the second audio is further based on the one or more characteristics of the sound in the environment.

13. The method of claim 8 , further comprising determining a topic of the communication session, wherein the determining the accuracy of the transcription of the second audio is further based on the topic.

14. A system comprising:

one or more processors; and

one or more non-transitory computer-readable media configured to store instructions that in response to being executed by the one or more processors cause the system to perform operations, the operations comprising:

obtaining audio of a communication session between a first device of a first user and a second device of a second user, the communication session configured for verbal communication such that the audio includes first audio captured by the first device and second audio captured by the second device;

identifying one or more first characteristics of the first audio captured by the first device;

obtaining a transcription of the second audio captured by the second device;

identifying one or more first words in the transcription indicating the first user lacks understanding of the second audio; and

based on the one or more first characteristics of the first audio and the one or more first words, determining an accuracy of the transcription of the second audio.

15. The system of claim 14 , wherein the first audio includes first speech of the first user and the first characteristics of the first audio include: a tone, a volume, a pitch, an inflection, a timbre, or a speed of the first speech.

16. The system of claim 14 , wherein the operations further comprise:

obtaining a second transcription of the first audio; and

identifying one or more second words in the second transcription indicating the first user lacks understanding of the second audio,

wherein the accuracy of the transcription of the second audio is determined further based on the second words.

17. The system of claim 14 , wherein the operations further comprise identifying one or more second characteristics of the second audio,

wherein determining the accuracy of the transcription of the second audio is further based on the one or more second characteristics of the second audio.

18. The system of claim 14 , wherein the one or more first characteristics of the first audio are identified at a first time, the operations further comprising:

identifying one or more second characteristics of the first audio at a second time different than the first time; and

determining a difference between the one or more first characteristics and the one or more second characteristics,

wherein the accuracy of the transcription of the second audio is determined based on the determined difference.

19. The system of claim 14 , wherein the operations further comprise obtaining one or more characteristics of sound in an environment of the first user,

wherein the determining the accuracy of the transcription of the second audio is further based on the one or more characteristics of the sound in the environment.

20. The system of claim 14 , wherein the operations further comprise determining a topic of the communication session, wherein the determining the accuracy of the transcription of the second audio is further based on the topic.

Assignments (3)
SECURITY INTEREST Recorded Apr 23, 2024
From: SORENSON COMMUNICATIONS, LLC; INTERACTIVECARE, LLC; CAPTIONCALL, LLC
To: OAKTREE FUND ADMINISTRATION, LLC, AS COLLATERAL AGENT
Reel/Frame 067573/0201 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: BOEKWEG, SCOTT
To: CAPTIONCALL, LLC
Reel/Frame 057653/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: CAPTIONCALL, LLC
To: SORENSON IP HOLDINGS, LLC
Reel/Frame 057653/0691 →