IP Library Granted Patent US 12,217,756
Granted Patent B2
US 12,217,756 · App. 17/465,509 · Granted Feb 4, 2025

Systems and methods for improved digital transcript creation using automated speech recognition

Inventors: Robert Ackerman (Boca Raton, FL); Anthony J. Vaglica (Silver Spring, MD); Holli Goldman (Richboro, PA); Amber Hickman (Swedesboro, NJ); Walter Barrett (Gibbsboro, NJ); Cameron Turner (Palo Alto, CA); Shawn Rutledge (Seattle, WA)
Assignee: AUDAX PRIVATE DEBT LLC
G10L15/26G06F17/18G06V40/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,756
App. No.
17/465,509
Granted
Feb 4, 2025
Kind
B2
Abstract

This disclosure relates generally to systems, methods, and computer readable media for providing improved insights and annotations to enhance recorded audio, video, and/or written transcriptions of testimony. For example, in some embodiments, a method is disclosed for correlating non-verbal cues recognized from an audio and/or video recording of testimony to the corresponding testimony transcript locations. In other embodiments, a method is disclosed for providing testimony-specific artificial intelligence-based insights and annotations to a testimony transcript, e.g., based on the use of machine learning, natural language processing, and/or other techniques. In still other embodiments, a method is disclosed for providing smart citations to a testimony transcript, e.g., which track the location of semantic constructs within the transcript over the course of various modifications being made to the transcript. In yet other embodiments, a method is disclosed for providing intelligent speaker identification-related insights and annotations to an audio recording of a testimony transcript.

Claims (57)

1. A method, comprising:

obtaining a video recording of a testimony given by a first deponent;

obtaining a transcript of the testimony;

scanning the video recording to locate one or more emotional cues and non-verbal cues;

linking the located one or more emotional cues and non-verbal cues to corresponding portions of the transcript;

updating the corresponding portions of the transcript with indications of the corresponding located one or more emotional cues and non-verbal cues;

updating the corresponding portions of the video recording with indications of the corresponding located one or more emotional cues and non-verbal cues, wherein:

at least one of the indications comprises a graphical overlay on a frame of the video recording visible during playback of the video recording;

the graphical overlay includes a visual indicator displaying derived insights corresponding to the located one or more emotional cues and non-verbal cues, wherein the derived insights comprise a bar graph depicting an amount of the one or more emotional cues and non-verbal cues; and

playing the video recording with the graphical overlay.

2. The method of claim 1 , wherein the located one or more emotional cues and non-verbal cues comprise at least one of the following: speech patterns, stress levels, response delay amounts, and movement anomalies.

3. The method of claim 1 , further comprising:

comparing the transcript to a different point in time in the transcript of the first deponent or to one or more additional transcripts given by the first deponent, wherein comparing the transcript includes correlating the located one or more emotional cues and non-verbal cues to a corresponding portion at the different point in time in the transcript of the first deponent or to corresponding portions of the one or more additional transcripts given by the first deponent.

4. The method of claim 1 , wherein the updating includes adding the indication to a page and line number of the transcript corresponding to the located one or more emotional cues and non-verbal cues.

5. The method of claim 1 , wherein the updating further includes detecting and tagging logical inconsistencies in the transcript as determined by a trained machine learning model.

6. The method of claim 1 , further comprising:

scanning the transcript to identify a semantic construct, wherein the semantic construct includes one or more words for tracking;

associating the semantic construct with a unique identifier; and

associating a page and line number in the transcript with the semantic construct and the unique identifier.

7. The method of claim 6 , wherein the semantic construct includes a word, a phrase, a sentence, a paragraph, a question, or an answer.

8. The method of claim 6 , further comprising:

automatically updating the page and line number associated with the semantic construct and the unique identifier on a condition that the transcript is updated.

9. The method of claim 8 , further comprising:

splitting the semantic construct into two or more new semantic constructs based on updates to the transcript;

updating the unique identifier associated with the semantic construct, wherein the updating includes creating new unique identifiers for each of the new semantic constructs.

10. The method of claim 9 , wherein updating the unique identifier further includes deleting the unique identifier for the semantic construct on a condition that the semantic construct is no longer used.

11. The method of claim 1 , wherein the located one or more emotional cues and non-verbal cues comprise anger, truthfulness, confidence, movement, or delay.

12. A non-transitory program storage device comprising instructions stored thereon to cause one or more processors to:

obtain a video recording of a testimony given by a first deponent;

obtain a transcript of the testimony;

scan the video recording to locate one or more emotional cues and non-verbal cues;

link the located one or more emotional cues and non-verbal cues to corresponding portions of the transcript;

update the corresponding portions of the transcript with indications of the corresponding located one or more emotional cues and non-verbal cues;

update the corresponding portions of the video recording with indications of the corresponding located one or more emotional cues and non-verbal cues, wherein:

at least one of the indications comprises a graphical overlay on a frame of the video recording visible during playback of the video recording;

the graphical overlay includes a visual indicator displaying derived insights corresponding to the located one or more emotional cues and non-verbal cues, wherein the derived insights comprise a bar graph depicting an amount of the one or more emotional cues and non-verbal cues; and

playing the video recording with the graphical overlay.

13. The non-transitory program storage device of claim 12 , wherein the located one or more emotional cues and non-verbal cues comprise at least one of the following: speech patterns, stress levels, response delay amounts, and movement anomalies.

14. The non-transitory program storage device of claim 12 , wherein the instructions further comprise instructions to cause the one or more processors to:

comparing the transcript to a different point in time in the transcript of the first deponent or to one or more additional transcripts given by the first deponent, wherein comparing the transcript includes correlating the located one or more emotional cues and non-verbal cues to a corresponding portion at the different point in time in the transcript of the first deponent or to corresponding portions of the one or more additional transcripts given by the first deponent.

15. A device, comprising:

a memory;

a display;

a user interface; and

one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to:

obtain a video recording of a testimony given by a first deponent;

obtain a transcript of the testimony;

scan the video recording to locate one or more emotional cues and nonverbal cues;

link the located one or more emotional cues and non-verbal cues to corresponding portions of the transcript;

update the corresponding portions of the transcript with indications of the corresponding located one or more emotional cues and non-verbal cues;

update the corresponding portions of the video recording with indications of the corresponding located one or more emotional cues and non-verbal cues, wherein:

at least one of the indications comprises a graphical overlay on a frame of the video recording visible during playback of the video recording;

the graphical overlay includes a visual indicator displaying derived insights corresponding to the located one or more emotional cues and non-verbal cues, wherein the derived insights comprise a bar graph depicting an amount of the one or more emotional cues and non-verbal cues; and

playing the video recording with the graphical overlay.

16. The device of claim 15 , wherein the located one or more emotional cues and non-verbal cues comprise at least one of the following: speech patterns, stress levels, response delay amounts, and movement anomalies.

17. The device of claim 15 , wherein the instructions further comprise instructions to cause the one or more processors to:

compare the transcript to a different point in time in the transcript of the first deponent or to one or more additional transcripts given by the first deponent, wherein comparing the transcript includes correlating the located one or more emotional cues and non-verbal cues to a corresponding portion at the different point in time in the transcript of the first deponent or to corresponding portions of the one or more additional transcripts given by the first deponent.

Assignments (1)
PATENT SECURITY AGREEMENT Recorded Nov 22, 2022
From: MAGNA LEGAL SERVICES, LLC; BARKLEY COURT REPORTERS, INC.
To: AUDAX PRIVATE DEBT LLC
Reel/Frame 061985/0232 →
Continuity (3)
Continuation 16570699 · Sep 13, 2019
Provisional Application 62730700 · Sep 13, 2018
Related Publication 20220059096A1 · Feb 24, 2022
References Cited (25)
US 6332122B1 · Ortega et al. · 2001 [cited by applicant]
US 6424946B1 · Tritschler et al. · 2002 [cited by applicant]
US 9990926B1 · Pearce · 2018 [cited by applicant]
US 10178301B1 · Welbourne et al. · 2019 [cited by applicant]
US 10304458B1 · Woo · 2019 [cited by applicant]
US 20050086226A1 · Krachman · 2005 [cited by examiner]
US 20060149558A1 · Kahn · 2006 [cited by examiner]
US 20070260457A1 · Bennett · 2007 [cited by examiner]
US 20090037171A1 · McFarland · 2009 [cited by examiner]
US 20090094029A1 · Koch et al. · 2009 [cited by applicant]
US 20090276215A1 · Hager · 2009 [cited by applicant]
US 20120011085A1 · Kocks et al. · 2012 [cited by applicant]
US 20160004679A1 · Grimm · 2016 [cited by examiner]
US 20160148043A1 · Bathiche · 2016 [cited by examiner]
US 20160260435A1 · Baard et al. · 2016 [cited by applicant]
US 20180174587A1 · Bermundo · 2018 [cited by applicant]
US 20180197548A1 · Palakodety et al. · 2018 [cited by applicant]
US 20180268305A1 · Dhondse · 2018 [cited by examiner]
US 20180293221A1 · Finkelstein et al. · 2018 [cited by applicant]
US 20180315249A1 · Suto et al. · 2018 [cited by applicant]
US 20180315429A1 · Taple · 2018 [cited by examiner]
US 20180358034A1 · Chakra et al. · 2018 [cited by applicant]
US 20190132265A1 · Nowak-Przygodzki et al. · 2019 [cited by applicant]
Kempter, Renato, et al. “EmotionWatch: Visualizing fine-grained emotions in event-related tweets.” Proceedings of the international AAAI conference on web and social media. vol. 8. No. 1. 2014. [cited by examiner]
Robinson, Raquel, et al. “All the feels: designing a tool that reveals streamers' biometrics to spectators.” Proceedings of the 12th International Conference on the Foundations of Digital Games. 2017. [cited by examiner]