IP Library Granted Patent US 12,125,487
Granted Patent B2
US 12,125,487 · App. 17/450,551 · Granted Oct 22, 2024

Method and system for conversation transcription with metadata

Inventors: Kiersten L. Bradley (Lafayette, CO); Ethan Coeytaux (Boulder, CO); Ziming Yin (Toronto, CA)
Assignee: SoundHound AI IP, LLC.
G10L15/26G06F40/134G06F40/166G06F40/284G10L15/02G10L15/063G10L15/07G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,125,487
App. No.
17/450,551
Granted
Oct 22, 2024
Kind
B2
Abstract

Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed and multiuser-editable transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript by one or more editors. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

Claims (63)

1. A computer-implemented method for automatic conversation transcription, comprising:

receiving audio streams from at least one audio source;

generating speech segments by segmenting the audio streams, wherein the segmenting is based on voice activity detection;

generating a plurality of text strings by transcribing the speech segments with acoustic models;

determining a plurality of speaker identities associated with the plurality of text strings based on a speaker diarization model;

assigning respective indicators to the plurality of text strings based on the plurality of speaker identities, wherein text strings associated with one speaker are assigned to the same indicator;

generating a transcript by combining the plurality of text strings associated with the respective indicators;

enabling editing the transcript via an editing application;

timestamping the plurality of text strings according to a common clock;

storing the timestamps associated with the text strings;

receiving a request from the editing application to play audio corresponding to a text string;

playing audio beginning at the timestamp corresponding to the requested text string, wherein the editing application is configured to continuously update the transcript according to the audio streams in real-time; and

displaying the transcript as text with line breaks, wherein vertical spacing between the text strings indicates an amount of break time between the speech segments, wherein the transcript text of a second speaker aligns with the transcript text of a first speaker according to their recorded timestamps.

2. The computer-implemented method of claim 1 , wherein the speaker diarization model is configured to utilize one or more diarization factors to determine the plurality of speaker identities.

3. The computer-implemented method of claim 2 , wherein the one or more diarization factors comprise audio channel data.

4. The computer-implemented method of claim 2 , wherein the one or more diarization factors comprise speech feature vectors data.

5. The computer-implemented method of claim 4 , further comprising:

determining the distance between the speech feature vectors of a group of speech segments is below a threshold; and

clustering the group of speech segments by assigning the same indicator to the group of speech segments.

6. The computer-implemented method of claim 2 , wherein the one or more diarization factors comprise at least one of acoustic beamforming data and speaker visual data.

7. The computer-implemented method of claim 1 , wherein, during joint editing of the transcript, the editing application is configured to identify a user with a unique marker displayed on the transcript.

8. The computer-implemented method of claim 1 , wherein the editing application is configured to assign various editing permissions to the plurality of users.

9. The computer-implemented method of claim 1 , further comprising:

determining a first speech segment overlaps with a second speech segment in time;

sending the transcript as text to a visual display with line breaks between speech segments; and

reducing spacing between the text strings associated with the first speech segment and the second speech segment in the transcript.

10. The computer-implemented method of claim 1 , comprising:

embedding hyperlinks within the plurality of text strings, wherein the hyperlinks are associated with corresponding speech segments of the audio streams; and

enabling, by receiving a selected hyperlink associated with a speech segment, a playback of relevant audio streams.

11. A computer-implemented method for automatic conversation transcription, comprising:

receiving audio streams from at least one audio source;

generating speech segments by segmenting the audio streams, wherein the segmenting is based on voice activity detection;

generating a plurality of text strings by transcribing the speech segments with a speech recognition system;

determining a plurality of speaker identities associated with the plurality of text strings based on a speaker diarization model;

assigning respective indicators to the plurality of text strings based on the plurality of speaker identities, wherein text strings associated with one speaker are assigned to the same indicator;

displaying, on a screen, video streams accompanying the audio streams;

capturing screenshots of the video streams accompanying the audio streams with timestamps based on a common clock;

generating a transcript by combining the plurality of text strings associated with the respective indicators and the screenshots;

enabling editing the transcript via an editing application;

timestamping the plurality of text strings according to a common clock;

storing the timestamps associated with the text strings;

receiving a request from the editing application to play audio corresponding to a text string;

playing audio beginning at the timestamp corresponding to the requested text string, wherein the editing application is configured to continuously update the transcript according to the audio streams in real-time; and

displaying the transcript as text with line breaks, wherein vertical spacing between the text strings indicates an amount of break time between the speech segments, wherein the transcript text of a second speaker aligns with the transcript text of a first speaker according to their recorded timestamps.

12. The computer-implemented method of claim 11 , wherein the speaker diarization model is configured to utilize one or more diarization factors to determine the plurality of speaker identities.

13. The computer-implemented method of claim 11 , wherein, during joint editing of the transcript, the editing application is configured to identify a user with a unique marker displayed on the transcript.

14. The computer-implemented method of claim 11 , wherein the editing application is configured to assign various editing permissions to the plurality of users.

15. The computer-implemented method of claim 11 , wherein capturing screenshots of video streams accompanying the audio streams further comprises:

detecting pixel changes on the screen displaying the video streams, wherein capturing screenshots is conditional upon the number of pixel changes being greater than a threshold.

16. The computer-implemented method of claim 11 , further comprising:

displaying, on a screen, the screenshots of video streams in a grid, wherein the screenshots of video streams are configured to associate with corresponding text strings and to represent the content of the video streams;

receiving a selection of a screenshot in the grid;

displaying the corresponding text strings based on the selected screenshot; and

playing audio associated with the corresponding text strings.

17. The computer-implemented method of claim 11 , wherein the acoustic models comprise at least one domain-specific language model.

18. The computer-implemented method of claim 11 , further comprising:

enabling, via the editing application, a global replacement of a term in the transcript.

19. The computer-implemented method of claim 11 , further comprising:

identifying a key phrase within a text string; and

tagging the key phrase as a hyperlink anchor corresponding to a URL associated with the key phrase.

20. The computer-implemented method of claim 11 , further comprising:

identifying an n-gram text as having a low frequency within a language model; and

tagging the n-gram text as a hyperlink anchor corresponding to a URL associated with a definition of the n-gram text.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2021
From: BRADLEY, KIERSTEN L.; COEYTAUX, ETHAN; YIN, ZIMING
To: SOUNDHOUND, INC.
Reel/Frame 057799/0806 →
Continuity (2)
Provisional Application 63198328 · Oct 12, 2020
Related Publication 20220115019A1 · Apr 14, 2022
Cited By (3)
US 12,444,419 US 12,513,374 US 12,632,643