IP Library Granted Patent US 9,710,819
Granted Patent B2
US 9,710,819 · App. 12/618,742 · Granted Jul 18, 2017

Real-time transcription system utilizing divided audio chunks

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,710,819
App. No.
12/618,742
Granted
Jul 18, 2017
Kind
B2
Abstract

A computing system accepts audio from one or more sources, parses the audio into chunks, and transcribes the chunks in substantially real time. Some transcription is performed automatically, while other transcription is performed by humans who listen to the audio and enter the words spoken and/or the intent of the caller (such as directions given to the system). The system provides for participants a user interface that is updated in substantially real time with the transcribed text from the audio stream(s). A single audio line can be used for simple transcription, and multiple audio lines are used to provide a real-time transcript of a conference call, deposition, or the like. A pool of analysts creates, checks, and/or corrects transcription, and callers/observers can even assist in the correction process through their respective user interfaces. Ads derived from the transcript are displayed together with the text in substantially real time.

Claims (47)

1. A system, comprising a processor and a memory in communication with the processor, the memory storing programming instructions executable by the processor to:

cause a first chunk of audio data to be played for a first analyst, where the first chunk of audio data represents a segment of a first audio stream, the segment associated with a participant in a conference call;

accept input from the first analyst sufficient to indicate a transcription of the segment of the first audio stream, the input being dependent upon a selected fidelity mode chosen from a plurality of fidelity modes, the plurality of fidelity modes having corresponding levels of fidelity and including all of:

a first verbatim interpreting mode wherein the first analyst provides a substantially verbatim transcription of the first audio stream;

a second text interpreting mode wherein the first analyst listens to the first chunk of audio data and provides input of text that has a substantially identical meaning to the words spoken in the first chunk of audio data; and

a third automatic transcription mode wherein the first analyst repeats the first chunk of audio as input for an automatic transcription subsystem responsive to the automatic transcription subsystem having a level of confidence below a threshold when provided with the first audio stream as input;

cause the transcription and an identity of the participant in the conference call to be displayed on a web-based user interface in substantially real time relative to the capture of the segment of the first audio stream;

determining that the transcription displayed on the user interface is not accurate; and

automatically, by a computer system and without user input, selecting, based on the levels of fidelity corresponding to the plurality of fidelity modes, a different fidelity mode in the plurality of fidelity modes having a corresponding level of fidelity higher than the selected fidelity mode.

2. The system of claim 1 , wherein the programming instructions are further executable to:

cause additional chunks of audio data to be played for a plurality of analysts that comprises the first analyst, where the first chunk and the additional chunks together represent all of the audio in the first audio stream;

accept input from the plurality of analysts sufficient to indicate a transcription of the first audio stream; and

cause the transcription to be displayed on the web-based user interface in substantially real time relative to the parsing of each of the chunk and additional chunks.

3. The system of claim 2 , wherein the first chunk and at least one of the additional chunks represent overlapping segments of the first audio stream.

4. The system of claim 2 , wherein the programming instructions are further executable to:

cause further audio chunks of audio data to be played for the plurality of analysts, where the further chunks represent a second audio stream associated with the first audio stream;

accept input from the plurality of analysts sufficient to indicate a transcription of the second audio stream; and

cause the transcription of the second audio stream to be displayed on the web-based user interface in substantially real time relative to when each of the further chunks was captured.

5. The system of claim 4 , wherein the first audio stream and the second audio stream are produced by different participants in the conference call.

6. The system of claim 1 , wherein the parsing chooses low volume points in the stream for the beginning and end of the chunk.

7. The system of claim 1 , wherein the audio stream is parsed from a video stream.

8. The system of claim 1 , wherein:

the selected fidelity mode is the text interpreting mode, and the input from the first analyst also indicates the analyst's interpretation of an intent of a person whose voice is represented by the first chunk of audio data; and

the programming instructions are further executable to respond automatically to the interpreted intent.

9. The system of claim 8 , wherein:

the interpreted intent is to suspend transcription of audio from all audio streams; and

the automatic response is to suspend transcription of audio from all audio streams until the system receives instruction from a user to resume.

10. The system of claim 9 , wherein:

while transcription is suspended, the system still causes the audio streams to be played for analysts, but

no transcription is accepted other than an interpretation of the intent of a person to resume transcription.

11. A system comprising a processor and a memory in communication with the processor, the memory storing programming instructions executable by the processor to:

play each of a plurality of audio chunks, each for at least one of a plurality of analysts, where the audio chunks were each captured from a single line associated with a participant of a conference call, and together represent speech arriving on all lines of the conference call;

accept, responsive to a selected fidelity mode chosen from a plurality of fidelity modes, input from the plurality of analysts, the input indicating a transcript of each of the plurality of chunks, and collectively indicating a transcript of speech on all lines of the conference call, the plurality of fidelity modes having corresponding levels of fidelity and including all of:

a first verbatim interpreting mode wherein a first analyst provides a substantially verbatim transcription of a first audio stream;

a second text interpreting mode wherein the first analyst listens to a first chunk of audio data and provides input of text that has a substantially identical meaning to the words spoken in the first chunk of audio data; and

a third automatic transcription mode wherein the first analyst repeats the first chunk of audio as input for an automatic transcription subsystem responsive to the automatic transcription subsystem having a level of confidence below a threshold when provided with the first audio stream as input;

cause the transcript to be displayed on a web-based user interface to at least one participant in the conference call, where the display of the transcript of each chunk occurs in substantially real time relative to the capture of that chunk and includes an identity of the participant of the conference call associated with the line from which the chunk was captured;

determine that the transcript displayed on the user interface is not accurate; and

automatically, by a computer system and without user input, selecting, based on the levels of fidelity corresponding to the plurality of fidelity modes, a different fidelity mode in the plurality of fidelity modes having a corresponding level of fidelity higher than the selected fidelity mode.

12. The system of claim 11 , wherein the programming instructions are further executable by the processor to:

accept input indicating that the transcript of a chunk is incorrect;

correct the transcript of that chunk; and

display the corrected transcript of that chunk.

13. The system of claim 12 , wherein the programming instructions executable by the processor to correct the transcript comprise instructions to accept input from an analyst in the plurality of analysts indicating a corrected transcript of the chunk.

14. The system of claim 11 , wherein the programming instructions are further executable by the processor to:

accept intent input from at least one of the plurality of analysts, where the intent input indicates the analyst's interpretation of an intent of a conference call participant; and

automatically respond to the intent of the conference call participant.

Assignments (17)
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT 036099/0682 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060544/0445 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 043039/0808 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060557/0636 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 29, 2017
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 043039/0808 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036099/0682 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
MERGER Recorded Dec 16, 2014
From: INTERACTIONS CORPORATION
To: INTERACTIONS LLC
Reel/Frame 034514/0066 →
SECURITY INTEREST Recorded Jun 27, 2014
From: INTERACTIONS CORPORATION
To: ORIX VENTURES, LLC
Reel/Frame 033247/0219 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2010
From: CLORAN, MICHAEL ERIC; HEITZMAN, DAVID PAUL; SHIELDS, MITCHELL GREGORY; GOETZ, JEROMEY RUSSELL
To: INTERACTIONS CORPORATION
Reel/Frame 025014/0123 →