IP Library › Granted Patent US 9,418,660
Granted Patent B2
US 9,418,660 · App. 14/156,032 · Granted Aug 16, 2016

Crowd sourcing audio transcription via re-speaking

Inventors: Matthias Paulik (San Jose, CA); Vivek Halder (Cupertino, CA); Ananth Sankar (Palo Alto, CA)
Assignee: Cisco Technology, Inc.
G10L15/26G06Q10/06311G10L15/04G10L15/07G10L15/32G10L25/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,418,660
App. No.
14/156,032
Granted
Aug 16, 2016
Kind
B2
Abstract

Speech audio that is intended for transcription into textual form is received. The received speech audio is divided into first speech segments. A plurality of speakers is identified. A speaker is configured for repeating in spoken form a first speech segment that the speaker has listened to. A subset of speakers is determined for sending each first speech segment. Each first speech segment is sent to the subset of speakers determined for the particular first speech segment. The second speech segments are received from the speakers. The second speech segment is a re-spoken version of a first speech segment that has been generated by a speaker by repeating in spoken form the first speech segment. The second speech segments are processed to generate partial transcripts. The partial transcripts are combined to generate a complete transcript that is a textual representation corresponding to the received speech audio.

Claims (40)

1. A method comprising:

receiving a speech audio intended for transcription to textual form at a job mapper on a data processing apparatus;

dividing, by the job mapper, the received speech audio into first speech segments;

identifying, by the job mapper, speakers for sending each first speech segment of the first speech segments;

sending, by the job mapper, each first speech segment to the speakers determined for a particular first speech segment;

receiving, at the job mapper, second speech segments from the speakers, wherein each second speech segment of the second speech segments is a re-spoken version of a first speech segment of the first speech segments that has been generated by one of the speakers by repeating in spoken form the first speech segment that the one of the speakers has listened to; and

processing, by the job mapper, the second speech segments to generate a complete transcript that is a textual representation corresponding to the received speech audio.

2. The method of claim 1 , wherein each of the speakers are configured for repeating in spoken form the first speech segment that a respective speaker has listened to.

3. The method of claim 1 , wherein processing the second speech segments to generate the complete transcript that is the textual representation corresponding to the received speech audio comprises:

processing the second speech segments to generate partial transcripts, wherein a partial transcript is a textual representation of a corresponding second speech segment; and

combining the partial transcripts to generate the complete transcript that is the textual representation corresponding to the received speech audio.

4. The method of claim 1 , wherein identifying speakers for sending the each first speech segment of the plurality of first speech segments comprise:

examining information associated with a plurality of speakers that are collected based on previous jobs;

mapping the each first speech segment to the plurality of speakers based on examining the collected information; and

determining the speakers from the plurality of speakers.

5. The method of claim 4 , wherein the collected information a speaker includes one or more of: an average time until the speaker provides a second speech segment, automatically-detected gender of the speaker, or accent of the speaker.

6. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving a speech audio intended for transcription to textual form;

dividing the received speech audio into first speech segments;

identifying speakers for sending each first speech segment;

sending the each first speech segment of a plurality of first speech segments to the speakers determined for a particular first speech segment;

receiving second speech segments from the speakers, wherein each second speech segment of the second speech segments is a re-spoken version of a first speech segment of the first speech segments that has been generated by one of the speakers by repeating in spoken form the first speech segment that the one of the speakers has listened to; and

processing the second speech segments to generate a complete transcript that is a textual representation corresponding to the received speech audio.

7. The system of claim 6 , wherein identifying speakers for sending the each first speech segment of the plurality of first speech segments further comprises:

identifying a plurality of speakers;

examining a user profile associated with each speaker;

mapping the each first speech segment to a subset of the plurality of speakers based on examining the user profiles; and

identifying the subset of the plurality speakers as the speakers for sending each first speech segment.

8. The system of claim 6 , wherein identifying speakers for sending the each first speech segment of the plurality of first speech segments further comprises:

identifying a plurality of speakers;

examining characteristics of each first speech segment;

mapping the each first speech segment to a subset of the plurality of speakers based on examining the characteristics of the first speech segment; and

identifying the subset of the plurality speakers as the speakers for sending the each first speech segment.

9. The system of claim 8 , wherein the characteristics of a person speaking the first speech segment include one or more of: a signal-to-noise ratio (SNR) of the first speech segment, a length of the first speech segment, gender of the person speaking the first speech segment, age of the person speaking the first speech segment, or accent of the person speaking the first speech segment.

10. The system of claim 6 , wherein processing the second speech segments to generate the complete transcript that is the textual representation corresponding to the received speech audio comprises:

processing, by individual automatic speech recognition (ASR) units associated with different speakers that have been trained for the associated speakers, the second speech segments.

11. The system of claim 10 , wherein processing the second speech segments to generate the complete transcript that is the textual representation corresponding to the received speech audio comprises:

storing the second speech segments received from the speakers; and

training the ASR units using the stored second speech segments.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2014
From: PAULIK, MATTHIAS; HALDER, VIVEK; SANKAR, ANANTH
To: CISCO TECHNOLOGY, INC.
Reel/Frame 032766/0218 →
Continuity (1)
Related Publication 20150199966A1 · Jul 16, 2015