IP Library Granted Patent US 8,756,064
Granted Patent B2
US 8,756,064 · App. 13/533,174 · Granted Jun 17, 2014

Method and system for creating frugal speech corpus using internet resources and conventional speech corpus

Inventors: SunilKumar Kopparapu (Mumbai, IN); Imran Ahmed Sheikh (Mumbai, IN)
Assignee: Tata Consultancy Services Limited
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,756,064
App. No.
13/533,174
Granted
Jun 17, 2014
Kind
B2
Abstract

A speech corpus creation method and system are disclosed. The method comprising identifying a publicly accessible first source of the first speech data and its corresponding first text transcription; extracting a second speech data of an accessible encoding format from the first speech data; extracting a second text transcription data with at least one encoding format from the first text transcription data; matching and aligning the transcription to the extracted second speech data at a sentence, word, phoneme level, or combination thereof to form a first and a second speech corpus; analyzing the text transcriptions in the second speech corpus to identify the short speech segments to produce a phonetically balanced, segmented, text aligned third speech corpus; and conditioning the third speech corpus by inserting a context and associated environment richer corpus therein the third speech corpus from at least one second source to form the final speech corpus.

Claims (30)

1. A speech corpus creation method, implementing extraction of a first speech data from at least one first source and mixing with at least one second source, the method comprising processor implemented steps of:

identifying at least one publicly accessible first source of the first speech data and its corresponding first text transcription;

extracting a second speech data of at least one accessible encoding format from the first speech data;

extracting a second text transcription data with at least one encoding format from the first text transcription data;

matching and aligning the transcription to the extracted second speech data at a sentence, word, phoneme level, or combination thereof to form a first and a second speech corpus;

analyzing the text transcriptions in the second speech corpus to identify the short speech segments to produce a phonetically balanced, segmented, text aligned third speech corpus; and

conditioning the third speech corpus by inserting a context and associated environment richer corpus therein the third speech corpus from at least one second source to form the final speech corpus.

2. A method as claimed in claim 1 , wherein speech data extractor engine extracts the first speech data along with transcription thereof from publicly accessible sources that are relevant to a desired corpus.

3. A method as claimed in claim 1 , wherein while aligning the speech, a two stage alignment process is carried out comprising a first level syllable matching adapted to match the extracted transcription to the corresponding extracted long speech data at a syllable level and a subsequent second level matching adapted to align, using automatic speech recognition engine, short speech segments of the long syllable aligned speech data at sentence, word or phoneme level.

4. A method as claimed in claim 1 , wherein the matching and aligning comprising steps of:

detecting plurality of syllables in the second speech data;

detecting plurality of syllables in the second text transcription data by employing a text syllable annotator;

annotating and indexing each detected syllable in the second speech data and in the second text transcription data;

aligning the syllable annotated second speech data with the syllable annotated second text data by matching the corresponding syllable indexes, to form a first syllable aligned speech corpus;

segmenting the said first aligned corpus into plurality of short speech segments of uniform length; and

aligning each short segment with the corresponding exacted text transcription to form a segmented text aligned second speech corpus, featuring alignment at sentence, word or phoneme level.

5. A method as claimed in claim 1 , wherein the third speech corpus, derived from a public source of speech data and its transcription, is conditioned with a context and associated environment richer corpus, collected using traditional procedure, to form the final speech corpus.

6. A speech corpus creation system, implementing extraction of a first speech data from at least one first source and mixing with at least one second source, the system comprising:

a speech data extractor adapted to extract a second speech data of at least one encoding format from the first speech data;

a text data extractor adapted to extract a second text transcription data of at least one encoding format from a first text transcription of the first speech data;

a speech alignment module adapted to match and align the first text transcription to the corresponding extracted long speech data in the first speech data, at a sentence word level, or combination thereof to form a first and a second speech corpus;

a phonetically balanced data extractor for analyzing the text transcriptions in the second speech corpus and to identify the short speech segments to form a phonetically balanced, segmented, text aligned third speech corpus; and

a compensator means adapted to identify at least one contextual gap in the third speech corpus and to condition the third speech corpus by inserting a context and associated environment richer corpus therein the third speech corpus from the at least one second source to form a final speech corpus.

7. A system as claimed in claim 6 , wherein speech data extractor engine extracts the first speech data along with transcription thereof from publicly accessible sources that are relevant to a desired corpus domain.

8. A system as claimed in claim 6 , wherein while aligning the speech, a two stage alignment is carried out comprising a first level syllable matching adapted to match the extracted transcription to the corresponding extracted long speech data at a syllable level and a subsequent second level of matching adapted to align, using automatic speech recognition, short speech segments of the syllable aligned long speech data, at least sentence, word and phoneme level.

9. A system as claimed in claim 6 , wherein the speech alignment module comprising of:

a speech syllable annotator adapted to annotate and index plurality of syllables in the second speech data;

a text syllable annotator adapted to annotate and index the syllables in the second text transcription data;

a syllable based aligner adapted to align the syllable indexed second speech data to the syllable indexed second text data by matching syllable indexes, to form a first syllable aligned long speech corpus, a long speech Segmenter adapted to segment the first syllable aligned long speech corpus into plurality of uniform segments; and

a short speech aligner adapted to align each short speech segment at least sentence, word and phoneme level with the corresponding transcription using an automatic speech recognition engine to form a segmented text aligned second speech corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2012
From: KOPPARAPU, SUNILKUMAR; SHEIKH, IMRAN AHMED
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 028445/0513 →
Priority Claims (1)
IN 2148/MUM/2011 · Jul 28, 2011 · national
Continuity (1)
Related Publication 20130030810A1 · Jan 31, 2013