IP Library Granted Patent US 10,068,565
Granted Patent B2
US 10,068,565 · App. 14/563,511 · Granted Sep 4, 2018

Method and apparatus for an exemplary automatic speech recognition system

Inventor: Fathy Yassa (Soquel, CA)
G10L15/065G10L13/00G10L13/02G10L13/033G10L15/063G10L15/183G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,068,565
App. No.
14/563,511
Granted
Sep 4, 2018
Kind
B2
Abstract

An exemplary computer system configured to train an ASR using the output from a TTS engine.

Claims (7)

1. An automatic speech recognition (ASR) system comprising:

a text input module configured to receive a speech corpus comprising prosody information of at least one speech file of a first speaker and phonetic transcriptions corresponding to the at least one speech file, wherein the prosody information and the phonetic transcription are generated or gathered specifically for the at least one speech file of the first speaker, and wherein the prosody information comprises pitch contours and time durations for all phonemes, diphones, or triphones based on the phonetic transcriptions;

a text-to-speech (TTS) engine configured to receive the specific prosody information and the phonetic transcriptions from the first speech input module, and synthesize, based on a generic TTS acoustic model, the at least one speech file of the first speaker into an audio waveform having a first prosody based on the prosody information, and output the audio waveform and the phonetic transcriptions;

an ASR engine configured to receive the audio waveform and the phonetic transcriptions output by the TTS engine, and create an ASR acoustic model through training on the audio waveform and the phonetic transcriptions output by the TTS engine by compiling the audio waveform and the phonetic transcriptions output by the TTS engine into statistical representations of phonetic units of the audio waveform based on the phonetic transcriptions, wherein the phonetic units comprise phonemes, diphones, or triphones;

a speech input module configured to receive an input speech of a second speaker;

a speech morphing module comprising an ASR module configured to morph the human speech of the second speaker having a second prosody, preserving word boundary information and types of breaks between words and changing pitch contours and time durations for all phonemes, diphones, or triphones, into morphed human speech having a prosody that is the same as the first prosody of the audio waveform of the at least one speech audio file of the first speaker output by the TTS engine, wherein the first and the second prosody comprise corresponding pitch contours and time durations for all phonemes, diphones, or triphones;

wherein the ASR engine receives the morphed human speech morphed by the speech morphing module, recognizes the morphed human speech based on the trained ASR acoustic model, and outputs text corresponding to the recognized morphed human speech.

Continuity (2)
Provisional Application 61913188 · Dec 6, 2013
Related Publication 20150161983A1 · Jun 11, 2015
Cited By (1)
US 12,505,830