IP Library Granted Patent US 9,202,465
Granted Patent B2
US 9,202,465 · App. 13/072,003 · Granted Dec 1, 2015

Speech recognition dependent on text message content

Inventors: Gaurav Talwar (Farmington Hills, MI); Xufang Zhao (Windsor, CA)
Assignee: General Motors LLC
G10L15/22G10L15/30G10L2015/227G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,202,465
App. No.
13/072,003
Granted
Dec 1, 2015
Kind
B2
Abstract

A method of automatic speech recognition. An utterance is received from a user in reply to a text message, via a microphone that converts the reply utterance into a speech signal. The speech signal is processed using at least one processor to extract acoustic data from the speech signal. An acoustic model is identified from a plurality of acoustic models to decode the acoustic data, and using a conversational context associated with the text message. The acoustic data is decoded using the identified acoustic model to produce a plurality of hypotheses for the reply utterance.

Claims (30)

1. A method of automatic speech recognition, comprising the steps of:

a) receiving a text message at a speech recognition client device;

b) processing the text message with conversational context-specific language models and emotional context-specific language models stored on the client device using at least one processor of the client device to identify a conversational context and an emotional context corresponding to the text message;

c) synthesizing speech from the text message;

d) communicating the synthesized speech via a loudspeaker of the client device to a user of the client device;

e) receiving a reply utterance in response to the text message from the user via a microphone of the client device that converts the reply utterance into a speech signal;

f) pre-processing the speech signal using the at least one processor to extract acoustic data from the received speech signal;

g) communicating the extracted acoustic data, the identified conversational context, and identified emotional context to a speech recognition server;

h) identifying an acoustic model of a plurality of acoustic models stored at the server to be used for decoding the acoustic data based on the identified conversational context, the identified emotional context, or both;

i) decoding the acoustic data using the identified acoustic model to produce a plurality of hypotheses for the reply utterance; and

j) post-processing the plurality of hypotheses to identify one of the hypotheses as the reply utterance;

k) presenting the identified hypothesis to the user;

l) seeking confirmation from the user that the identified hypothesis is correct;

m) outputting the identified hypothesis as at least part of a reply text message if the user confirms that the identified hypothesis is correct; otherwise

n) using the emotional context to improve identification of the acoustic model, and repeating steps e) through m).

2. A method of automatic speech recognition, comprising the steps of:

a) receiving a text message at a speech recognition client device;

b) processing the text message with conversational context-specific language models and emotional context-specific language models stored on the client device using at least one processor of the client device to identify a conversational context and emotional context corresponding to the text message;

c) synthesizing speech from the text message;

d) communicating the synthesized speech via a loudspeaker of the client device to a user of the client device;

e) receiving a reply utterance in response to the text message from the user via a microphone of the client device that converts the reply utterance into a speech signal;

f) pre-processing the speech signal using the at least one processor to extract acoustic data from the received speech signal;

g) identifying an acoustic model of a plurality of acoustic models to decode the acoustic data, using the identified conversational context and emotional context associated with the text message;

h) decoding the acoustic data using the identified acoustic model to produce a plurality of hypotheses for the reply utterance;

i) determining whether a confidence value associated with at least one of the plurality of hypotheses for the reply utterance is greater or less than a confidence threshold;

j) communicating the extracted acoustic data, the conversational context, and the emotional context to a speech recognition server, if the confidence value is determined to be less than the confidence threshold, otherwise post-processing the plurality of hypotheses to identify one of the hypotheses as the reply utterance, and outputting from the client device the identified hypothesis as at least part of a reply text message;

k) identifying at the server, an acoustic model of a plurality of acoustic models stored at the server to decode the acoustic data, using the identified conversational context, the emotional context, or both;

l) decoding the acoustic data using the acoustic model identified at the server to produce a plurality of hypotheses for the reply utterance;

m) post-processing the plurality of hypotheses to identify one of the hypotheses as the reply utterance; and

n) outputting from the server the identified hypothesis as at least part of a reply text message.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2014
From: WILMINGTON TRUST COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 034183/0436 →
SECURITY AGREEMENT Recorded Jun 22, 2012
From: GENERAL MOTORS LLC
To: WILMINGTON TRUST COMPANY
Reel/Frame 028423/0432 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2011
From: TALWAR, GAURAV; ZHAO, XUFANG
To: GENERAL MOTORS LLC
Reel/Frame 026127/0345 →
Continuity (1)
Related Publication 20120245934A1 · Sep 27, 2012