IP Library Granted Patent US 9,704,484
Granted Patent B2
US 9,704,484 · App. 14/420,587 · Granted Jul 11, 2017

Speech recognition method and speech recognition device

Inventor: Shiro Iwai (Niiza, JP)
Assignee: HONDA ACCESS CORP.
G10L15/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,704,484
App. No.
14/420,587
Granted
Jul 11, 2017
Kind
B2
Abstract

Provided is a speech recognition device that executes a speech recognition method capable of improving speech recognition accuracy. The speech recognition device includes a trigger generation unit for generating a trigger signal on the basis of at least mouth movement and a speech recognition unit which extracts a sound signal on the basis of the trigger signal and starts speech recognition for speech in the extracted sound signal. When the trigger generation unit is generating a trigger signal solely on the basis of opening of the mouth, the trigger generation unit generates the trigger signal so as to precede the opening of the mouth by a predetermined period. Alternatively, when the trigger generation unit is generating a trigger signal on the basis of opening of the mouth and changes in eye orientation, the trigger generation unit generates the trigger signal from the moment any of the above occurs.

Claims (36)

1. A system for recognizing speech, said system comprising:

a microcomputer comprising a trigger generation section and a speech recognition section;

the trigger generation section configured to generate a trigger signal based on at least a condition of present-absent of opening of a mouth; and

the speech recognition section configured to, in response to the trigger signal, receive audio signals and start speech recognition relative to the received audio signals,

wherein, when the trigger generation section generates the trigger signal based solely on the condition of present-absent of opening of the mouth, the trigger generation section generates the trigger signal for a predetermined time duration retroactively from a time point at which the condition of present-absent of opening of the mouth is present, and,

when the trigger generation section generates the trigger signal based on the condition of the present-absent of opening of the mouth and generates another trigger signal based on present-absent of a change in a view direction of eyes, and/or present-absent of a change in an orientation of a face, the trigger generation section generates the trigger signal and the another trigger signal when one of the present-absent conditions is present, and

when only the trigger signal based on the condition of the present-absent of opening of the mouth is generated, the speech recognition section utilizes the trigger signal as is, and when the another trigger signal based on the condition of the present-absent of the change in the view direction of eyes and/or the present-absent of the change in an orientation of the face is also generated, the speech recognition section utilizes the another trigger signal.

2. The speech recognition device of claim 1 , wherein the speech recognition section is configured to perform the speech recognition based on the condition of present-absent of opening of the mouth, using the trigger signal in advance and, when an outcome of the speech recognition by the speech recognition section indicates an error, corrects the trigger signal, and the corrected trigger signal comprises a trigger signal that is generated for the predetermined duration of time retroactively from the time point at which the condition of present-absent of opening of the mouth is present.

3. The speech recognition device of claim 1 , wherein the predetermined time is a period of 2-3 seconds.

4. A speech recognition device comprising:

a microcomputer configured to generate a trigger signal based on at least a condition of present-absent of opening of a mouth, and in response to the trigger signal, receive audio signals and start speech recognition relative to the received audio signals, wherein,

the microcomputer is further configured to generate another trigger signal based on present-absent of the change in a view direction of eyes and/or present-absent of the change in an orientation of the face,

when the trigger signal based on the condition of the present-absent of opening of the mouth and the another trigger signal based on the condition of present-absent of the change in the view direction of eyes and/or present-absent of the change in an orientation of the face are generated, the microcomputer is configured to start speech recognition using the another trigger signal, and

when an outcome of the speech recognition indicates an error, the microcomputer is configured to generate the trigger signal for a predetermined time duration retroactively from a time point at which the condition of present-absent of opening of the mouth is present, and restart the speech recognition with the trigger signal.

5. The speech recognition device of claim 4 , wherein the predetermined time is a period of 2-3 seconds.

6. A computer-implemented speech recognition method comprising the steps of:

generating a trigger signal via a processor, said trigger signal based solely on motion of a first facial organ or based on the motion of the first facial organ and motion of a second facial organ different from the first facial organ; and

starting speech recognition relative to audio signals via the processor in response to the trigger signal,

wherein the first facial organ is a mouth, and

when only the trigger signal based on the motion of the first facial organ is generated, the trigger signal, as is, is utilized to start speech recognition, and when the trigger signal based on the motion of the second facial organ is also generated, the trigger signal based on the motion of the second facial organ is utilized to start the speech recognition.

7. The computer-implemented speech recognition method of claim 6 , wherein the second facial organ comprises an eye and/or a face.

8. The computer-implemented speech recognition method of claim 7 , wherein the motion of the mouth is present-absent of opening of the mouth, the motion of the eye is present-absent of a change in a view direction of the eye, and the motion of the face is present-absent of a change in an orientation of the face.

9. A computer-implemented speech recognition method comprising the steps of:

generating a trigger signal via a processor based on a condition of present-absent of motion of a mouth;

taking in audio signals via the processor in response to the trigger signal and starting speech recognition relative to the audio signals taken in,

wherein the trigger signal is generated from a predetermined time duration retroactively from a point in time at which the condition of present-absent of motion of the mouth is present, and

when only the trigger signal based on the condition of the present-absent of opening of the mouth is generated, the speech recognition is started using the trigger signals as is, and when another trigger signal based on a condition of present-absent of change in a view direction of eyes and/or present-absent of change in an orientation of a face is also generated, the speech recognition is started using the another trigger signal.

10. The computer-implemented speech recognition method of claim 9 , wherein, when the speech recognition encounters an error, the trigger signal is generated for the predetermined time duration retroactively from the time point at which the condition of present-absent of motion of the mouth is present.

11. The computer-implemented speech recognition method of claim 9 , wherein the predetermined time duration is a period of 2-3 seconds.

12. A computer-implemented speech recognition method comprising the steps of:

generating a trigger signal via a processor from a point in time at which a condition of present-absent of motion of a mouth is present;

taking in audio signals via the processor in response to the trigger signal and starting speech recognition relative to the audio signals taken in; and

judging via the processor whether an outcome of the speech recognition indicates an error,

wherein, when the trigger signal based on the condition of the present-absent of opening of the mouth and another trigger signal based on a condition of present-absent of change in a view direction of eyes and/or present-absent of change in an orientation of a face are generated, the speech recognition is started using the another trigger signal, and

when the outcome of the speech recognition indicates an error, the trigger signal is generated for a predetermined time duration retroactively from a point in time at which the condition of present-absent of motion of the mouth is present, and the speech recognition is restarted in response to the trigger signal.

13. The computer-implemented speech recognition method of claim 12 , wherein the predetermined time duration is a period of 2-3 seconds.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2015
From: IWAI, SHIRO
To: HONDA ACCESS CORP.
Reel/Frame 035402/0001 →
Priority Claims (1)
JP 2012-178701 · Aug 10, 2012 · national
Continuity (1)
Related Publication 20150206535A1 · Jul 23, 2015