IP Library › Granted Patent US 12,482,459
Granted Patent B2
US 12,482,459 · App. 17/897,352 · Granted Nov 25, 2025

Speech recognition system, acoustic processing method, and non-temporary computer-readable medium

Inventors: Yui Sudo (Wako, JP); Kazuhiro Nakadai (Wako, JP); Muhammad Shakeel (Wako, JP)
Assignee: HONDA MOTOR CO., LTD.
G10L15/197G10L15/02G10L15/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,459
App. No.
17/897,352
Granted
Nov 25, 2025
Kind
B2
Abstract

The speech recognition that is disclosed analyzes an acoustic feature for each subframe of an audio signal; provides a first model configured to determine a hidden state for each frame consisting of multiple subframes on the basis of the acoustic feature; provides a second model configured to determine a hidden state for each frame consisting of multiple subframes on the basis of the acoustic feature; and provides a third model configured to determine an utterance content on the basis of a sequence of the hidden states of each block consisting of multiple frames belonging to a voice segment.

Claims (53)

1 . A speech recognition system comprising:

a processor and a memory, the processor coupled to the memory,

the processor is configured to:

input an audio signal;

calculate an acoustic feature for each subframe of the audio signal;

calculate, by using a first model, a hidden state series for each frame consisting of multiple subframes on the basis of the acoustic feature;

specify, by using a second model, whether a voice segment or a non-voice segment for each block on the basis of the hidden state series, the block consisting of a plurality of frames;

calculate, by using a third model, a probability for an utterance content candidate on the basis of a sequence of the hidden state provided series for each block having a single voice segment to specify an utterance content; and

train the third model to calculate the probability for the utterance content candidate on the basis of hidden state series; wherein

the processor is configured to:

specify a first frame subsequent to the non-voice segment as a beginning of the voice segment,

specify a second frame prior to a succeeding non-voice segment as an end of the voice segment,

adjust block arrangement of the audio signal, by concatenating one or more frames up to the end of the voice segment in a first block with the end of the voice segment to a second block proceeding to the first block, concatenating one or more frames from the beginning of the voice segment in a third block with the beginning of the voice segment to a fourth block subsequent to the third block,

search for recognition results indicating an utterance content for each block arrangement of the audio signal based on the probability for the utterance content candidate calculated by using the third model, and

output the recognition results to an external device.

2 . The speech recognition system according to claim 1 , wherein

the processor is configure to divide, by using the second model, block comprising two or more voice segments into two or more blocks, the two or more blocks containing respective voice segments.

3 . The speech recognition system according to claim 2 ,

wherein the processor is configured to

calculate for each frame a probability that the frame belongs to a voice segment on the basis of the hidden state, the probability being as a voice segment probability;

specify, a segment having consecutive inactive frames in which the number of inactive frames is more than a predetermined threshold frame number as the non-voice segment, each of the inactive frames having voice segment probability equals to or less than a predetermined probability threshold; and

specify a segment having the consecutive inactive frames as voice segments.

4 . The speech recognition system according to claim 1 , wherein

the first model comprises a first-stage model and a second-stage model,

the first-stage model being designed for converting an acoustic feature for each subframe to a frame feature for each frame, and

the second-stage model being designed for specifying the hidden state series on the basis of the frame feature.

5 . The speech recognition system according to claim 1 , wherein

the third model is designed for calculating an estimated probability for each candidate of the utterance content corresponding to a hidden state series up to the latest block forming a voice segment, and specifying the utterance content with the highest estimated probability.

6 . A non-transitory computer-readable medium storing instructions at a speech recognition system, the instructions executed by a processor cause the speech recognition system to:

input an audio signal;

calculate an acoustic feature for each subframe of the audio signal;

calculate, by using a first model, a hidden state series for each frame consisting of multiple subframes on the basis of the acoustic feature;

specify, by using a second model, whether a voice segment or a non-voice segment for each block on the basis of the hidden state series, the block consisting of a plurality of frames;

calculate, by using a third model, a probability for an utterance content candidate on the basis of a sequence of the hidden state series provided for each block having a single voice segment to specify an utterance content; and

train the third model to calculate the probability for the utterance content candidate on the basis of hidden state series; wherein

the instructions cause the speech recognition system to:

specify a first frame subsequent to the non-voice segment as a beginning of the voice segment,

specify a second frame prior to a succeeding non-voice segment as an end of the voice segment,

adjust block arrangement of the audio signal, by concatenating one or more frames up to the end of the voice segment in a first block with the end of the voice segment to a second block preceding to the first block, concatenating one or more frames from the beginning of the voice segment in a third block with the beginning of the voice segment to a fourth block subsequent to the third block,

search for recognition results indicating an utterance content for each block arrangement of the audio signal based on the probability for the utterance content candidate calculated by using the third model, and

output the recognition results to an external device.

7 . A method for speech recognition, comprising the steps of:

inputting an audio signal;

calculating an acoustic feature for each subframe of the audio signal;

calculating, by using a first model, a hidden state series for each frame consisting of multiple subframes on the basis of the acoustic feature;

specifying, by using a second model, whether a voice segment or a non-voice segment for each block on the basis of the hidden state series, the block consisting of a plurality of frames;

calculating, by using a third model, a probability for an utterance content candidate on the basis of a sequence of the hidden state series provided for each block having a single voice segment to specify an utterance content; and

training the third model to calculate the probability for the utterance content candidate on the basis of hidden state series; further comprising the steps of:

specifying a first frame subsequent to the non-voice segment as a beginning of the voice segment,

specifying a second frame prior to a succeeding non-voice segment as an end if the voice segment,

adjusting block arrangement of the audio signal, by concatenating one or more frames up to the end of the voice segment in a first block with the end of the voice segment to a second block preceding to the first block, concatenating one or more frames from the beginning of the voice segment in a third block with the beginning of the voice segment to a fourth block subsequent to the third block,

searching for recognition results indicating an utterance content for each of arranged blocks of the audio signal based on the probability calculated by using the third model, and

outputting the recognition results to an external device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2022
From: SUDO, YUI; NAKADAI, KAZUHIRO; SHAKEEL, MUHAMMAD
To: HONDA MOTOR CO., LTD.
Reel/Frame 062020/0913 →
Continuity (1)
Related Publication 20240071379A1 · Feb 29, 2024
References Cited (5)
US 5613037A · Sukkar · 1997 [cited by examiner]
US 6629070B1 · Nagasaki · 2003 [cited by examiner]
US 20220230627A1 · Chang · 2022 [cited by examiner]
WO 2018207390 · 2018 [cited by applicant]
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, “Joint CTC-Attention Based End-To-End Speech Recognition Using Multi-Task Learning”, 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) … [cited by applicant]