IP Library Granted Patent US 12,014,730
Granted Patent B2
US 12,014,730 · App. 17/322,238 · Granted Jun 18, 2024

Voice processing method, electronic device, and storage medium

Inventor: Xiangyan Xu (Beijing, CN)
Assignee: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
G10L15/20G10L15/02G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,014,730
App. No.
17/322,238
Granted
Jun 18, 2024
Kind
B2
Abstract

A voice processing method includes: collecting a voice signal by a microphone of an electronic device, and signal-processing the collected voice signal to obtain a first voice frame segment; performing voice recognition on the first voice frame segment to obtain a first recognition result; in response to the first recognition result not matching a target content and a plurality of tokens in the first recognition result meeting a preset condition, performing frame compensation on the first voice frame segment to obtain a second voice frame segment; and performing voice recognition on the second voice frame segment to obtain a second recognition result. A matching degree between the second recognition result and the target content is greater than a matching degree between the first recognition result and the target content.

Claims (44)

1. A voice processing method, comprising:

collecting a voice signal by a microphone of an electronic device, and performing voice activity detection (VAD) processing of the collected voice signal to obtain a first voice frame segment;

detecting whether a phoneme is a filler phoneme starting from the first phoneme of the first voice frame segment, wherein the filler phoneme indicates non-keywords;

in response to determining that a probability of the phoneme being the filler phoneme is greater than a probability of the phoneme being not the filler phoneme, skipping the filler phoneme and performing voice recognition on the first voice frame segment to obtain a first recognition result;

in response to the first recognition result not matching a target content and a number of tokens in the first recognition result whose matching probability to the target content is greater than a second set threshold exceeding a preset number, determining that there is a VAD processing truncation in the first voice frame segment;

in response to the VAD processing truncation in the first voice frame segment,

estimating a target frame length compensated for the first voice frame segment according to a length of historical target content counted,

determining a next voice frame segment in the collected voice signal adjacent to the first voice frame segment,

obtaining, in the collected voice signal, a third voice frame segment having the target frame length and including complete token units, from a start position of the next voice frame segment, and

obtaining a second voice frame segment by splicing the third voice frame segment behind the first voice frame segment; and

performing voice recognition on the second voice frame segment to obtain a second recognition result, wherein a matching degree between the second recognition result and the target content is greater than a matching degree between the first recognition result and the target content.

2. The method of claim 1 , wherein obtaining the second voice frame segment comprises:

obtaining a fourth voice frame segment having a set frame length and including complete token units from a start position of the next voice frame segment, and obtaining the second voice frame segment by splicing the fourth voice frame segment behind the first voice frame segment.

3. The method of claim 1 , wherein a unit of each token of the plurality of tokens comprises at least one of a word, a phone, a monophone and a triphone.

4. An electronic device, comprising:

a microphone configured to collect a voice signal;

a processor; and

a memory storing instructions that when executed by the processor, control the processor to perform voice activity detection (VAD) processing of the collected voice signal to obtain a first voice frame segment;

detect whether a phoneme is a filler phoneme starting from the first phoneme of the first voice frame segment, wherein the filler phoneme indicates non-keywords;

in response to determining that a probability of the phoneme being the filler phoneme is greater than a probability of the phoneme being not the filler phoneme, skip the filler phoneme and perform voice recognition on the first voice frame segment to obtain a first recognition result;

in response to the first recognition result not matching a target content and a number of tokens in the first recognition result whose matching probability to the target content is greater than a second set threshold exceeding a preset number, determine that there is a VAD processing truncation in the first voice frame segment;

in response to the VAD processing truncation in the first voice frame segment,

estimate a target frame length compensated for the first voice frame segment according to a length of historical target content counted,

determine a next voice frame segment in the collected voice signal adjacent to the first voice frame segment,

obtain, in the collected voice signal, a third voice frame segment having the target frame length and including complete token units, from a start position of the next voice frame segment, and

obtain a second voice frame segment by splicing the third voice frame segment behind the first voice frame segment; and

perform voice recognition on the second voice frame segment to obtain a second recognition result, wherein a matching degree between the second recognition result and the target content is greater than a matching degree between the first recognition result and the target content.

5. The device of claim 4 , wherein the processor is further configured to:

obtain a fourth voice frame segment having a set frame length and including complete token units from a start position of the next voice frame segment, and obtain the second voice frame segment by splicing the fourth voice frame segment behind the first voice frame segment.

6. The device of claim 4 , wherein a unit of each token of the plurality of tokens comprises at least one of a word, a phone, a monophone and a triphone.

7. A non-transitory computer-readable storage medium having instructions stored thereon, when the instructions are executed by a processor in an electronic device, a voice processing method is implemented, the method comprising:

collecting a voice signal by a microphone of an electronic device, and performing voice activity detection (VAD) processing of the collected voice signal to obtain a first voice frame segment;

detecting whether a phoneme is a filler phoneme starting from the first phoneme of the first voice frame segment, wherein the filler phoneme indicates non-keywords;

in response to determining that a probability of the phoneme being the filler phoneme is greater than a probability of the phoneme being not the filler phoneme, skipping the filler phoneme and performing voice recognition on the first voice frame segment to obtain a first recognition result;

in response to the first recognition result not matching a target content and a number of tokens in the first recognition result whose matching probability to the target content is greater than a second set threshold exceeding a preset number, determining that there is a VAD processing truncation in the first voice frame segment;

in response to the VAD processing truncation in the first voice frame segment,

estimating a target frame length compensated for the first voice frame segment according to a length of historical target content counted,

determining a next voice frame segment in the collected voice signal adjacent to the first voice frame segment,

obtaining, in the collected voice signal, a third voice frame segment having the target frame length and including complete token units, from a start position of the next voice frame segment, and

obtaining a second voice frame segment by splicing the third voice frame segment behind the first voice frame segment; and

performing voice recognition on the second voice frame segment to obtain a second recognition result, wherein a matching degree between the second recognition result and the target content is greater than a matching degree between the first recognition result and the target content.

8. The method of claim 7 , wherein obtaining the second voice frame segment comprises:

obtaining a fourth voice frame segment having a set frame length and including complete token units from a start position of the next voice frame segment, and obtaining the second voice frame segment by splicing the fourth voice frame segment behind the first voice frame segment.

9. The method of claim 7 , wherein a unit of each token of the plurality of tokens comprises at least one of a word, a phone, a monophone and a triphone.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: XU, XIANGYAN
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 056262/0665 →
Priority Claims (1)
CN 202011324752.5 · Nov 23, 2020 · national
Continuity (1)
Related Publication 20220165258A1 · May 26, 2022
Cited By (1)
US 12,315,497