IP Library Granted Patent US 12,094,468
Granted Patent B2
US 12,094,468 · App. 17/838,500 · Granted Sep 17, 2024

Speech detection method, prediction model training method, apparatus, device, and medium

Inventors: Yi Gao (Shanghai, CN); Weiran Nie (Shanghai, CN); Youjia Huang (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L15/25G06V10/82G06V40/172G10L15/05G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,094,468
App. No.
17/838,500
Granted
Sep 17, 2024
Kind
B2
Abstract

A speech detection method includes performing recognition on a photographed face image using a model to predict whether a user intends to continue speaking, and to determine whether a collected audio signal is a speech end point with reference to a prediction result.

Claims (82)

1. A method comprising:

obtaining an audio signal and a face image, wherein a first photographing time point of the face image is the same as a first collection time point of the audio signal;

inputting the face image into a prediction model to predict whether a user intends to continue speaking;

processing the face image using the prediction model to obtain a prediction result;

outputting the prediction result; and

determining that the audio signal is a speech end point when the prediction result indicates that the user does not intend to continue speaking.

2. The method of claim 1 , wherein processing the face image comprises:

extracting a key point from the face image;

processing the key point to obtain action features of the face image;

classifying the action features to obtain confidence degrees respectively corresponding to different types; and

determining the prediction result based on the confidence degrees.

3. The method of claim 1 , further comprising obtaining the prediction model through training based on a first sample face image and a second sample face image, wherein the first sample face image is marked with a first label indicating that a sample user intends to continue speaking, wherein the first label is based on a first sample audio signal, wherein a second collection time point of the first sample audio signal and a first collection object of the first sample audio signal are the same as a second photographing time point of the first sample face image and a first photographing object of the first sample face image, wherein the second sample face image is marked with a second label indicating that the sample user does not intend to continue speaking, wherein the second label is based on a second sample audio signal, and wherein a third collection time point of the second sample audio signal and a second collection object of the second sample audio signal are the same as a third photographing time point of the second sample face image and a second photographing object of the second sample face image.

4. The method of claim 3 , wherein the first sample audio signal meets a condition, and wherein the condition comprises at least one of:

a voice activity detection (VAD) result corresponding to the first sample audio signal is firstly updated from a speaking state to a silent state, and secondly updated from the silent state to the speaking state;

a trailing silence duration of the first sample audio signal is less than a first threshold and greater than a second threshold, wherein the first threshold is greater than the second threshold;

a first confidence degree of a text information combination is greater than a second confidence degree of first text information, wherein the text information combination is of the first text information and second text information, wherein the first text information indicates first semantics of a previous sample audio signal of the first sample audio signal, wherein the second text information indicates second semantics of a next sample audio signal of the first sample audio signal, wherein the first confidence degree indicates a first probability that the text information combination is a first complete statement, and wherein the second confidence degree indicates a second probability that the first text information is a second complete statement; or

the first confidence degree is greater than a third confidence degree of the second text information, wherein the third confidence degree indicates a third probability that the second text information is a third complete statement.

5. The method of claim 3 , wherein the second sample audio signal meets a condition, and wherein the condition comprises a voice activity detection (VAD) result corresponding to the second sample audio signal is updated from a speaking state to a silent state.

6. The method of claim 3 , wherein the first sample face image meets a condition, and wherein the condition comprises:

inputting the first sample face image into a first classifier in the prediction model, wherein the first classifier is configured to predict a first probability that the first sample face image comprises an action;

inputting the first sample face image into a second classifier in the prediction model, wherein the second classifier is configured to predict a second probability that the first sample face image does not comprise the action; and

identifying that the first probability is greater than the second probability.

7. The method of claim 3 , wherein the second sample audio signal meets a condition, and wherein the condition comprises a trailing silence duration of the second sample audio signal is greater than a threshold.

8. The method of claim 1 , wherein after obtaining the audio signal and the face image, the method further comprises:

performing speech recognition on the audio signal to obtain text information corresponding to the audio signal;

performing syntax analysis on the text information to obtain a first analysis result indicating whether the text information is a first complete statement;

determining that the audio signal is not the speech end point when the first analysis result indicates that the text information is not the first complete statement; and

inputting the face image into the prediction model when the first analysis result indicates that the text information is the first complete statement.

9. The method according of claim 8 , wherein performing syntax analysis comprises:

performing word segmentation on the text information to obtain a plurality of first words;

performing, for each of the first words, the syntax analysis on a second word in the first words to obtain a second analysis result corresponding to the second word, wherein the second analysis result indicates whether the second word and a previous word of the second word form a second complete statement;

determining that the text information is a third complete statement when a third analysis result corresponding to one of the first words indicates that a fourth complete statement is formed; and

determining that the text information is not the third complete statement when a fourth analysis result corresponding to each of the first words indicates that the fourth complete statement is not formed.

10. A method for training a prediction model for speech detection, wherein the method comprises:

obtaining a sample audio signal set and a sample face image set;

processing, based on a first sample audio signal in the sample audio signal set, a third sample face image in the sample face image set, to obtain a first sample face image marked with a first label, wherein the first label indicates that a sample user intends to continue speaking, and wherein a first photographing time point of the first sample face image and a first photographing object of the first sample face image are the same as a first collection time point of the first sample audio signal and a first collection object of the first sample audio signal;

processing, based on a second sample audio signal in the sample audio signal set, a fourth sample face image in the sample face image set to obtain a second sample face image marked with a second label, wherein the second label indicates that the sample user does not intend to continue speaking, and wherein a second photographing time point of the second sample face image and a second photographing object of the second sample face image are the same as a second collection time point of the second sample audio signal and a second collection object of the second sample audio signal; and

performing model training using the first sample face image and the second sample face image to obtain the prediction model to predict whether a user intends to continue speaking.

11. The method of claim 10 , wherein the first sample audio signal meets a condition, and wherein the condition comprises at least one of:

a voice activity detection (VAD) result corresponding to the first sample audio signal is firstly updated from a speaking state to a silent state, and secondly updated from the silent state to the speaking state;

a trailing silence duration of the first sample audio signal is less than a first threshold and greater than a second threshold, wherein the first threshold is greater than the second threshold;

a first confidence degree of a text information combination is greater than a second confidence degree of first text information, wherein the text information combination is of the first text information and second text information, wherein the first text information indicates first semantics of a previous sample audio signal of the first sample audio signal, wherein the second text information indicates second semantics of a next sample audio signal of the first sample audio signal, wherein the first confidence degree indicates a first probability that the text information combination is a first complete statement, and wherein the second confidence degree indicates a second probability that the first text information is a second complete statement; or

the first confidence degree is greater than a third confidence degree of the second text information, wherein the third confidence degree indicates a third probability that the second text information is a third complete statement.

12. The method of claim 10 , wherein the second sample audio signal meets a condition, and wherein the condition comprises at least one of:

a voice activity detection (VAD) result corresponding to the second sample audio signal is updated from a speaking state to a silent state; or

a trailing silence duration of the second sample audio signal is greater than a threshold.

13. The method according to claim 10 , wherein the first sample face image meets a condition, and wherein the condition comprises:

inputting the first sample face image into a first classifier in the prediction model, wherein the first classifier is configured to predict a first probability that the first sample face image comprises an action;

inputting the first sample face image into a second classifier in the prediction model, wherein the second classifier is configured to predict a second probability that the first sample face image does not comprise the action; and

identifying that the first probability is greater than the second probability.

14. An apparatus comprising:

a memory configured to store instructions; and

a processor coupled to the memory, wherein the instructions cause the processor to be configured to:

obtain an audio signal and a face image, wherein a first photographing time point of the face image is the same as a first collection time point of the audio signal;

input the face image into a prediction model to predict whether a user intends to continue speaking;

process the face image using the prediction model to obtain a prediction result;

output the prediction result; and

determine that the audio signal is a speech end point when the prediction result indicates that the user does not intend to continue speaking.

15. The apparatus of claim 14 , wherein the instructions further cause the processor to be configured to obtain the prediction model through training based on a first sample face image and a second sample face image, wherein the first sample face image is marked with a first label indicating that a sample user intends to continue speaking, wherein the first label is based on a first sample audio signal, wherein a second collection time point of the first sample audio signal and a first collection object of the first sample audio signal are the same as a second photographing time point of the first sample face image and a first photographing object of the first sample face image, wherein the second sample face image is marked with a second label, indicating that the sample user does not intend to continue speaking, wherein the second label is based on a second sample audio signal, and wherein a third collection time point of the second sample audio signal and a second collection object of the second sample audio signal are the same as a third photographing time point of the second sample face image and a second photographing object of the second sample face image.

16. The apparatus of claim 15 , wherein the first sample audio signal meets a condition, and wherein the condition comprises at least one of:

a voice activity detection (VAD) result corresponding to the first sample audio signal is firstly updated from a speaking state to a silent state, and secondly updated from the silent state to the speaking state;

a trailing silence duration of the first sample audio signal is less than a first threshold and greater than a second threshold, wherein the first threshold is greater than the second threshold;

a first confidence degree of a text information combination is greater than a second confidence degree of first text information, wherein the text information combination is of the first text information and second text information, wherein the first text information indicates first semantics of a previous sample audio signal of the first sample audio signal, wherein the second text information indicates second semantics of a next sample audio signal of the first sample audio signal, wherein the first confidence degree indicates a first probability that the text information combination is a first complete statement, and wherein the second confidence degree indicates a second probability that the first text information is a second complete statement; or

the first confidence degree is greater than a third confidence degree of the second text information, wherein the third confidence degree indicates a third probability that the second text information is a third complete statement.

17. The apparatus of claim 15 , wherein the second sample audio signal meets a condition, and wherein the condition comprises at least one of:

a voice activity detection (VAD) result corresponding to the second sample audio signal is updated from a speaking state to a silent state; or

a trailing silence duration of the second sample audio signal is greater than a threshold.

18. The apparatus of claim 15 , wherein the first sample face image meets a condition, and wherein the condition comprises:

inputting the first sample face image into a first classifier in the prediction model, wherein the first classifier is configured to predict a first probability that the first sample face image comprises an action;

inputting the first sample face image into a second classifier in the prediction model, wherein the second classifier is configured to predict a second probability that the first sample face image does not comprise the action; and

identifying that the first probability is greater than the second probability.

19. The apparatus of claim 14 , wherein the instructions further cause the processor to be configured to:

perform speech recognition on the audio signal to obtain text information corresponding to the audio signal;

perform syntax analysis on the text information to obtain a first analysis result indicating whether the text information is a first complete statement;

determine that the audio signal is not the speech end point when the first analysis result indicates that the text information is not the first complete statement; and

input the face image into the prediction model when the first analysis result indicates that the text information is the first complete statement.

20. A computer program product comprising computer-executable instructions that are stored on a non-transitory computer-readable medium and that, when executed by a processor, cause an apparatus to:

obtain an audio signal and a face image, wherein a first photographing time point of the face image is the same as a first collection time point of the audio signal;

input the face image into a prediction model to predict whether a user intends to continue speaking;

process the face image using the prediction model to obtain a prediction result;

output the prediction result; and

determine that the audio signal is a speech end point when the prediction result indicates that the user does not intend to continue speaking.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2024
From: HUAWEI TECHNOLOGIES CO., LTD.
To: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 069335/0967 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2024
From: GAO, YI; NIE, WEIRAN; HUANG, YOUJIA
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 068043/0345 →