IP Library Patent Application 16601630
Patent Application
App. No. 16/601,630

METHOD, DEVICE AND APPARATUS FOR RECOGNIZING VOICE SIGNAL, AND STORAGE MEDIUM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/601,630
Abstract

A method, device and apparatus for recognizing a voice signal, and a storage medium are provided. The method includes: collecting a voice signal; extracting the voiceprint feature of the voice signal; comparing the voiceprint feature with a pre-stored reference voiceprint feature; and recognizing a content of the voice signal with a voice recognition model, in response to a consistence of the voiceprint feature with the pre-stored reference voiceprint feature. Embodiments of the present application can improve the accuracy of recognizing voice signals.

Claims (51)

1 . A method for recognizing a voice signal, comprising:

collecting a voice signal;

extracting a voiceprint feature of the voice signal;

comparing the voiceprint feature with a pre-stored reference voiceprint feature; and

recognizing a content of the voice signal with a voice recognition model, in response to a consistence of the voiceprint feature with the pre-stored reference voiceprint feature.

2 . The method according to claim 1 , further comprising: prestoring at least one reference voiceprint feature,

wherein the comparing the voiceprint feature with a pre-stored reference voiceprint feature comprises:

comparing the voiceprint feature with the reference voiceprint feature, to determine whether the voiceprint feature is consistent with the reference voiceprint feature.

3 . The method according to claim 2 , further comprising: determining at least one reference voiceprint feature by:

acquiring at least one user's voice signal;

extracting a voiceprint feature of the user's voice signal; and

determining the voiceprint feature of the user's voice signal as the reference voiceprint feature.

4 . The method according to claim 2 , further comprising: pre-establishing at least one voice recognition model corresponding to the at least one reference voiceprint feature,

wherein the recognizing the content of the voice signal with a voice recognition model comprises:

determining a voice recognition model corresponding to the reference voiceprint feature, in response to a consistence of the voiceprint feature with the reference voiceprint feature; and

recognizing the content of the voice signal with the determined voice recognition model.

5 . The method according to claim 4 , wherein the pre-establishing at least one voice recognition model corresponding to the at least one reference voiceprint feature comprises:

training the voice recognition model corresponding to the reference voiceprint feature, by using a user's voice signal having the reference voiceprint feature and real text information of the user's voice signal,

wherein the training the voice recognition model corresponding to the reference voiceprint feature comprises:

inputting the user's voice signal into the voice recognition model;

comparing text information outputted by the voice recognition model with the real text information, to obtain a comparison result; and

adjusting parameters of the voice recognition model according to the comparison result.

6 . An apparatus for recognizing a voice signal, comprising:

one or more processors; and

a storage device configured to store one or more programs, wherein

the one or more programs, when executed by the one or more processors, cause the one or more processors to:

collect a voice signal;

extract a voiceprint feature of the voice signal;

compare the voiceprint feature with a pre-stored reference voiceprint feature; and

recognize a content of the voice signal with a voice recognition model, in response to a consistence of the voiceprint feature with the pre-stored reference voiceprint feature.

7 . The apparatus according to claim 6 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

prestore at least one reference voiceprint feature, and

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

compare the voiceprint feature with the reference voiceprint feature, to determine whether the voiceprint feature is consistent with the reference voiceprint feature.

8 . The apparatus according to claim 7 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

determine at least one reference voiceprint feature by:

acquiring at least one user's voice signal;

extracting a voiceprint feature of the user's voice signal; and

determining the voiceprint feature of the user's voice signal as the reference voiceprint feature.

9 . The apparatus according to claim 7 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

pre-establish at least one voice recognition model corresponding to the at least one reference voiceprint feature, and

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

determine a voice recognition model corresponding to the reference voiceprint feature, in response to a consistence of the voiceprint feature with the reference voiceprint feature; and

recognize the content of the voice signal with the determined voice recognition model.

10 . The apparatus according to claim 9 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

train the voice recognition model corresponding to the reference voiceprint feature, by using a user's voice signal having the reference voiceprint feature and real text information of the user's voice signal, and

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

input the user's voice signal into the voice recognition model;

compare text information outputted by the voice recognition model with the real text information, to obtain a comparison result; and

adjust parameters of the voice recognition model according to the comparison result.

11 . A non-transitory computer-readable storage medium comprising computer executable instructions stored thereon, wherein the executable instructions, when executed by a processor, causes the processor to implement the method of claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2019
From: LIU, YONG; ZHOU, JI; XUE, XIANGDONG; WANG, PENG; ZHAO, LIFENG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051803/0735 →