IP Library › Patent Application 16601631
Patent Application
App. No. 16/601,631

VOICE INTERACTION METHOD, APPARATUS AND DEVICE, AND STORAGE MEDIUM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/601,631
Abstract

A voice interaction method, apparatus and device, and a computer-readable storage medium are provided. The method includes: receiving a voice signal to be detected within a preset time period; performing a voice identification on the voice signal to be detected, to obtain a text to be detected; and performing a first detection on the text to be detected, and providing a response according to the text to be detected in response to determining that the first detection is passed. In the embodiments, the misrecognition rate of a voice signal during a voice interaction is reduced, thereby improving user experience.

Claims (50)

1 . A voice interaction method, comprising:

receiving a voice signal to be detected within a preset time period;

performing a voice identification on the voice signal to be detected, to obtain a text to be detected; and

performing a first detection on the text to be detected, and providing a response according to the text to be detected in response to determining that the first detection is passed.

2 . The voice interaction method according to claim 1 , wherein the providing a response according to the text to be detected in response to determining that the first detection is passed comprises:

performing a second detection on the text to be detected in response to determining that the first detection is passed; and

performing the response according to the text to be detected, in response to determining that the second detection is passed.

3 . The voice interaction method according to claim 2 , wherein

the performing a first detection on the text to be detected comprises: performing a grammar and/or semantic detection on the text to be detected, with a preset first detection model; and

the performing a second detection on the text to be detected comprises: performing a contextual logic relation detection on the text to be detected, with a preset second detection model.

4 . The voice interaction method according to claim 3 , wherein the method further comprises establishing the first detection model by:

training the first detection model by using a plurality of instruction texts and a plurality of non-instruction texts; wherein

the plurality of instruction texts are texts associated with voice instructions, and the plurality of non-instruction texts are texts associated with voice signals other than voice instructions.

5 . The voice interaction method according to claim 4 , wherein the performing a first detection on the text to be detected comprises:

inputting the text to be detected into the first detection model; and

predicting that the text to be detected is an instruction text with the first detection model, and determining that the first detection is passed; or predicting that the text to be detected is a non-instruction text with the first detection model, and determining that the first detection is not passed.

6 . The voice interaction method according to claim 3 , wherein the method further comprises establishing the second detection model by:

training the second detection model by using a plurality of sets of voice interactive texts and a plurality of sets of non-voice interactive texts; wherein

each set of voice interactive texts comprises texts associated with voice instructions in at least two rounds of voice interactions and responses to the texts, and a contextual logic relation exists between the at least two rounds of voice interactions; and

each set of the non-voice interactive texts comprises texts associated with at least two voice instructions between which no contextual logic relation exists.

7 . The voice interaction method according to claim 3 , wherein the performing a second detection on the text to be detected comprises:

inputting the text to be detected, a historical instruction text associated with a historical voice instruction of the text to be detected and a historical response to the historical instruction text into the second detection model; and

predicting that the text to be detected has a contextual logic relation with the historical instruction text and the historical response with the second detection model, and determining that the second detection is passed; or predicting that the text to be detected has no contextual logic relation with the historical instruction text and the historical response with the second detection model, and determining that the second detection is not passed.

8 . A voice interaction apparatus, comprising:

one or more processors; and

a memory for storing one or more programs, wherein

the one or more programs are executed by the one or more processors to enable the one or more processors to:

receive a voice signal to be detected within a preset time period;

perform a voice identification on the voice signal to be detected, to obtain a text to be detected; and

perform a first detection on the text to be detected, and provide a response according to the text to be detected in response to determining that the first detection is passed.

9 . The voice interaction apparatus according to claim 8 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to:

perform a second detection on the text to be detected in response to determining that the first detection is passed; and

perform the response according to the text to be detected, in response to determining that the second detection is passed.

10 . The voice interaction apparatus according to claim 9 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to:

perform a grammar and/or semantic detection on the text to be detected, with a preset first detection model; and

perform a contextual logic relation detection on the text to be detected, with a preset second detection model.

11 . The voice interaction apparatus according to claim 10 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to establish the first detection model by:

training the first detection model by using a plurality of instruction texts and a plurality of non-instruction texts; wherein

the plurality of instruction texts are texts associated with voice instructions, and the plurality of non-instruction texts are texts associated with voice signals other than voice instructions.

12 . The voice interaction apparatus according to claim 11 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to

input the text to be detected into the first detection model; and

predict that the text to be detected is an instruction text with the first detection model, and determine that the first detection is passed; or predict that the text to be detected is a non-instruction text with the first detection model, and determine that the first detection is not passed.

13 . The voice interaction apparatus according to claim 10 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to establish the second detection model by:

training the second detection model by using a plurality of sets of voice interactive texts and a plurality of sets of non-voice interactive texts; wherein

each set of voice interactive texts comprises texts associated with voice instructions in at least two rounds of voice interactions and responses to the texts, and a contextual logic relation exists between the at least two rounds of voice interactions; and

each set of the non-voice interactive texts comprises texts associated with at least two voice instructions between which no contextual logic relation exists.

14 . The voice interaction apparatus according to claim 10 , wherein the one or more programs are executed by the one or more processors to enable the one or more processors to:

input the text to be detected, a historical instruction text associated with a historical voice instruction of the text to be detected and a historical response to the historical instruction text into the second detection model; and

predict that the text to be detected has a contextual logic relation with the historical instruction text and the historical response with the second detection model, and determine that the second detection is passed; or predict that the text to be detected has no contextual logic relation with the historical instruction text and the historical response with the second detection model, and determine that the second detection is not passed.

15 . A non-transitory computer-readable storage medium, in which a computer program is stored, wherein the computer program, when executed by a processor, causes the processor to implement the method of claim 1 .

Assignments (2)
PATENT LICENSE AGREEMENT Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: SHANGHAI XIAODU TECHNOLOGY CO., LTD.
Reel/Frame 056427/0560 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2019
From: ZHANG, GANG; ZHU, KAIHUA; GAO, CONG; WANG, DAN
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 050816/0550 →