IP Library Patent Application 16557917
Patent Application
App. No. 16/557,917

METHOD, COMPUTER DEVICE AND STORAGE MEDIUM FOR IMPEMENTING SPEECH INTERACTION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/557,917
Abstract

The present disclosure provides a method, apparatus, computer device and storage medium for implementing speech interaction, wherein the method comprises: a content server obtaining a user's speech information from a client device, and completing the speech interaction in a first manner; the first manner comprises: sending the speech information to an automatic speech recognition server and obtaining a partial speech recognition result returned by the automatic speech recognition server each time; after determining that voice activity detection starts and if it is determined through semantic understanding that the partial speech recognition result obtained each time already includes entire content that the user hopes to express, taking the partial speech recognition result as a final speech recognition result, obtaining a response speech corresponding to the final speech recognition result, and returning the response speech to the client device. The solution of the present disclosure can be applied to improve the speech interaction response speed.

Claims (26)

1 . A method for implementing speech interaction, wherein the method comprises:

a content server obtaining a user's speech information from a client device, and completing the speech interaction in a first manner;

the first manner comprises: sending the speech information to an automatic speech recognition server and obtaining a partial speech recognition result returned by the automatic speech recognition server each time; after determining that voice activity detection starts and if it is determined through semantic understanding that the obtained partial speech recognition result already includes entire content that the user hopes to express, taking the partial speech recognition result as a final speech recognition result, obtaining a response speech corresponding to the final speech recognition result, and returning the response speech to the client device.

2 . The method according to claim 1 , wherein

the method further comprises:

for the partial speech recognition result obtained each time before and after the start of the voice activity detection, respectively obtaining a search result corresponding to the partial speech recognition result, and sending the search result to a Text To Speech server for speech synthesis;

upon obtaining the final speech recognition result, taking a speech synthesis result obtained according to the final speech recognition result as the response speech.

3 . The method according to claim 1 , wherein

the method further comprises:

after the content server obtaining the user's speech information, obtaining the user's expression attribute information;

if it is determined according to the expression attribute information that the user is a user who expresses content completely at one time, completing the speech interaction in the first manner.

4 . The method according to claim 3 , wherein

the method further comprises:

if it is determined according to the expression attribute information that the user is a user who does not express content completely at one time, completing the speech interaction in a second manner;

the second manner comprises:

sending the speech information to the automatic speech recognition server, and obtaining a partial speech recognition result returned by the automatic speech recognition server each time;

for the partial speech recognition result obtained each time, respectively obtaining a search result corresponding to the partial speech recognition result, and sending the search result to the Text To Speech server for speech synthesis;

upon determining that the voice activity detection ends, taking the finally-obtained speech syntheses result as the response speech, and returning the response speech to the client device.

5 . The method according to claim 3 , wherein

the method further comprises: determining the user's expression attribute information by analyzing the user's past speaking expression habits.

6 . A computer device, comprising a memory, a processor and a computer program which is stored on the memory and runs on the processor, wherein the processor, upon executing the program, implements a method for implementing speech interaction, wherein the method comprises:

a content server obtaining a user's speech information from a client device, and completing the speech interaction in a first manner;

the first manner comprises: sending the speech information to an automatic speech recognition server and obtaining a partial speech recognition result returned by the automatic speech recognition server each time; after determining that voice activity detection starts and if it is determined through semantic understanding that the obtained partial speech recognition result already includes entire content that the user hopes to express, taking the partial speech recognition result as a final speech recognition result, obtaining a response speech corresponding to the final speech recognition result, and returning the response speech to the client device.

7 . A computer-readable storage medium on which a computer program is stored, wherein the program, when executed by a processor, implements a method for implementing speech interaction, wherein the method comprises:

a content server obtaining a user's speech information from a client device, and completing the speech interaction in a first manner;

the first manner comprises: sending the speech information to an automatic speech recognition server and obtaining a partial speech recognition result returned by the automatic speech recognition server each time; after determining that voice activity detection starts and if it is determined through semantic understanding that the obtained partial speech recognition result already includes entire content that the user hopes to express, taking the partial speech recognition result as a final speech recognition result, obtaining a response speech corresponding to the final speech recognition result, and returning the response speech to the client device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2019
From: YUAN, CHAO; CHANG, XIANTANG; CHEN, HUAILIANG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051333/0104 →