IP Library Granted Patent US 11,741,952
Granted Patent B2
US 11,741,952 · App. 16/847,852 · Granted Aug 29, 2023

Voice skill starting method, apparatus, device and storage medium

Inventors: Guangya Zhu (Beijing, CN); Xiao Zhou (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G10L15/22G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,952
App. No.
16/847,852
Granted
Aug 29, 2023
Kind
B2
Abstract

The present application disclosures a voice skill starting method, an apparatus, a device, and a storage medium, and relates to the field of artificial intelligence. An implementation scheme is that: the method is applied to an electronic device including at least one third-party voice skill and a built-in voice skill, and the electronic device is currently in the built-in voice skill. The method includes: receiving a current demand instruction of a user; judging whether the current demand instruction belongs to an entry demand instruction corresponding to the third-party voice skill according to a mapping relationship in response to the current demand instruction; and switching from the built-in voice skill to the third-party voice skill if it is determined that the current demand instruction belongs to the entry demand instruction corresponding to the third-party voice skill.

Claims (88)

1. A voice skill starting method, wherein the method is applied to an electronic device comprising at least one third-party voice skill and a built-in voice skill, and the electronic device is currently in the built-in voice skill, and the method comprises:

receiving a current demand instruction of a user;

judging whether the current demand instruction belongs to an entry demand instruction corresponding to the third-party voice skill according to a mapping relationship in response to the current demand instruction, wherein the mapping relationship is a pre-established mapping relationship between the third-party voice skill and the entry demand instruction, the mapping relationship is determined according to skill-associated characteristic data of a first historical demand instruction under each third-party voice skill, and the first historical demand instruction is located in the built-in voice skill, the skill-associated characteristic data includes skill-associated entry characteristic data and skill satisfaction data, the skill-associated entry characteristic data is data associated with the third-party voice skill and representing a characteristic of the entry demand instruction, the skill satisfaction data indicates characteristic data of whether the third-party voice skill satisfies a second historical demand instruction, and the second historical demand instruction is a historical demand instruction filtered from the first historical demand instruction according to the skill-associated entry characteristic data; and

switching from the built-in voice skill to the third-party voice skill if it is determined that the current demand instruction belongs to the entry demand instruction corresponding to the third-party voice skill;

before the receiving a current demand instruction of a user, further comprising:

obtaining the first historical demand instruction in the built-in voice skill;

determining the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill;

obtaining the entry demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated characteristic data; and

establishing the mapping relationship between each third-party voice skill and the corresponding entry demand instruction;

wherein the determining the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill, comprises:

determining skill-associated entry characteristic data of the first historical demand instruction under each third-party voice skill; and

determining skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill;

wherein the skill satisfaction data comprises content satisfaction or interaction satisfaction, and the determining skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill, comprises:

judging whether each third-party voice skill is a resource skill;

if a third-party voice skill is the resource skill, determining a content satisfaction of the second historical demand instruction under that third-party voice skill; and

if a third-party voice skill is a non-resource skill, determining an interaction satisfaction of the second historical demand instruction under that third-party voice skill;

wherein the determining the content satisfaction of the second historical demand instruction under that third-party voice skill, comprises:

obtaining a first playback resource duration of the second historical demand instruction under that third-party voice skill and a second playback resource duration of the second historical demand instruction under the built-in voice skill; and

determining the content satisfaction according to the first playback resource duration and the second playback resource duration.

2. The method according to claim 1 , after the obtaining the entry demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated characteristic data, further comprising:

judging whether there is one third-party voice skill corresponding to each entry demand instruction;

if there are multiple third-party voice skills corresponding to an entry demand instruction, determining an occurrence frequency of that entry demand instruction in each corresponding third-party voice skill; and

determining a third-party voice skill with the highest occurrence frequency as the third-party voice skill that has the mapping relationship with that entry demand instruction.

3. The method according to claim 1 , wherein the obtaining the entry demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated characteristic data, comprises:

obtaining the second historical demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated entry characteristic data; and

obtaining the entry demand instruction corresponding to each third-party voice skill by filtering the second historical demand instruction according to the skill satisfaction data.

4. A voice skill starting apparatus, wherein the apparatus is located in an electronic device comprising at least one third-party voice skill and a built-in voice skill, and the electronic device is currently in the built-in voice skill, and the apparatus comprises:

at least one processor; and

a memory, communicatively connected with the at least one processor; wherein

the memory stores instructions executable by the at least one processor, and the instruction is executed by the at least one processor to enable the at least one processor to:

receive a current demand instruction of a user;

judge whether the current demand instruction belongs to an entry demand instruction corresponding to the third-party voice skill according to a mapping relationship in response to the current demand instruction, wherein the mapping relationship is a pre-established mapping relationship between the third-party voice skill and the entry demand instruction, the mapping relationship is determined according to skill-associated characteristic data of a first historical demand instruction under each third-party voice skill, and the first historical demand instruction is located in the built-in voice skill, the skill-associated characteristic data includes skill-associated entry characteristic data and skill satisfaction data, the skill-associated entry characteristic data is data associated with the third-party voice skill and representing a characteristic of the entry demand instruction, the skill satisfaction data indicates characteristic data of whether the third-party voice skill satisfies a second historical demand instruction, and the second historical demand instruction is a historical demand instruction filtered from the first historical demand instruction according to the skill-associated entry characteristic data; and

switch from the built-in voice skill to the third-party voice skill if it is determined that the current demand instruction belongs to the entry demand instruction corresponding to the third-party voice skill;

wherein the instruction is executed by the at least one processor to enable the at least one processor, before the receiving a current demand instruction of a user, to:

obtain the first historical demand instruction in the built-in voice skill;

determine the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill;

obtain the entry demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated characteristic data; and

establish the mapping relationship between each third-party voice skill and the corresponding entry demand instruction;

wherein the instruction is executed by the at least one processor to enable the at least one processor, when determining the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill, to:

determine skill-associated entry characteristic data of the first historical demand instruction under each third-party voice skill; and

determine skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill;

wherein the skill satisfaction data comprises content satisfaction or interaction satisfaction, and the instruction is executed by the at least one processor to enable the at least one processor, when determining skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill, to:

judge whether each third-party voice skill is a resource skill;

if a third-party voice skill is the resource skill, determine a content satisfaction of the second historical demand instruction under that third-party voice skill; and

if a third-party voice skill is a non-resource skill, determine an interaction satisfaction of the second historical demand instruction under that third-party voice skill;

wherein the instruction is executed by the at least one processor to enable the at least one processor, when determining the content satisfaction of the second historical demand instruction under that third-party voice skill, to:

obtain a first playback resource duration of the second historical demand instruction under that third-party voice skill and a second playback resource duration of the second historical demand instruction under the built-in voice skill; and

determine the content satisfaction according to the first playback resource duration and the second playback resource duration.

5. The voice skill starting apparatus according to claim 4 , wherein the instruction is executed by the at least one processor to enable the at least one processor to:

judge whether there is one third-party voice skill corresponding to each entry demand instruction;

if there are multiple third-party voice skills corresponding to an entry demand instruction, determine an occurrence frequency of that entry demand instruction in each corresponding third-party voice skill; and

determine a third-party voice skill with the highest occurrence frequency as the third-party voice skill that has the mapping relationship with that entry demand instruction.

6. The voice skill starting apparatus according to claim 4 , wherein the instruction is executed by the at least one processor to enable the at least one processor to:

obtain the second historical demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated entry characteristic data; and

obtain the entry demand instruction corresponding to each third-party voice skill by filtering the second historical demand instruction according to the skill satisfaction data.

7. The voice skill starting apparatus according to claim 6 , wherein the instruction is executed by the at least one processor to enable the at least one processor, when obtaining the second historical demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated entry characteristic data, to:

input the skill-associated entry characteristic data of the first historical demand instruction into a trained-to-converged classification model corresponding to each third-party voice skill for classifying the first historical demand instruction by the classification model to obtain the second historical demand instruction corresponding to each third-party voice skill.

8. The voice skill starting apparatus according to claim 7 , wherein the instruction is executed by the at least one processor to enable the at least one processor, before inputting the skill-associated entry characteristic data of the first historical demand instruction into the trained-to-converged classification model corresponding to each third-party voice skill, to:

obtain a training sample of each classification model, wherein the training sample is a demand instruction sample, and the demand instruction sample has an identifier for identifying whether a demand instruction is capable of being used as the entry demand instruction corresponding to the third-party voice skill; and

train the corresponding classification model by using skill-associated entry characteristic data of the training sample until convergence, to obtain each trained-to-converged classification model.

9. The voice skill starting apparatus according to claim 4 , wherein the skill-associated entry characteristic data comprises: entry behavior characteristic data, skill correlation characteristic data, and entry grammar characteristic data.

10. The voice skill starting apparatus according to claim 4 , wherein the instruction is executed by the at least one processor to enable the at least one processor, when determining the interaction satisfaction of the second historical demand instruction under that third-party voice skill, to:

obtain text of a multi-round conversation corresponding to the second historical demand instruction under that third-party voice skill;

determine a skill response satisfaction and a skill response repetition rate of the second historical demand instruction under that third-party voice skill according to the text of the multi-round conversation; and

determine the interaction satisfaction according to the skill response satisfaction and the skill response repetition rate.

11. The voice skill starting apparatus according to claim 10 , wherein the instruction is executed by the at least one processor to enable the at least one processor, when obtaining the entry demand instruction by filtering the second historical demand instruction according to the skill satisfaction data, to:

if a third-party voice skill is the resource skill, obtain the entry demand instruction by filtering the second historical demand instruction according to a content satisfaction of the second historical demand instruction under that third-party voice skill; and

if a third-party voice skill is a non-resource skill, obtain the entry demand instruction by filtering the second historical demand instruction according to an interaction satisfaction of the second historical demand instruction under that third-party voice skill after it is determined that there is no playback resource record corresponding to the second historical demand instruction in the built-in voice skill.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method according to claim 1 .

13. A voice skill starting method, wherein the method is applied to an electronic device comprising at least one third-party voice skill and a built-in voice skill, and the method comprises:

obtaining a current demand instruction of a user;

judging whether the current demand instruction belongs to an entry demand instruction corresponding to the third-party voice skill according to a mapping relationship, wherein the mapping relationship is a pre-established mapping relationship between the third-party voice skill and the entry demand instruction, and the mapping relationship is determined according to skill-associated characteristic data of a first historical demand instruction under each third-party voice skill, the skill-associated characteristic data includes skill-associated entry characteristic data and skill satisfaction data, the skill-associated entry characteristic data is data associated with the third-party voice skill and representing a characteristic of the entry demand instruction, the skill satisfaction data indicates characteristic data of whether the third-party voice skill satisfies a second historical demand instruction, and the second historical demand instruction is a historical demand instruction filtered from the first historical demand instruction according to the skill-associated entry characteristic data; and

starting the third-party voice skill if it is determined that the current demand instruction belongs to the entry demand instruction corresponding to the third-party voice skill;

before the obtaining a current demand instruction of a user, further comprising:

obtaining the first historical demand instruction in the built-in voice skill;

determining the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill;

obtaining the entry demand instruction corresponding to each third-party voice skill by filtering the first historical demand instruction according to the skill-associated characteristic data; and

establishing the mapping relationship between each third-party voice skill and the corresponding entry demand instruction;

wherein the determining the skill-associated characteristic data of the first historical demand instruction under each third-party voice skill, comprises:

determining skill-associated entry characteristic data of the first historical demand instruction under each third-party voice skill; and

determining skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill;

wherein the skill satisfaction data comprises content satisfaction or interaction satisfaction, and the determining skill satisfaction data of the second historical demand instruction in the first historical demand instruction under each third-party voice skill, comprises:

judging whether each third-party voice skill is a resource skill;

if a third-party voice skill is the resource skill, determining a content satisfaction of the second historical demand instruction under that third-party voice skill; and

if a third-party voice skill is a non-resource skill, determining an interaction satisfaction of the second historical demand instruction under that third-party voice skill;

wherein the determining the content satisfaction of the second historical demand instruction under that third-party voice skill, comprises:

obtaining a first playback resource duration of the second historical demand instruction under that third-party voice skill and a second playback resource duration of the second historical demand instruction under the built-in voice skill; and

determining the content satisfaction according to the first playback resource duration and the second playback resource duration.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2020
From: ZHU, GUANGYA; ZHOU, XIAO
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 052388/0785 →
Priority Claims (1)
CN 201910809147.8 · Aug 29, 2019 · national
Continuity (1)
Related Publication 20210065707A1 · Mar 4, 2021