IP Library › Granted Patent US 11,461,317
Granted Patent B2
US 11,461,317 · App. 17/359,578 · Granted Oct 4, 2022

Method, apparatus, system, device, and storage medium for answering knowledge questions

Inventors: Xiexiong Lin (Hangzhou, CN); Jianshan He (Hangzhou, CN); Taifeng Wang (Hangzhou, CN)
Assignee: ALIPAY (HANGZHOU) INFORMATION TECHNOLOGY CO., LTD.
G06F16/245G06F16/285G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,317
App. No.
17/359,578
Granted
Oct 4, 2022
Kind
B2
Abstract

Embodiments of the present specification disclose a method, an apparatus, a system, a device, and a storage medium for answering user questions, including: obtaining a user question; encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector, wherein the structured data further comprises a data level, wherein the data level comprises knowledge for answering questions structured according to the schema level; retrieving one or more candidate sub-graphs related to the user question from the structured data; encoding the one or more candidate sub-graphs to obtain a second feature vector; performing multi-task classification for the user question based on the first feature vector and the second feature vector; and obtaining answer content for the user question based on a result of the multi-task classification.

Claims (80)

1. A method for answering user questions, comprising:

obtaining a user question;

encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector, wherein the structured data further comprises a data level, a plurality of sub-graphs according to topics of user questions, a table, and a knowledge graph, and wherein the data level comprises knowledge for answering questions structured according to the schema level;

retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data, comprises:

performing knowledge positioning on the table based on a topic of the user question, and positioning to each column whose column name matches the topic of the user question; and

for each positioned column whose column name matches the topic of the user question, retrieving a candidate sub-graph having the column name of the column as a center node of the candidate sub-graph and values in the column as adjacent nodes of the center node;

encoding the one or more candidate sub-graphs to obtain a second feature vector;

performing multi-task classification for the user question based on the first feature vector and the second feature vector; and

obtaining answer content for the user question based on a result of the multi-task classification.

2. The method of claim 1 , wherein the encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector comprises:

constructing a standard input text based on the user question and the schema level; and

encoding the standard input text using a self-encoding language model to obtain the first feature vector, wherein the first feature vector contains a vector expression of each element in the user question and a vector expression of each element in the schema level.

3. The method of claim 2 , wherein:

the constructing a standard input text based on the user question and the schema level comprises:

performing structure unification processing on the schema level of the table and the schema level of the knowledge graph to obtain a unified data structure; and

constructing a standard input structure based on the unified data structure and the user question to obtain the standard input text.

4. The method according to claim 1 , wherein the retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data comprises:

performing content understanding on content of the user question to obtain the topic of the user question; and

retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question.

5. The method according to claim 4 , wherein the retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question comprises:

performing knowledge positioning on the knowledge graph based on the topic of the user question, and positioning to each topic entity of the knowledge graph corresponding to the topic; and

for each positioned topic entity, retrieving a candidate sub-graph using a range of one or more hops of the topic entity.

6. The method according to claim 1 , wherein the performing multi-task classification for the user question based on the first feature vector and the second feature vector comprises:

mapping the user question into a structured query language statement, and dividing the user question into multiple sub-tasks based on the structured query language statement;

for each of the multiple sub-tasks, processing the first feature vector and the second feature vector through a corresponding task network in a multi-task classifier to obtain a sub- classification result for the sub-task; and

combining sub-classification results of all of the multiple sub-tasks to obtain the result of the multi-task classification.

7. A system for answering user questions, comprising a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor to cause the system to perform operations comprising:

obtaining a user question;

encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector, wherein the structured data further comprises a data level, a plurality of sub-graphs according to topics of user questions, a table, and a knowledge graph, and wherein the data level comprises knowledge for answering questions structured according to the schema level;

retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data, comprises:

performing knowledge positioning on the table based on a topic of the user question, and positioning to each column whose column name matches the topic of the user question; and

for each positioned column whose column name matches the topic of the user question, retrieving a candidate sub-graph having the column name of the column as a center node of the candidate sub-graph and values in the column as adjacent nodes of the center node;

encoding the one or more candidate sub-graphs to obtain a second feature vector;

performing multi-task classification for the user question based on the first feature vector and the second feature vector; and

obtaining answer content for the user question based on a result of the multi-task classification.

8. The system of claim 7 , wherein the encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector comprises:

constructing a standard input text based on the user question and the schema level; and

encoding the standard input text using a self-encoding language model to obtain the first feature vector, wherein the first feature vector contains a vector expression of each element in the user question and a vector expression of each element in the schema level.

9. The system of claim 8 , wherein:

the constructing a standard input text based on the user question and the schema level comprises:

performing structure unification processing on the schema level of the table and the schema level of the knowledge graph to obtain a unified data structure; and

constructing a standard input structure based on the unified data structure and the user question to obtain the standard input text.

10. The system of claim 7 , wherein the retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data comprises:

performing content understanding on content of the user question to obtain the topic of the user question; and

retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question.

11. The system of claim 10 , wherein

the retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question comprises:

performing knowledge positioning on the knowledge graph based on the topic of the user question, and positioning to each topic entity of the knowledge graph corresponding to the topic; and

for each positioned topic entity, retrieving a candidate sub-graph using a range of one or more hops of the topic entity.

12. The system of claim 7 , wherein the performing multi-task classification for the user question based on the first feature vector and the second feature vector comprises:

mapping the user question into a structured query language statement, and dividing the user question into multiple sub-tasks based on the structured query language statement;

for each of the multiple sub-tasks, processing the first feature vector and the second feature vector through a corresponding task network in a multi-task classifier to obtain a sub-classification result for the sub-task; and

combining sub-classification results of all of the multiple sub-tasks to obtain the result of the multi-task classification.

13. A non-transitory computer-readable storage medium for answering user questions, configured with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

obtaining a user question;

encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector, wherein the structured data further comprises a data level, a plurality of sub-graphs according to topics of user questions, a table, and a knowledge graph, and wherein the data level comprises knowledge for answering questions structured according to the schema level;

retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data, comprises:

performing knowledge positioning on the table based on a topic of the user question, and positioning to each column whose column name matches the topic of the user question; and

for each positioned column whose column name matches the topic of the user question, retrieving a candidate sub-graph using the column name of the column as a center node of the candidate sub-graph and values in the column as adjacent nodes of the center node;

encoding the one or more candidate sub-graphs to obtain a second feature vector;

performing multi-task classification for the user question based on the first feature vector and the second feature vector; and

obtaining answer content for the user question based on a result of the multi-task classification.

14. The medium of claim 13 , wherein the encoding the user question and a schema level of pre-constructed structured data to obtain a first feature vector comprises:

constructing a standard input text based on the user question and the schema level; and

encoding the standard input text using a self-encoding language model to obtain the first feature vector, wherein the first feature vector contains a vector expression of each element in the user question and a vector expression of each element in the schema level.

15. The medium of claim 14 , wherein:

the constructing a standard input text based on the user question and the schema level comprises:

performing structure unification processing on the schema level of the table and the schema level of the knowledge graph to obtain a unified data structure; and

constructing a standard input structure based on the unified data structure and the user question to obtain the standard input text.

16. The medium of claim 13 , wherein the retrieving one or more candidate sub-graphs related to the user question from the plurality of sub-graphs of the structured data comprises:

performing content understanding on content of the user question to obtain the topic of the user question; and

retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question.

17. The medium of claim 16 , wherein

the retrieving, from the plurality of sub-graphs of the structured data, one or more candidate sub-graphs matching the topic of the user question comprises:

performing knowledge positioning on the knowledge graph based on the topic of the user question, and positioning to each topic entity of the knowledge graph corresponding to the topic; and

for each positioned topic entity, retrieving a candidate sub-graph using a range of one or more hops of the topic entity.

18. The medium of claim 13 , wherein the performing multi-task classification for the user question based on the first feature vector and the second feature vector comprises:

mapping the user question into a structured query language statement, and dividing the user question into multiple sub-tasks based on the structured query language statement;

for each of the multiple sub-tasks, processing the first feature vector and the second feature vector through a corresponding task network in a multi-task classifier to obtain a sub-classification result for the sub-task; and

combining sub-classification results of all of the multiple sub-tasks to obtain the result of the multi-task classification.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2021
From: LIN, XIEXIONG; HE, JIANSHAN; WANG, TAIFENG
To: ALIPAY (HANGZHOU) INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 056679/0737 →
Priority Claims (1)
CN 202010632352.4 · Jul 3, 2020 · national
Continuity (1)
Related Publication 20220004547A1 · Jan 6, 2022