IP Library › Granted Patent US 11,461,613
Granted Patent B2
US 11,461,613 · App. 16/704,512 · Granted Oct 4, 2022

Method and apparatus for multi-document question answering

Inventors: Julien Perez (Grenoble, FR); Arnaud Sors (Grenoble, FR)
Assignee: NAVER CORPORATION
G06N3/004G06F16/93G06F16/953G06N3/0427G06N3/0445G06N3/0454G06N3/0472G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,613
App. No.
16/704,512
Granted
Oct 4, 2022
Kind
B2
Abstract

A computer implemented method for multi-document question answering is performed on a server communicating with a client device over a network. The method includes receiving a runtime question from a client device and retrieving runtime documents concerning the runtime question using a search engine. Runtime answer samples are identified in the retrieved runtime documents. A neural network model, trained using distant supervision and distance based ranking loss, is used to compute runtime scores from runtime question data representing the runtime question and from a runtime answer sample representing a first portion of text from the corpus of documents, where each runtime score represents a probability an answer to the runtime question is present in the runtime answer samples. A runtime answer is selected from the runtime answer samples corresponding to the highest runtime score sent to the client device.

Claims (65)

1. A computer implemented method for training a neural network model to compute a score representing a probability an answer to a predetermined question is present in a portion of text from a corpus of documents, the method comprising:

(a) computing a first training score using the neural network model from training question data representing a predetermined question and from a first training answer sample representing a first portion of text from the corpus of documents, where the first training answer sample is labelled as containing an answer to the predetermined question;

(b) repeating (a) for a plurality of first training answer samples which have been respectively labelled as containing an answer to the predetermined question to compute a first set of training scores;

(c) computing a second training score using the neural network model from the training question data and from a second training answer sample representing a second portion of text from the corpus of documents, where the second training answer sample is labelled as not containing an answer to the predetermined question;

(d) repeating (c) for a plurality of second training answer samples which have been respectively labelled as not containing an answer to the predetermined question to compute a second set of training scores;

(e) computing a loss using the first set of training scores and the second set of training scores by: (i) determining a minimum score among the first set of training scores; (ii) for each training score in second set of training scores, computing a training term that is a difference between the determined minimum score and each training score; and (iii) summing the training terms, wherein the loss is the sum of the training terms; and

(f) updating the neural network model using the loss.

2. The method of claim 1 , wherein computing the second training score further comprises summing a predetermined margin and the difference between the determined minimum score and the second training score.

3. The method of claim 1 , wherein the neural network model is trained with an adaptive subsampling layer configured to transform a variable size first matrix into a fixed size second matrix of a given size n×m, the size of the first matrix being greater than the size of the second matrix, the training with the adaptative subsampling layer further comprising:

partitioning the rows, respectively the columns, of the first matrix into n, respectively m, parts; and

subsampling one value from each partition in order to create the nxm second matrix.

4. The method of claim 1 , wherein the neural network model is trained using:

a first block comprising at least one layer;

a second block arranged in output of the first block, and comprising an adaptative subsampling layer; and

a fully connected neural network (FCNN) arranged in output of the second block.

5. The method of claim 4 , wherein the neural network model is trained using a BiLSTM layer for at least one layer of the first block.

6. The method of claim 1 , wherein the training is a distant supervision training.

7. The method of claim 1 , further comprising:

finding in a Q&A database an entry associating the predetermined question with a predetermined answer,

searching for the predetermined answer in the corpus of documents using a search engine,

labelling an answer sample from the corpus of documents to be input in the neural network model:

as containing an answer to the predetermined question whenever the predetermined answer is found by a search engine in a portion of text, and

as not containing an answer to the predetermined question whenever the predetermined answer is not found by the search engine in a portion of text.

8. The method of claim 7 , wherein each portion of text is a paragraph or shorter than a paragraph, for example a sentence.

9. Method of claim 8 , wherein the corpus of documents is an encyclopaedia database.

10. A computer implemented method performed on a server communicating with a client device over a network, comprising:

(A) receiving a runtime question from a client device;

(B) retrieving runtime documents concerning the runtime question using a search engine;

(C) identifying runtime answer samples in the retrieved runtime documents;

(D) computing using a neural network model runtime scores from runtime question data representing the runtime question and from a runtime answer sample representing a first portion of text from a corpus of documents, where each runtime score represents a probability an answer to the runtime question is present in the runtime answer samples;

(E) selecting a runtime answer from the runtime answer samples corresponding to the highest runtime score; and

(F) sending the runtime answer to the client device;

wherein the neural network model is trained by:

(a) computing a first training score using the neural network model from training question data representing a predetermined question and from a first training answer sample representing a first portion of text from the corpus of documents, where the first training answer sample is labelled as containing an answer to the predetermined question;

(b) repeating (a) for a plurality of first training answer samples which have been respectively labelled as containing an answer to the predetermined question to compute a first set of training scores;

(c) computing a second training score using the neural network model from the training question data and from a second training answer sample representing a second portion of text from the corpus of documents, where the second training answer sample is labelled as not containing an answer to the predetermined question;

(d) repeating (c) for a plurality of second training answer samples which have been respectively labelled as not containing an answer to the predetermined question to compute a second set of training scores;

(e) computing a loss using the first set of training scores and the second set of training scores by: (i) determining a minimum score among the first set of training scores; (ii) for each training score in second set of training scores, computing a training term that is a difference between the determined minimum score and each training score; and (iii) summing the training terms, wherein the loss is the sum of the training terms; and

(f) updating the neural network model using the loss.

11. The method of claim 10 , wherein computing the second training score further comprises summing a predetermined margin and the difference between the determined minimum score and the second training score.

12. The method of claim 10 , wherein the neural network model is trained with an adaptive subsampling layer configured to transform a variable size first matrix into a fixed size second matrix of a given size nxm, the size of the first matrix being greater than the size of the second matrix, the training with the adaptative subsampling layer further comprising:

partitioning the rows, respectively the columns, of the first matrix into n, respectively m, parts; and

subsampling one value from each partition in order to create the nxm second matrix.

13. The method of claim 10 , wherein the neural network model is trained using:

a first block comprising at least one layer;

a second block arranged in output of the first block, and comprising the adaptative subsampling layer; and

a fully connected neural network (FCNN) arranged in output of the second block.

14. The method of claim 10 , wherein the neural network model is trained using a BiLSTM layer for at least one layer of the first block.

15. The method of claim 10 , wherein the training is a distant supervision training.

16. The method of claim 10 , further comprising:

finding in a Q&A database an entry associating the predetermined question with a predetermined answer,

searching for the predetermined answer in the corpus of documents using a search engine,

labelling an answer sample from the corpus of documents to be input in the neural network model:

as containing an answer to the predetermined question whenever the predetermined answer is found by a search engine in a portion of text, and

as not containing an answer to the predetermined question whenever the predetermined answer is not found by the search engine in a portion of text.

17. The method of claim 16 , wherein each portion of text is a paragraph or shorter than a paragraph, for example a sentence.

18. The method of claim 17 , wherein the corpus of documents is an encyclopaedia database.

19. A computer implemented method for multi-document question answering, comprising:

receiving a runtime question at server from a client device;

retrieving documents concerning the runtime question using a search engine;

identifying portions of text in the retrieved documents;

computing a runtime score for the identified portions of text using a neural network model trained using distant supervision and distance based ranking loss;

selecting the portions of text corresponding to the highest score to provide an answer; and

sending from the server the answer to the client device.

20. The method of claim 19 , wherein computing the runtime score for the identified portions of text further comprises using the neural network model combines an interaction matrix, Weaver blocks and adaptive subsampling to map to a fixed size representation vector to scores.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2019
From: PEREZ, JULIEN; SORS, ARNAUD
To: NAVER CORPORATION
Reel/Frame 051197/0913 →
Continuity (1)
Related Publication 20210174161A1 · Jun 10, 2021
Cited By (1)
US 12,347,426