IP Library Granted Patent US 11,468,238
Granted Patent B2
US 11,468,238 · App. 16/676,174 · Granted Oct 11, 2022

Data processing systems and methods

Inventors: Mitul Tiwari (Mountain View, CA); Ravi Narasimhan Raj (Los Altos, CA); Madhusudan Mathihalli (Saratoga, CA); Kaushik Rangadurai (Sunnyvale, CA); Srivatsava Daruru (Mountain View, CA); Quaizar Vohra (Cupertino, CA); Deepak Bobbarjung (Sunnyvale, CA); Abhisaar Yadav (Los Altos, CA)
Assignee: ServiceNow Inc.
G06F40/289G06F40/205G06F40/247G06F40/253G06F40/30G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,468,238
App. No.
16/676,174
Granted
Oct 11, 2022
Kind
B2
Abstract

Example data processing systems and methods are described. In one implementation, a system accesses a corpus of data and analyzes the data contained in the corpus of data to identify multiple documents. The system generates vector indexes for the multiple documents such that the vector indexes allow a computing system to quickly access the plurality of documents and identify an answer to a question associated with the corpus of data.

Claims (42)

1. A method of generating an utterance, the method comprising:

paraphrasing, by a computing system, a portion of a document using at least two different paraphrasing techniques to generate a plurality of candidate paraphrases;

selecting, by the computing system, relevant paraphrases from the plurality of candidate paraphrases by:

filtering the plurality of candidate paraphrases, and

performing de-duplication of the plurality of candidate paraphrases using vector representations of the sentence and the plurality of candidate paraphrases determined via a neural sentence encoder of the computing system, wherein one or more candidate paraphrases are discarded based on computed cosine similarities between the vector representations; and

selecting, by the computing system, the utterance from the relevant paraphrases.

2. The method of claim 1 , further comprising applying extractive summarization to select the portion of the document from best text portions of the document.

3. The method of claim 1 , wherein the at least two different paraphrasing techniques include full backtranslation using a neural machine translation model.

4. The method of claim 1 , wherein the at least two different paraphrasing techniques include noun/verb phrase backtranslation.

5. The method of claim 1 , wherein the at least two different paraphrasing techniques include synonym replacement.

6. The method of claim 1 , wherein the at least two different paraphrasing techniques include phrase replacement.

7. The method of claim 1 , wherein de-duplication comprises: determining, via the neural sentence encoder of the computing system, a first vector representation of a portion of the document and a respective vector representation of each of the plurality of candidate paraphrases; and discarding the one or more candidate paraphrases based on a computed respective cosine similarity between the first vector representation and the respective vector representation of each of the plurality of candidate paraphrases.

8. The method of claim 7 , wherein discarding the one or more candidate paraphrases comprises:

discarding a particular candidate paraphrase if the respective cosine similarity is less than 0.5.

9. The method of claim 7 , wherein discarding the one or more candidate paraphrases comprises:

discarding a particular candidate paraphrase if the respective cosine similarity is greater than 0.95.

10. The method of claim 1 , comprising:

performing token-based de-duplication of the plurality of candidate paraphrases.

11. The method of claim 1 , wherein filtering the plurality of candidate paraphrases includes removing irrelevant sentences.

12. A method of generating an utterance, the method comprising:

paraphrasing, by a computing system, a sentence of a document using at least two different paraphrasing techniques to generate a plurality of candidate paraphrases;

selecting, by the computing system, relevant paraphrases from the plurality of candidate paraphrases by:

filtering the plurality of candidate paraphrases; and

performing de-duplication of the plurality of candidate paraphrases by:

determining, via a neural sentence encoder of the computing system, a first vector representation of the sentence and a respective vector representation for each of the plurality of candidate paraphrases,

computing a respective cosine similarity between the first vector representation and the respective vector representation of each of the plurality of candidate paraphrases, and

discarding one or more of the plurality of candidate paraphrases based on the cosine similarities; and

selecting, by the computing system, the utterance from the relevant paraphrases.

13. The method of claim 12 , wherein the at least two different paraphrasing techniques include at least two of full backtranslation, noun/verb phrase backtranslation, synonym replacement, and phrase replacement.

14. The method of claim 12 , further comprising discarding a particular paraphrase if the respective cosine similarity is less than 0.5 or greater than 0.95.

15. The method of claim 12 , wherein filtering the plurality of candidate paraphrases includes removing irrelevant sentences.

16. The method of claim 12 , comprising:

applying extractive summarization to select the sentence from best text portions of the document.

17. A method of generating utterances, the method comprising:

paraphrasing, by a computing system, a portion of a document using at least two different paraphrasing techniques to generate a plurality of candidate paraphrases;

de-duplicating, by the computing system, the plurality of candidate paraphrases using vector representations determined by a neural sentence encoder of the computing system for the portion of the document and for each of the plurality of candidate paraphrases, wherein one or more candidate paraphrases are discarded based on computed cosine similarities between the vector representations; and

selecting, by the computing system, the utterances from the plurality of candidate paraphrases.

18. The method of claim 17 , comprising:

removing a particular candidate paraphrase from the plurality of candidate paraphrases when a respective computed cosine similarity between a vector representation of the portion of the document and a vector representation of the particular candidate paraphrase is less than 0.5 or greater than 0.95.

19. The method of claim 17 , wherein, prior to de-duplicating, the method comprises:

filtering, by the computing system, the plurality of candidate paraphrases to remove irrelevant paraphrases.

20. The method of claim 17 , wherein the at least two paraphrasing techniques include full backtranslation via a neural machine translation model, noun/verb phrase backtranslation via a neural parser, or a combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2019
From: TIWARI, MITUL; RAJ, RAVI NARASIMHAN; MATHIHALLI, MADHUSUDAN; RANGADURAI, KAUSHIK; DARURU, SRIVATSAVA; VOHRA, QUAIZAR; BOBBARJUNG, DEEPAK; YADAV, ABHISAAR
To: RUPERT LABS INC. (DBA PASSAGE AI)
Reel/Frame 050935/0641 →
Continuity (1)
Related Publication 20210133251A1 · May 6, 2021
Cited By (1)
US 12,614,035