IP Library Granted Patent US 11,775,762
Granted Patent B1
US 11,775,762 · App. 17/127,708 · Granted Oct 3, 2023

Data comparision using natural language processing models

Inventors: Erica K. Mason (San Francisco, CA); Apoorv Sharma (San Francisco, CA); Andy Mahdavi (San Francisco, CA); Daniel Faddoul (San Francisco, CA)
Assignee: States Title, LLC
G06F40/295G06F16/3347G06F40/284G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,762
App. No.
17/127,708
Granted
Oct 3, 2023
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for natural language processing. One of the methods includes the steps of receiving a first set of labeled data from a first data source; receiving a text string from a second data source; performing natural language processing on the text string to extract particular text portions and generate a second set of labeled data; performing a comparison between the first set of labeled data and the second set of labeled data; and generating an output based on the comparison.

Claims (55)

1. A method comprising:

receiving a first set of labeled data from a first data source;

receiving a text string from a second data source corresponding to one or more particular fields of a document;

performing natural language processing on the text string to extract particular text portions and generate a second set of labeled data;

determining, from the second set of labeled data, first text portions associated with vesting information, wherein determining the first text portions associated with the vesting information comprises comparing the particular text portions with entries in a domain knowledge base corresponding to particular vesting types;

evaluating the second set of labeled data for consistency with a vesting type corresponding to the first text portions;

in response to determining consistency, performing a comparison between the first set of labeled data and the second set of labeled data; and

generating an output based on the comparison.

2. The method of claim 1 , wherein performing natural language processing comprises:

tokenizing the text string into individual words and punctuation marks; and

transforming each token into a word vector representation of the word comprising a string of numbers that represent the semantic meaning of the token.

3. The method of claim 2 , further comprising:

performing parts of speech tagging using the word vectors to label each token with a particular part of speech;

performing name entity recognition using the word vectors to label tokes as corresponding to particular entities; and

using the outputs of the parts of speech tagging and name entity recognition to perform domain phrase matching.

4. The method of claim 1 , wherein prior to performing the comparison, performing one or more pre-comparison checks for internal consistency in one or more of the respective first and second sets of labeled data.

5. The method of claim 4 , wherein the comparison is performed in response to a determination that the data is consistent.

6. The method of claim 1 , wherein performing the comparison between the first set of labeled data and the second set of labeled data comprises applying a similarity function to content from the first data source and the second data source to calculate a similarity score and comparing the similarity score to a specified threshold.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving a first set of labeled data from a first data source;

receiving a text string from a second data source corresponding to one or more particular fields of a document;

performing natural language processing on the text string to extract particular text portions and generate a second set of labeled data;

determining, from the second set of labeled data, first text portions associated with vesting information, wherein determining the first text portions associated with the vesting information comprises comparing the particular text portions with entries in a domain knowledge base corresponding to particular vesting types;

evaluating the second set of labeled data for consistency with a vesting type corresponding to the first text portions;

in response to determining consistency, performing a comparison between the first set of labeled data and the second set of labeled data; and

generating an output based on the comparison.

8. The system of claim 7 , wherein performing natural language processing comprises:

tokenizing the text string into individual words and punctuation marks; and

transforming each token into a word vector representation of the word comprising a string of numbers that represent the semantic meaning of the token.

9. The system of claim 8 , the operations further comprising:

performing parts of speech tagging using the word vectors to label each token with a particular part of speech;

performing name entity recognition using the word vectors to label tokes as corresponding to particular entities; and

using the outputs of the parts of speech tagging and name entity recognition to perform domain phrase matching.

10. The system of claim 7 , wherein prior to performing the comparison, performing one or more pre-comparison checks for internal consistency in one or more of the respective first and second sets of labeled data.

11. The system of claim 10 , wherein the comparison is performed in response to a determination that the data is consistent.

12. The system of claim 7 , wherein performing the comparison between the first set of labeled data and the second set of labeled data comprises applying a similarity function to content from the first data source and the second data source to calculate a similarity score and comparing the similarity score to a specified threshold.

13. One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a first set of labeled data from a first data source;

receiving a text string from a second data source corresponding to one or more particular fields of a document;

performing natural language processing on the text string to extract particular text portions and generate a second set of labeled data;

determining, from the second set of labeled data, first text portions associated with vesting information, wherein determining the first text portions associated with the vesting information comprises comparing the particular text portions with entries in a domain knowledge base corresponding to particular vesting types;

evaluating the second set of labeled data for consistency with a vesting type corresponding to the first text portions;

in response to determining consistency, performing a comparison between the first set of labeled data and the second set of labeled data; and

generating an output based on the comparison.

14. The computer storage media of claim 13 , wherein performing natural language processing comprises:

tokenizing the text string into individual words and punctuation marks; and

transforming each token into a word vector representation of the word comprising a string of numbers that represent the semantic meaning of the token.

15. The computer storage media of claim 14 , further comprising:

performing parts of speech tagging using the word vectors to label each token with a particular part of speech;

performing name entity recognition using the word vectors to label tokes as corresponding to particular entities; and

using the outputs of the parts of speech tagging and name entity recognition to perform domain phrase matching.

16. The computer storage media of claim 13 , wherein prior to performing the comparison, performing one or more pre-comparison checks for internal consistency in one or more of the respective first and second sets of labeled data.

17. The computer storage media of claim 16 , wherein the comparison is performed in response to a determination that the data is consistent.

18. The computer storage media of claim 13 , wherein performing the comparison between the first set of labeled data and the second set of labeled data comprises applying a similarity function to content from the first data source and the second data source to calculate a similarity score and comparing the similarity score to a specified threshold.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2024
From: STATES TITLE, LLC
To: DOMA TECHNOLOGY LLC
Reel/Frame 069659/0478 →
RELEASE OF SECURITY INTEREST Recorded Sep 30, 2024
From: HUDSON STRUCTURED CAPITAL MANAGEMENT LTD.
To: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
Reel/Frame 068742/0095 →
RELEASE OF SECURITY INTEREST Recorded Sep 30, 2024
From: ALTER DOMUS (US) LLC
To: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
Reel/Frame 068742/0050 →
SECURITY INTEREST Recorded May 3, 2024
From: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
To: ALTER DOMUS (US) LLC
Reel/Frame 067312/0857 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2021
From: FADDOUL, DANIEL
To: STATES TITLE, LLC
Reel/Frame 057744/0024 →
CHANGE OF NAME Recorded Oct 8, 2021
From: STATES TITLE, INC.
To: STATES TITLE, LLC
Reel/Frame 057759/0088 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: MASON, ERICA K.; SHARMA, APOORV; MAHDAVI, ANDY
To: STATES TITLE, INC.
Reel/Frame 055236/0813 →
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBERS 10255550, 10510009 AND 10755184 TO PATENT NUMBERS 10255550, 10510009 AND 10755184 PREVIOUSLY RECORDED ON REEL 054804 FRAME 0211. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Jan 8, 2021
From: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
To: HUDSON STRUCTURED CAPITAL MANAGEMENT LTD
Reel/Frame 056322/0310 →
SECURITY INTEREST Recorded Jan 5, 2021
From: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
To: HUDSON STRUCTURED CAPITAL MANAGEMENT LTD.
Reel/Frame 054812/0286 →
SECURITY INTEREST Recorded Jan 4, 2021
From: STATES TITLE HOLDING, INC.; TITLE AGENCY HOLDCO, LLC
To: HUDSON STRUCTURED CAPITAL MANAGEMENT LTD.
Reel/Frame 054804/0211 →
Cited By (3)
US 12,307,216 US 12,488,195 US 12,639,433