IP Library Granted Patent US 11,055,327
Granted Patent B2
US 11,055,327 · App. 16/225,268 · Granted Jul 6, 2021

Unstructured data parsing for structured information

Inventor: Kasper Sørensen (Seattle, WA)
Assignee: QUADIENT TECHNOLOGIES FRANCE
G06F16/313G06F16/353G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,055,327
App. No.
16/225,268
Granted
Jul 6, 2021
Kind
B2
Abstract

Systems and methods are provided for computerized, automatic processing of unstructured text to extract bits of contact data to which the extracted text can be linked or attributed. Unstructured text is received and text segments within the text are enriched with metadata labels. A machine learning system is trained on, and used to parse feature values for the text segments and the metadata labels to classify text and generate structured text from the unstructured text.

Claims (93)

1. A method comprising:

receiving unstructured text;

identifying a plurality of text segments within the unstructured text and enriching the unstructured text by assigning one or more metadata labels to each of the text segments;

calculating one or more feature values for each text segment for at least one of the metadata labels associated with at least one of the text segments;

providing the feature values, metadata labels, and text segments to a trained machine learning system;

receiving, from the machine learning system, an indication of a classification for at least one of the text segments;

based upon the indication of the classification for the at least one of the text segments, identifying the at least one of the text segments as a portion of contact data;

receiving, from the machine learning system, a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generating structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

2. The method of claim 1 , wherein the step of assigning one or more metadata labels to each of the text segments comprises:

processing each text segment of the text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning the one or more metadata labels to the each text segment.

3. The method of claim 2 , wherein the at least one enrichment source is selected from the group consisting of: a dictionary of name parts; a pattern matching system; an address-matching library; and an external contact information service.

4. The method of claim 3 , wherein the at least one enrichment source comprises a plurality of different enrichment sources, each enrichment source independently selected from the group consisting of: a dictionary of name parts; a pattern matching system; an address-matching library; and an external contact information service.

5. The method of claim 1 , wherein the structured contact information includes only text segments having a confidence score above a minimum confidence level.

6. The method of claim 1 , further comprising:

processing at least a portion of the structured contact information for the entity with a domain expert module; and

based upon the processing, generating revised structured contact information for the entity.

7. The method of claim 1 , further comprising:

providing the structure contact information for the entity and structured contact information for a plurality of other entities to a contact information database.

8. A system comprising:

a processor configured to:

receive unstructured text;

identify a plurality of text segments within the unstructured text and enriching the unstructured text by assigning one or more metadata labels to each of the text segments;

calculate one or more feature values for each text segment for at least one of the metadata labels associated with at least one of the text segments;

provide the feature values, metadata labels, and text segments to a trained machine learning system;

receive, from the machine learning system, an indication of a classification for at least one of the text segments;

identify the at least one of the text segments as a portion of contact data based upon the indication of the classification for the at least one of the text segments;

provide a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generate structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

9. The system of claim 8 , further comprising a computer-based system configured to execute the machine learning system.

10. The system of claim 9 , further comprising a database storing training data to train the machine learning system.

11. The system of claim 8 , wherein the step of assigning one or more metadata labels to each of the text segments comprises:

processing each text segment of the text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning the one or more metadata labels to the each text segment.

12. The system of claim 11 , wherein the at least one enrichment source is selected from the group consisting of: a dictionary of name parts; a pattern matching system; an address-matching library; and an external contact information service.

13. The system of claim 12 , wherein the at least one enrichment source comprises a plurality of different enrichment sources, each enrichment source independently selected from the group consisting of: a dictionary of name parts; a pattern matching system; an address-matching library; and an external contact information service.

14. The system of claim 8 , wherein the processor is further configured to:

process at least a portion of the structured contact information for the entity with a domain expert module; and

based upon the processing, generate revised structured contact information for the entity.

15. The system of claim 8 , wherein the processor is further configured to:

provide the structure contact information for the entity and structured contact information for a plurality of other entities to a contact information database.

16. A non-transitory computer-readable storage medium storing a plurality of instructions that, when executed, cause a processor to:

receive unstructured text;

identify a plurality of text segments within the unstructured text and enriching the unstructured text by assigning one or more metadata labels to each of the text segments;

calculate one or more feature values for each text segment for at least one of the metadata labels associated with at least one of the text segments;

provide the feature values, metadata labels, and text segments to a trained machine learning system;

receive, from the machine learning system, an indication of a classification for at least one of the text segments;

identify the at least one of the text segments as a portion of contact data based upon the indication of the classification for the at least one of the text segments;

provide a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generate structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the step of assigning one or more metadata labels to each of the text segments comprises:

processing each text segment of the text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning the one or more metadata labels to the each text segment.

18. The non-transitory computer-readable storage medium of claim 16 , the plurality of instructions, when executed, further causing the processor to:

process at least a portion of the structured contact information for the entity with a domain expert module; and

based upon the processing, generate revised structured contact information for the entity.

19. The non-transitory computer-readable storage medium of claim 16 , the plurality of instructions, when executed, further causing the processor to:

provide the structure contact information for the entity and structured contact information for a plurality of other entities to a contact information database.

20. A method comprising:

receiving unstructured text;

identifying a plurality of text segments within the unstructured text and enriching the unstructured text by:

processing each text segment of the plurality of text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning one or more metadata labels to the each text segment;

calculating one or more feature values for each text segment;

providing the feature values, metadata labels, and text segments to a trained machine learning system;

receiving, from the machine learning system, an indication of a classification for at least one of the text segments;

based upon the indication of the classification for the at least one of the text segments, identifying the at least one of the text segments as a portion of contact data;

receiving, from the machine learning system, a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generating structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

21. A system comprising:

a processor configured to:

receive unstructured text;

identify a plurality of text segments within the unstructured text and enriching the unstructured text by:

processing each text segment of the plurality of text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning the one or more metadata labels to the each text segment;

calculate one or more feature values for each text segment;

provide the feature values, metadata labels, and text segments to a trained machine learning system;

receive, from the machine learning system, an indication of a classification for at least one of the text segments;

identify the at least one of the text segments as a portion of contact data based upon the indication of the classification for the at least one of the text segments;

provide a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generate structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

22. A non-transitory computer-readable storage medium storing a plurality of instructions that, when executed, cause a processor to:

receive unstructured text;

identify a plurality of text segments within the unstructured text and enriching the unstructured text by:

processing each text segment of the text segments with at least one enrichment source; and

based upon an output of the at least one enrichment source, assigning the one or more metadata labels to the each text segment;

calculate one or more feature values for each text segment;

provide the feature values, metadata labels, and text segments to a trained machine learning system;

receive, from the machine learning system, an indication of a classification for at least one of the text segments;

identify the at least one of the text segments as a portion of contact data based upon the indication of the classification for the at least one of the text segments;

provide a confidence score for each of the classifications, wherein the at least one of the text segments is identified as a portion of contact data based upon a confidence score assigned to the classification of the at least one text segment; and

generate structured contact information for an entity based on the confidence score for each of the classifications and the text segments.

Assignments (2)
CHANGE OF NAME Recorded Dec 12, 2022
From: NEOPOST TECHNOLOGIES
To: QUADIENT TECHNOLOGIES FRANCE
Reel/Frame 062226/0973 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2019
From: SØRENSEN, KASPER
To: NEOPOST TECHNOLOGIES
Reel/Frame 048017/0782 →