IP Library › Granted Patent US 11,809,477
Granted Patent B1
US 11,809,477 · App. 17/994,854 · Granted Nov 7, 2023

Topic focused related entity extraction

Inventors: Pallabi Ghosh (St. Augustine, FL); Sparsh Gupta (San Diego, CA)
Assignee: Intuit Inc.
G06F16/38
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,809,477
App. No.
17/994,854
Granted
Nov 7, 2023
Kind
B1
Abstract

This disclosure relates to extracting entities from unstructured text. The unstructured text is segmented into structured segments with one or more instances, that belong to different topics, with a topic segmentation model. Each instances of the structured segment is operated on by an entity extraction model to extract entities, and the extracted entities associated with each topic is produced in a computer-readable format. The relations between extracted entities associated with each topic may be identified.

Claims (30)

1. A computer-implemented method for extracting text information from an electronic document, comprising:

obtaining unstructured text of the electronic document;

segmenting the unstructured text into structured segments with one or more instances, belonging to different topics, with a topic segmentation model;

extracting entities from each instance of the structured segments with an entity extraction model, wherein extracting entities from each instance of the structured segments comprises operating on a first segment to extract all entities from each instance in the first segment before operating on a second segment; and

producing extracted entities associated with each topic in a computer-readable format.

2. The computer-implemented method of claim 1 , wherein one or more of the different topics includes multiple sub-sections, and wherein the extracted entities are associated with each sub-section within each topic.

3. The computer-implemented method of claim 1 , wherein entities are assigned to each topic, wherein extracting entities from each instance of the structured segments comprises extracting only entities that are related to a topic associated with each respective structured segment.

4. The computer-implemented method of claim 1 , wherein producing the extracted entities associated with each topic in the computer-readable format comprises identifying relations between the extracted entities associated with each topic.

5. The computer-implemented method of claim 1 , wherein the topic segmentation model comprises at least one of SECTOR, Transformer2, and BeamSeg, and the entity extraction model comprises at least one of a few-shot learning model, Named-Entity Recognition (NER), Generative Pre-trained Transformer 3 (GPT3), and T5 (Text-to-Text Transfer Transformer).

6. A system for extracting text information from an electronic document, comprising:

one or more processors; and

a memory coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

obtain unstructured text of the electronic document;

segment the unstructured text into structured segments with one or more instances, belonging to different topics, with a topic segmentation model;

extract entities from each instance of the structured segments with an entity extraction model, wherein the system is caused to extract entities from each instance of the structured segments by being caused to operate on a first segment to extract all entities from each instance in the first segment before operating on a second segment; and

produce extracted entities associated with each topic in a computer-readable format.

7. The system of claim 6 , wherein one or more of the different topics includes multiple sub-sections, and wherein the extracted entities are associated with each sub-section within each topic.

8. The system of claim 6 , wherein entities are assigned to each topic,

wherein the system is caused to extract entities from each instance of the structured segments by being caused to extract only entities that are related to a topic associated with each respective structured segment.

9. The system of claim 6 , wherein the system is caused to produce the extracted entities associated with each topic in the computer-readable format by being caused to identify relations between the extracted entities associated with each topic.

10. The system of claim 6 , wherein the topic segmentation model comprises at least one of SECTOR, Transformer2, and BeamSeg, and the entity extraction model comprises at least one of a few-shot learning model, Named-Entity Recognition (NER), Generative Pre-trained Transformer 3 (GPT3), and T5 (Text-to-Text Transfer Transformer).

11. A system for extracting text information from an electronic document, comprising:

an interface configured to obtain unstructured text of the electronic document;

a topic segmentation model configured to segment the unstructured text into structured segments, with one or more instances, belonging to different topics; and

an entity extraction model configured to extract entities from each instance of the structured segments and to produce extracted entities associated with each topic in a computer-readable format, wherein the entity extraction model is configured to extract entities from each instance of the structured segments by operating on a first segment to extract all entities from each instance in the first segment before operating on a second segment.

12. The system of claim 11 , wherein one or more of the different topics includes multiple sub-sections, and wherein the extracted entities are associated with each sub-section within each topic.

13. The system of claim 11 , wherein entities are assigned to each topic,

wherein the entity extraction model is configured to extract entities from each instance of the structured segments by extracting only entities that are related to a topic associated with each respective structured segment.

14. The system of claim 11 , wherein entity extraction model is configured to produce extracted entities associated with each topic in the computer-readable format by identifying relations between the extracted entities associated with each topic.

15. The system of claim 11 , wherein the topic segmentation model comprises at least one of SECTOR, Transformer2, and BeamSeg, and the entity extraction model comprises at least one of a few-shot learning model, Named-Entity Recognition (NER), Generative Pre-trained Transformer 3 (GPT3), and T5 (Text-to-Text Transfer Transformer).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2022
From: GHOSH, PALLABI; GUPTA, SPARSH
To: INTUIT INC.
Reel/Frame 061891/0928 →
Cited By (1)
US 12,462,101