IP Library Granted Patent US 11,727,708
Granted Patent B2
US 11,727,708 · App. 17/711,596 · Granted Aug 15, 2023

Sectionizing documents based on visual and language models

Inventor: Kunling Geng (Milpitas, CA)
Assignee: Ciitizen, LLC
G06V30/416G06F40/258G06F40/30G06V30/10G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,708
App. No.
17/711,596
Granted
Aug 15, 2023
Kind
B2
Abstract

Some embodiments provide a program that receives a request to sectionize a document, uses a visual model to identify a set of candidate section headers in the document, and uses a language model to determine a type of section header for at least one candidate section header in the set of candidate section headers in the document. Some embodiments provide a program that receives a request to anonymize data in a document, uses a visual model to identify a set of candidate confidential sections in the document that are each predicted to include a collection of confidential data, uses a language model to identify terms in each candidate confidential section that are determined to be confidential data, analyzes the document to identify a set of terms in the document based on the identified terms in the set of candidate confidential sections, and redacts the set of terms in the document.

Claims (53)

1. A non-transitory machine-readable medium comprising instructions stored thereon that, when executed by a processor, cause the processor to:

receive a request to anonymize data in a document;

identify, using a first visual model, a set of candidate confidential sections in the document that are each predicted to include a collection of confidential data;

identify, using a first language model, confidential data in each candidate confidential section of the set of candidate confidential sections;

validate, using the first language model, that the set of candidate confidential sections comprises the collection of confidential data;

identify data in the document based on the identified confidential data in the set of candidate confidential sections; and

redact the data in the document, wherein the document is a document that has previously been sectionalized using a second visual model and a second language model.

2. The non-transitory machine-readable medium of claim 1 , wherein the instructions that cause the processor to identify the set of candidate confidential sections further comprise instructions to:

identify, using the first visual model, dimensions of bounding boxes encompassing text in the set of candidate confidential sections, locations of the bounding boxes in the document, and confidence scores corresponding to the bounding boxes.

3. The non-transitory machine-readable medium of claim 1 , wherein the first visual model receives an image of the document as an input to identify the set of candidate confidential sections in the document.

4. The non-transitory machine-readable medium of claim 1 , wherein the instructions to identify the confidential data in each candidate confidential section comprise instructions that cause the processor to identify, using the first language model, key-value pairs.

5. The non-transitory machine-readable medium of claim 1 , wherein the instructions to identify the confidential data in each candidate confidential section comprise instructions that cause the processor to:

extract the confidential data;

parse the confidential data; and

generate a data structure comprising the confidential data.

6. The non-transitory machine-readable medium of claim 1 , wherein the instructions to identify data in the document comprise instructions to execute a fuzzy matching algorithm to identify one or more references to the confidential data in the document.

7. The non-transitory machine-readable medium of claim 1 , further comprising instructions that cause the processor, prior to identifying a set of candidate confidential sections in the document, to sectionalize the document using the second visual model and the second language model.

8. A system comprising:

a set of processing units; and

a non-transitory machine-readable medium comprising instructions that, when executed by at least one processing unit in the set of processing units, cause the at least one processing unit to:

receive a request to anonymize data in a document;

identify, using a first visual model, a set of candidate confidential sections in the document that are each predicted to include a collection of confidential data;

identify, using a first language model, confidential data in each candidate confidential section of the set of candidate confidential sections;

validate, using the first language model, that the set of candidate confidential sections comprises the collection of confidential data;

identify data in the document based on the identified confidential data in the set of candidate confidential sections; and

redact the data in the document, wherein the document is a document that has previously been sectionalized using a second visual model and a second language model.

9. The system of claim 8 , wherein the instructions that cause the at least one processing unit to identify the set of candidate confidential sections comprise instructions to:

identify, using the first visual model, dimensions of bounding boxes encompassing text in the set of candidate confidential sections, locations of the bounding boxes in the document, and confidence scores corresponding to the bounding boxes.

10. The system of claim 8 , wherein the first visual model receives an image of the document as an input to identify the set of candidate confidential sections in the document.

11. The system of claim 8 , wherein the instructions to identify the confidential data in each candidate confidential section comprise instructions that cause the at least one processing unit to identify, using the first language model, key-value pairs.

12. The system of claim 8 , wherein the instructions to identify the confidential data in each candidate confidential section comprise instructions that cause the at least one processing unit to:

extract the confidential data;

parse the confidential data; and

generate a data structure comprising the confidential data.

13. The system of claim 8 , wherein the instructions to identify data in the document comprise instructions to execute a fuzzy matching algorithm to identify one or more references to the confidential data in the document.

14. The system of claim 8 , wherein the instructions further comprise instructions that cause the at least one processing unit, prior to identifying a set of candidate confidential sections in the document, to sectionalize the document using the second visual model and the second language model.

15. A method comprising:

receiving, by a processor, a request to anonymize data in a document;

identifying, using a first visual model implemented by the processor, a set of candidate confidential sections in the document that are each predicted to include a collection of confidential data;

identifying, using a first language model implemented by the processor, confidential data in each candidate confidential section of the set of candidate confidential sections;

validating, using the first language model, that the set of candidate confidential sections comprises the collection of confidential data;

identifying, by the processor, data in the document based on the identified confidential data in the set of candidate confidential sections; and

redacting the data in the document, wherein the document is a document that has previously been sectionalized using a second visual model and a second language model.

16. The method of claim 15 , wherein the identifying the set of candidate confidential sections further comprises:

identifying, using the first visual model implemented by the processor, dimensions of bounding boxes encompassing text in the set of candidate confidential sections, locations of the bounding boxes in the document, and confidence scores corresponding to the bounding boxes.

17. The method of claim 15 , wherein the first visual model receives an image of the document as an input to identify the set of candidate confidential sections in the document.

18. The method of claim 15 , wherein the identifying confidential data in each candidate confidential section comprises identifying, using the first language model implemented by the processor, key-value pairs.

19. The method of claim 15 , wherein the identifying confidential data in each candidate confidential section comprises:

extracting, by the processor, the confidential data;

parsing, by the processor, the confidential data; and

generating, by the processor, a data structure comprising the confidential data.

20. The method of claim 15 , wherein the identifying data in the document uses a fuzzy matching algorithm to identify one or more references to the confidential data in the document.

21. The method of claim 15 , further comprising, prior to the identifying a set of candidate confidential sections in the document, sectionalizing the document using the second visual model and the second language model.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: INVITAE CORPORATION; CIITIZEN, LLC
To: CITIZEN HEALTH, INC.
Reel/Frame 066087/0060 →
RELEASE OF SECURITY INTEREST AT R/F 63787/0148 Recorded Dec 14, 2023
From: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
To: CIITIZEN, LLC
Reel/Frame 066017/0791 →
SECURITY INTEREST Recorded Mar 7, 2023
From: CIITIZEN, LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 062907/0924 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: GENG, KUNLING
To: CIITIZEN CORP.
Reel/Frame 061830/0835 →
MERGER AND CHANGE OF NAME Recorded Nov 18, 2022
From: CIITIZEN CORPORATION; CAYMAN MERGER SUB B LLC
To: CIITIZEN, LLC
Reel/Frame 061830/0926 →
Continuity (2)
Continuation 16702394 · Dec 3, 2019
Related Publication 20220230465A1 · Jul 21, 2022
Cited By (7)
US 12,242,806 US 12,493,740 US 12,541,638 US 12,566,916 US 12,626,058 US 12,632,669 US 12,705,420