IP Library Granted Patent US 9,158,833
Granted Patent B2
US 9,158,833 · App. 12/610,937 · Granted Oct 13, 2015

System and method for obtaining document information

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,158,833
App. No.
12/610,937
Granted
Oct 13, 2015
Kind
B2
Abstract

A method and system for determining at least one target value of at least one target in at least one document, comprising: determining, utilizing at least one scoring application; at least one possible target value, wherein the at least one scoring application utilizes information from at least one training document, and applying the information, utilizing the at least one scoring application, on the at least one new document to determine at least one value of the at least one target on the at least one new document.

Claims (117)

1. A method for determining at least one target value of at least one target in at least one document, comprising:

determining, utilizing at least one scoring application, target position information from at least one training document;

using the target position information, utilizing at least one localization module, the using comprising:

finding at least one reference;

creating at least one reference vector for each reference;

performing variance filtering on the at least one reference and the at least one reference vector from each document to obtain any similar references and any similar reference vectors from all documents;

using the any similar references and the any similar reference vectors to create at least one dynamic variance network, wherein the at least one dynamic variance network is a set comprising the at least one reference and the at least one reference vector tying each reference of the set to the at least one target; and

applying the target position information, utilizing the at least one scoring application, on at least one new document to determine the at least one target value of the at least one target on the at least one new document, wherein the target position information comprises the at least one reference within the training document and the at least one reference vector, each reference vector tying the at least one reference of the at least one references to the at least one target, each reference vector comprising content data indicating at least one intrinsic content type.

2. The method of claim 1 , further comprising at least one additional scoring application utilizing:

(a) information comprising at least one position of the at least one target in the at least one training document;

(b) format information and possible variation format information for the at least one target in the at least one training document; or

(c) any combination thereof.

3. The method of claim 1 , further comprising applying at least one document classifier to the at least one new document.

4. The method of claim 1 , wherein the variance filtering further comprises:

comparing, utilizing the at least one localization module, the any similar references to the at least one reference on the at least one new document to determine if there are any matching references; and

using, utilizing the at least one localization module, the any similar reference vectors corresponding to any matching references to determine the at least one target on the at least one new document.

5. The method of claim 1 , wherein the at least one reference comprises:

at least one character string,

at least one word;

at least one number;

at least one alpha-numeric representation;

at least one token;

at least one blank space;

at least one logo; or

at least one text fragment; or

any combination thereof.

6. The method of claim 1 , wherein at least one position of the at least one target is used to obtain and/or confirm information about the target.

7. The method of claim 1 , wherein the at least one reference comprises: a typo, an OCR mistake, or an alternate spelling, or any combination thereof; but the at least one reference is still used as a reference because of the at least one reference's position.

8. The method of claim 1 , wherein the any similar reference vectors can be: positionally similar; content similar; or type similar; or any combination thereof.

9. The method of claim 1 , wherein similarity across the any similar references and the any similar reference vectors is configurable.

10. The method of claim 1 , wherein strict and/or fuzzy matching can be utilized to match the any similar references to the at least one reference in the at least one new document.

11. The method of claim 10 , wherein the following characteristics of the at least one reference are taken into account: font; font size; style; or any combination thereof.

12. The method of claim 1 , wherein the at least one reference is: merged with at least one other reference; and/or split into at least two references.

13. The method of claim 1 , wherein the at least one dynamic variance network is dynamically adapted during document processing.

14. The method of claim 1 , wherein the at least one dynamic variance network is used for:

reference correction;

document classification;

page separation;

recognition of document modification;

document summarization; or

document compression;

or any combination thereof.

15. The method of claim 1 , wherein other information is also utilized to determine the at least one target value, the other information comprising:

format information and possible variations of the format information; and/or

key word information related to the at least one target.

16. The method of claim 1 , wherein the at least one intrinsic content type comprises:

a word;

a number;

a combination of letters and numbers;

a number and punctuation string;

an optical character recognition mistake;

a word in a different language from at least one other word in the at least one document;

a word found in a dictionary;

a word not found in a dictionary;

a font type;

a font size; or

a font property; or

a combination thereof.

17. A system for determining at least one target value of at least one target in at least one document, comprising:

at least one processor, wherein the at least one processor is configured for:

determining, utilizing at least one scoring application, target position information from at least one training document;

using the target position information, utilizing at least one localization module, the using comprising:

finding at least one reference;

creating at least one reference vector for each reference;

performing variance filtering on the at least one reference and the at least one reference vector from each document to obtain any similar references and any similar reference vectors from all documents;

using the any similar references and the any similar reference vectors to create at least one dynamic variance network, wherein the at least one dynamic variance network is a set comprising the at least one reference and the at least one reference vector tying each reference of the set to the at least one target; and

applying the target position information, utilizing the at least one scoring application, on at least one new document to determine the at least one target value of the at least one target on the at least one new document, wherein the target position information comprises the at least one reference within the training document and the at least one reference vector, each reference vector tying the at least one reference to the at least one target, each reference vector comprising content data indicating at least one intrinsic content type.

18. The system of claim 17 , wherein the processor is further configured for utilizing at least one additional scoring application for:

(a) information comprising at least one position of the at least one target in the at least one training document;

(b) format information and possible variation format information for the at least one target in the at least one training document; or

(c) any combination thereof.

19. The system of claim 17 , further comprising applying at least one document classifier to the at least one new document.

20. The system of claim 17 , wherein the variance filtering further comprises:

comparing, utilizing the at least one localization module, the any similar references to the at least one reference on the at least one new document to determine if there are any matching references; and

using, utilizing the at least one localization module, the any similar reference vectors corresponding to any matching references to determine the at least one target on the at least one new document.

21. The system of claim 17 , wherein the at least one reference comprises:

at least one character string,

at least one word;

at least one number;

at least one alpha-numeric representation;

at least one token;

at least one blank space;

at least one logo; or

at least one text fragment; or

any combination thereof.

22. The system of claim 17 , wherein at least one position of the at least one target is used to obtain and/or confirm information about the target.

23. The system of claim 17 , wherein the at least one reference comprises: a typo, an OCR mistake, or an alternate spelling, or any combination thereof; but the at least one reference is still used as a reference because of the at least one reference's position.

24. The system of claim 17 , wherein the any similar reference vectors can be: positionally similar; content similar; or type similar; or any combination thereof.

25. The system of claim 17 , wherein similarity across the any similar references and the any similar reference vectors is configurable.

26. The system of claim 17 , wherein strict and/or fuzzy matching can be utilized to match the any similar references to the at least one reference in the at least one new document.

27. The system of claim 26 , wherein the following characteristics of the at least one reference are taken into account: font; font size; style; or any combination thereof.

28. The system of claim 17 , wherein the at least one reference is: merged with at least one other reference; and/or split into at least two references.

29. The system of claim 17 , wherein the at least one dynamic variance network is dynamically adapted during document processing.

30. The system of claim 17 , wherein the at least one dynamic variance network is used for:

reference correction;

document classification;

page separation;

recognition of document modification;

document summarization; or

document compression;

or any combination thereof.

31. The system of claim 17 , wherein other information is also used to determine the target value, the other information comprising:

format information and possible variations of the formal information; and/or

key word information related to the at least one target.

32. The system of claim 17 , wherein the at least one intrinsic content type comprises:

a word;

a number;

a combination of letters and numbers;

a number and punctuation string;

an optical character recognition mistake;

a word in a different language from at least one other word in the at least one document;

a word found in a dictionary;

a word not found in a dictionary;

a font type;

a font size; or

a font property; or

a combination thereof.

Assignments (11)
SECURITY INTEREST Recorded Jan 17, 2024
From: HYLAND SWITZERLAND SARL
To: GOLUB CAPITAL MARKETS LLC, AS COLLATERAL AGENT
Reel/Frame 066339/0304 →
RELEASE OF SECURITY INTEREST RECORDED AT REEL/FRAME 045430/0593 Recorded Sep 24, 2023
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT, A BRANCH OF CREDIT SUISSE
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 065020/0806 →
RELEASE OF SECURITY INTEREST RECORDED AT REEL/FRAME 045430/0405 Recorded Sep 24, 2023
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT, A BRANCH OF CREDIT SUISSE
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 065018/0421 →
CHANGE OF NAME Recorded Feb 20, 2019
From: KOFAX INTERNATIONAL SWITZERLAND SÀRL
To: HYLAND SWITZERLAND SÀRL
Reel/Frame 048389/0380 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT (FIRST LIEN) Recorded Feb 23, 2018
From: KOFAX INTERNATIONAL SWITZERLAND SARL
To: CREDIT SUISSE
Reel/Frame 045430/0405 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT (SECOND LIEN) Recorded Feb 23, 2018
From: KOFAX INTERNATIONAL SWITZERLAND SARL
To: CREDIT SUISSE
Reel/Frame 045430/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2017
From: LEXMARK INTERNATIONAL TECHNOLOGY SARL
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 042919/0841 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2017
From: LEXMARK ENTERPRISE SOFTWARE, LLC
To: LEXMARK INTERNATIONAL TECHNOLOGY SARL
Reel/Frame 042743/0504 →
CHANGE OF NAME Recorded Jun 8, 2017
From: PERCEPTIVE SOFTWARE, LLC
To: LEXMARK ENTERPRISE SOFTWARE, LLC
Reel/Frame 042734/0970 →
MERGER Recorded Jun 6, 2017
From: BRAINWARE INC.
To: PERCEPTIVE SOFTWARE, LLC
Reel/Frame 042617/0217 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2010
From: URBSCHAT, HARRY; MEIER, RALPH; WANSCHURA, THORSTEN; HAUSMANN, JOHANNES
To: BRAINWARE, INC.
Reel/Frame 024452/0391 →