IP Library Granted Patent US 8,321,357
Granted Patent B2
US 8,321,357 · App. 12/570,412 · Granted Nov 27, 2012

Method and system for extraction

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,321,357
App. No.
12/570,412
Granted
Nov 27, 2012
Kind
B2
Abstract

A system and method for extracting information from at least one document in at least one set of documents, the method comprising: generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score.

Claims (104)

1. A method for extracting information from at least one document in at least one set of documents, the method comprising:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score;

determining, using at least one features processor, negative features and positive features based on N-gram statistics;

determining, using at least one negative features processor, whether negative features apply to the at least one possible match;

deleting, using at least one deleting processor, any possible match to which the negative feature applies from the at least one possible match list;

determining, using at least one positive features processor, whether any of the possible matches are positive features; and

re-ordering, using at least one re-ordering processor, the possible matches in the at least one possible match list based on the information learned from determining whether any of the possible matches are positive features.

2. The method of claim 1 , wherein the at least one attribute score and the at least one localization score are based on:

spatial feature criteria;

contextual feature criteria;

relational feature criteria; or

derived feature criteria; or

any combination thereof.

3. The method of claim 2 , wherein the spatial feature criteria is used to determine areas where the at least one target entry is most likely to be found.

4. The method of claim 2 , wherein the contextual feature criteria weighs information about at least one possible target entry in the neighborhood of the at least one target entry.

5. The method of claim 2 , wherein the relational feature criteria is used to determine at least one area where and within which the at least one target entry is likely to be found.

6. The method of claim 2 , wherein the derived feature criteria is generated by mathematical transformations between any combination of the spatial feature criteria, the contextual feature criteria, and the relational feature criteria.

7. The method of claim 1 , wherein at least one processor can comprise: the at least one ranking and/or matching processor, the at least one features processor, the at least one negative features processor, the at least one deleting processor, the at least one positive features processor, or the at least one re-ordering processor, or any combination thereof.

8. The method of claim 1 , further comprising:

learning characteristics of the at least one set of documents from sample documents;

using the learned characteristics to find similar information in the at least one set of documents.

9. The method of claim 8 , wherein the learned characteristics apply to at least one unknown document and/or at least one different document type.

10. The method of claim 1 , further comprising:

validating information in the at least one document to determine if the information is consistent.

11. The method of claim 10 , wherein the validating comprises internal validating and/or external validating.

12. The method of claim 1 , wherein the ranked possible match list based on the at least one attribute score and the at least one localization score takes into account information related to:

text features;

geometric features;

graphic features;

feature conversion; or

any combination thereof.

13. A method for extracting information from at least one document in at least one set of documents, the method comprising:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score.

14. The method of claim 13 , wherein the at least one attribute score and the at least one localization score are based on:

spatial feature criteria;

contextual feature criteria;

relational feature criteria; or

derived feature criteria; or

any combination thereof.

15. The method of claim 14 , wherein the spatial feature criteria is used to determine areas where the at least one target entry is most likely to be found.

16. The method of claim 14 , wherein the contextual feature criteria weighs information about at least one possible target entry in the neighborhood of the at least one target entry.

17. The method of claim 14 , wherein the relational feature criteria is used to determine at least one area where and within which the at least one target entry is likely to be found.

18. The method of claim 14 , wherein the derived feature criteria is generated by mathematical transformations between any combination of the spatial feature criteria, the contextual feature criteria, and the relational feature criteria.

19. The method of claim 13 , wherein at least one processor can comprise: the at least one ranking and/or matching processor, the at least one features processor, the at least one negative features processor, the at least one deleting processor, the at least one positive features processor, or the at least one re-ordering processor, or any combination thereof.

20. The method of claim 13 , further comprising:

learning characteristics of the at least one set of documents from sample documents;

using the learned characteristics to find similar information in the at least one set of documents.

21. The method of claim 20 , wherein the learned characteristics apply to at least one unknown document and/or at least one different document type.

22. The method of claim 13 , further comprising:

validating information in the at least one document to determine if the information is consistent.

23. The method of claim 22 , wherein the validating comprises internal validating and/or external validating.

24. The method of claim 13 , wherein the ranked possible match list based on the at least one attribute score and the at least one localization score takes into account information related to:

text features;

geometric features;

graphic features;

feature conversion; or

any combination thereof.

25. A method for extracting information from at least one document in at least one set of documents, the method comprising:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score;

determining, using at least one features processor, positive features based on N-gram statistics; and

re-ordering, using at least one re-ordering processor, the possible matches in the at least one possible match list based on the information learned from determining whether any of the possible matches are positive features.

26. The method of claim 25 , wherein the at least one attribute score and the at least one localization score are based on:

spatial feature criteria;

contextual feature criteria;

relational feature criteria; or

derived feature criteria; or

any combination thereof.

27. The method of claim 26 , wherein the spatial feature criteria is used to determine areas where the at least one target entry is most likely to be found.

28. The method of claim 26 , wherein the contextual feature criteria weighs information about at least one possible target entry in the neighborhood of the at least one target entry.

29. The method of claim 26 wherein the relational feature criteria is used to determine at least one area where and within which the at least one target entry is likely to be found.

30. The method of claim 26 , wherein the derived feature criteria is generated by mathematical transformations between any combination of the spatial feature criteria, the contextual feature criteria, and the relational feature criteria.

31. The method of claim 25 , wherein at least one processor can comprise: the at least one ranking and/or matching processor, the at least one features processor, or the at least one re-ordering processor, or any combination thereof.

32. The method of claim 25 further comprising:

learning characteristics of the at least one set of documents from sample documents;

using the learned characteristics to find similar information in the at least one set of documents.

33. The method of claim 32 , wherein the learned characteristics apply to at least one unknown document and/or at least one different document type.

34. The method of claim 25 , further comprising:

validating information in the at least one document to determine if the information is consistent.

35. The method of claim 34 , wherein the validating comprises internal validating and/or external validating.

36. The method of claim 25 , wherein the ranked possible match list based on the at least one attribute score and the at least one localization score takes into account information related to:

text features;

geometric features;

graphic features;

feature conversion; or

any combination thereof.

37. A computer system for extracting information from at least one document in at least one set of documents, the system comprising:

at least one processor;

wherein the at least one processor is configured to perform:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score;

determining, using at least one features processor, negative features and positive features based on N-gram statistics;

determining, using at least one negative features processor, whether negative features apply to the at least one possible match;

deleting, using at least one deleting processor, any possible match to which the negative feature applies from the at least one possible match list;

determining, using at least one positive features processor, whether any of the possible matches are positive features; and

re-ordering, using at least one re-ordering processor, the possible matches in the at least one possible match list based on the information learned from determining whether any of the possible matches are positive features.

38. A computerized system for extracting information from at least one document in at least one set of documents, the system comprising:

at least one processor;

wherein the processor is configured to perform:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score.

39. A computerized system for extracting information from at least one document in at least one set of documents, the system comprising:

at least one processor;

wherein the processor is configured to perform:

generating, using at least one ranking and/or matching processor, at least one ranked possible match list comprising at least one possible match for at least one target entry on the at least one document, the at least one ranked possible match list based on at least one attribute score and at least one localization score;

determining, using at least one features processor, positive features based on N-gram statistics; and

re-ordering, using at least one re-ordering processor, the possible matches in the at least one possible match list based on the information learned from determining whether any of the possible matches are positive features.

Assignments (11)
SECURITY INTEREST Recorded Jan 17, 2024
From: HYLAND SWITZERLAND SARL
To: GOLUB CAPITAL MARKETS LLC, AS COLLATERAL AGENT
Reel/Frame 066339/0304 →
RELEASE OF SECURITY INTEREST RECORDED AT REEL/FRAME 045430/0593 Recorded Sep 24, 2023
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT, A BRANCH OF CREDIT SUISSE
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 065020/0806 →
RELEASE OF SECURITY INTEREST RECORDED AT REEL/FRAME 045430/0405 Recorded Sep 24, 2023
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT, A BRANCH OF CREDIT SUISSE
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 065018/0421 →
CHANGE OF NAME Recorded Feb 20, 2019
From: KOFAX INTERNATIONAL SWITZERLAND SÀRL
To: HYLAND SWITZERLAND SÀRL
Reel/Frame 048389/0380 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT (FIRST LIEN) Recorded Feb 23, 2018
From: KOFAX INTERNATIONAL SWITZERLAND SARL
To: CREDIT SUISSE
Reel/Frame 045430/0405 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT (SECOND LIEN) Recorded Feb 23, 2018
From: KOFAX INTERNATIONAL SWITZERLAND SARL
To: CREDIT SUISSE
Reel/Frame 045430/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2017
From: LEXMARK INTERNATIONAL TECHNOLOGY SARL
To: KOFAX INTERNATIONAL SWITZERLAND SARL
Reel/Frame 042919/0841 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2017
From: LEXMARK ENTERPRISE SOFTWARE, LLC
To: LEXMARK INTERNATIONAL TECHNOLOGY SARL
Reel/Frame 042743/0504 →
CHANGE OF NAME Recorded Jun 8, 2017
From: PERCEPTIVE SOFTWARE, LLC
To: LEXMARK ENTERPRISE SOFTWARE, LLC
Reel/Frame 042734/0970 →
MERGER Recorded Jun 6, 2017
From: BRAINWARE INC.
To: PERCEPTIVE SOFTWARE, LLC
Reel/Frame 042617/0217 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2010
From: LAPIR, GENNADY; URBSCHAT, HARRY; MEIER, RALPH; WANSCHURA, THORSTEN; HAUSMANN, JOHANNES
To: BRAINWARE, INC.
Reel/Frame 024409/0904 →