IP Library Granted Patent US 8,832,108
Granted Patent B1
US 8,832,108 · App. 13/865,350 · Granted Sep 9, 2014

Method and system for classifying documents that have different scales

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,832,108
App. No.
13/865,350
Granted
Sep 9, 2014
Kind
B1
Abstract

Classifying documents that have different scales is described. Instances are counted for each character size in documents. Character sizes for the first document and the second document are selected based on the instance count for each character size. Scales are calculated based on ratios of each first character size relative to each second character size. Scale products are calculated based on each instance count for each character size range for the first character sizes multiplied by each instance count for each corresponding character size range for the second character sizes. The corresponding character size range is based on a corresponding scale. Scale scores are calculated based on summing each of the scale products for each scale. A scale is selected based a highest scale score. The second document may be classified with the first document based on a comparison of first document location information and second document location information. The second document location information is based on the scale.

Claims (58)

1. A system for classifying documents that have different scales, the system comprising:

one or more processors; and

a non-transitory computer readable medium storing a plurality of instructions, which when executed, cause the one or more processors to:

count instances for each character size in a first document and instances for each character size in a second document;

select a first plurality of character sizes for the first document and a second plurality of character sizes for the second document, based on a corresponding count of instances associated with each corresponding character size;

calculate a plurality of scales, wherein each scale of the plurality of scales is based on a corresponding ratio of a corresponding one of the first plurality of character sizes relative to a corresponding one of the second plurality of character sizes;

calculate a plurality of scale products based on each corresponding count of instances for each character size range associated with the first plurality of character sizes multiplied by each corresponding count of instances for each corresponding character size range associated with the second plurality of character sizes, wherein the corresponding character size range is based on a corresponding one of the plurality of scales;

calculate a plurality of scale scores based on summing each of the plurality of scale products associated with each corresponding one of the plurality of scales;

select a scale of the plurality of scales based a highest one of the plurality of scale scores associated with a corresponding one the plurality of scales;

determine whether the second document is in a class associated with the first document based on a comparison of location information associated with the first document and location information associated with the second document, wherein the location information associated with second document is based on the scale; and

classify the second document in the class associated with the first document in response to a determination that the second document is in the class associated with the first document.

2. The system of claim 1 , wherein the first document and the second document comprise digitized optical character recognition data.

3. The system of claim 1 , further comprising instructions, which when executed, cause the one or more processors to classify the second document in a class associated with another document in response to a determination that the second document is not in the class associated with the first document.

4. The system of claim 1 , wherein the comparison of the location information associated with the first document and the location information associated with the second document comprises:

generating a plurality of word pairs, wherein each word pair comprises a word from the first document and a corresponding word from the second document;

computing, for each word pair, first location information for the word that indicates a location of the word in the first document relative to other words in the first document;

computing, for each word pair, second location information for the corresponding word that indicates a location of the corresponding word in the second document relative to other words in the second document; and

comparing the first location information and the second location information.

5. The system of claim 4 , wherein the word pairs comprise keywords associated with the second document based on a comparison of the second document with at least one of a class and a template.

6. The system of claim 1 , wherein the second document location information is also based on at least one of a scale associated with a next highest one of the plurality of scale scores and a scale of one.

7. The system of claim 1 , wherein the first document is associated with a template in response to a comparison to classify documents similar to a document associated with the template.

8. A computer-implemented method for classifying documents that have different scales, the method comprising:

counting, by a server computer, instances for each character size in a first document and instances for each character size in a second document;

selecting, by the server computer, a first plurality of character sizes for the first document and a second plurality of character sizes for the second document, based on a corresponding count of instances associated with each corresponding character size;

calculating, by the server computer, a plurality of scales, wherein each scale of the plurality of scales is based on a corresponding ratio of a corresponding one of the first plurality of character sizes relative to a corresponding one of the second plurality of character sizes;

calculating, by the server computer, a plurality of scale products based on each corresponding count of instances for each character size range associated with the first plurality of character sizes multiplied by each corresponding count of instances for each corresponding character size range associated with the second plurality of character sizes, wherein the corresponding character size range is based on a corresponding one of the plurality of scales;

calculating, by the server computer, a plurality of scale scores based on summing each of the plurality of scale products associated with each corresponding one of the plurality of scales;

selecting, by the server computer, a scale of the plurality of scales based a highest one of the plurality of scale scores associated with a corresponding one the plurality of scales;

determining, by the server computer, whether the second document is in a class associated with the first document based on a comparison of location information associated with the first document and location information associated with the second document, wherein the location information associated with second document is based on the scale; and

classifying, by the server computer, the second document in the class associated with the first document in response to a determination that the second document is in the class associated with the first document.

9. The computer-implemented method of claim 8 , wherein the first document and the second document comprise digitized optical character recognition data.

10. The computer-implemented method of claim 8 , wherein the further comprises classifying the second document in a class associated with another document in response to a determination that the second document is not in the class associated with the first document.

11. The computer-implemented method of claim 8 , wherein the comparison of the location information associated with the first document and the location information associated with the second document comprises:

generating a plurality of word pairs, wherein each word pair comprises a word from the first document and a corresponding word from the second document;

computing, for each word pair, first location information for the word that indicates a location of the word in the first document relative to other words in the first document;

computing, for each word pair, second location information for the corresponding word that indicates a location of the corresponding word in the second document relative to other words in the second document; and

comparing the first location information and the second location information.

12. The computer-implemented method of claim 11 , wherein the word pairs comprise keywords associated with the second document based on a comparison of the second document with at least one of a class and a template.

13. The computer-implemented method of claim 8 , wherein the second document location information is also based on at least one of a scale associated with a next highest one of the plurality of scale scores and a scale of one.

14. The computer-implemented method of claim 8 , wherein the first document is associated with a template in response to a comparison to classify documents similar to a document associated with the template.

15. A computer program product, comprising a non-transitory computer-readable medium having a computer-readable program code embodied therein to be executed by one or more processors, the program code including instructions to:

count instances for each character size in a first document and instances for each character size in a second document;

select a first plurality of character sizes for the first document and a second plurality of character sizes for the second document, based on a corresponding count of instances associated with each corresponding character size;

calculate a plurality of scales, wherein each scale of the plurality of scales is based on a corresponding ratio of a corresponding one of the first plurality of character sizes relative to a corresponding one of the second plurality of character sizes;

calculate a plurality of scale products based on each corresponding count of instances for each character size range associated with the first plurality of character sizes multiplied by each corresponding count of instances for each corresponding character size range associated with the second plurality of character sizes, wherein the corresponding character size range is based on a corresponding one of the plurality of scales;

calculate a plurality of scale scores based on summing each of the plurality of scale products associated with each corresponding one of the plurality of scales;

select a scale of the plurality of scales based a highest one of the plurality of scale scores associated with a corresponding one the plurality of scales;

determine whether the second document is in a class associated with the first document based on a comparison of location information associated with the first document and location information associated with the second document, wherein the location information associated with second document is based on the scale; and

classify the second document in the class associated with the first document in response to a determination that the second document is in the class associated with the first document.

16. The computer program product of claim 15 , wherein the first document and the second document comprise digitized optical character recognition data.

17. The computer program product of claim 15 , wherein the further comprises classifying the second document in a class associated with another document in response to a determination that the second document is not in the class associated with the first document.

18. The computer program product of claim 15 , wherein the comparison of the location information associated with the first document and the location information associated with the second document comprises:

generating a plurality of word pairs, wherein each word pair comprises a word from the first document and a corresponding word from the second document;

computing, for each word pair, first location information for the word that indicates a location of the word in the first document relative to other words in the first document;

computing, for each word pair, second location information for the corresponding word that indicates a location of the corresponding word in the second document relative to other words in the second document; and

comparing the first location information and the second location information.

19. The computer program product of claim 18 , wherein the word pairs comprise keywords associated with the second document based on a comparison of the second document with at least one of a class and a template.

20. The computer program product of claim 15 , wherein the second document location information is also based on at least one of a scale associated with a next highest one of the plurality of scale scores and a scale of one.

Assignments (12)
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 063559/0805) Recorded Jun 21, 2024
From: BARCLAYS BANK PLC
To: OPEN TEXT CORPORATION
Reel/Frame 067807/0069 →
SECURITY INTEREST Recorded Aug 30, 2023
From: OPEN TEXT CORPORATION
To: THE BANK OF NEW YORK MELLON
Reel/Frame 064761/0008 →
SECURITY INTEREST Recorded May 7, 2023
From: OPEN TEXT CORPORATION
To: BARCLAYS BANK PLC
Reel/Frame 063559/0805 →
SECURITY INTEREST Recorded May 7, 2023
From: OPEN TEXT CORPORATION
To: BARCLAYS BANK PLC
Reel/Frame 063559/0831 →
SECURITY INTEREST Recorded May 7, 2023
From: OPEN TEXT CORPORATION
To: BARCLAYS BANK PLC
Reel/Frame 063559/0839 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (045455/0001) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO ASAP SOFTWARE EXPRESS, INC.); DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC CORPORATION (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MAGINATICS LLC); EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); SCALEIO LLC
Reel/Frame 061753/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2017
From: EMC CORPORATION
To: OPEN TEXT CORPORATION
Reel/Frame 041579/0133 →
RELEASE OF SECURITY INTEREST Recorded Jan 23, 2017
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: EMC CORPORATION
Reel/Frame 041073/0443 →
PATENT RELEASE (REEL:40134/FRAME:0001) Recorded Jan 23, 2017
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
To: EMC CORPORATION, AS GRANTOR
Reel/Frame 041073/0136 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 040134/0001 →
SECURITY AGREEMENT Recorded Sep 21, 2016
From: ASAP SOFTWARE EXPRESS, INC.; AVENTAIL LLC; CREDANT TECHNOLOGIES, INC.; DELL USA L.P.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL SOFTWARE INC.; DELL SYSTEMS CORPORATION; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; FORCE10 NETWORKS, INC.; MAGINATICS LLC; MOZY, INC.; SCALEIO LLC; SPANNING CLOUD APPS LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 040136/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2013
From: SAMPSON, STEVEN
To: EMC CORPORATION
Reel/Frame 030285/0005 →