IP Library Granted Patent US 12,190,622
Granted Patent B2
US 12,190,622 · App. 16/951,485 · Granted Jan 7, 2025

Document clusterization

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,622
App. No.
16/951,485
Granted
Jan 7, 2025
Kind
B2
Abstract

A computer-implemented method for document clusterization, comprising: receiving an input document; determining, by evaluating a document similarity function, a plurality of similarity measures, wherein each similarity measure of the plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of a plurality of clusters of documents; based on the plurality of similarity measures, determining that the input document does not belong to any of the clusters of documents of the plurality of clusters of documents; creating a new cluster of documents; and associating the input document with the new cluster of documents.

Claims (43)

1. A computer-implemented method for document clusterization, comprising:

receiving an input document;

determining, by evaluating a first document similarity function, a first plurality of similarity measures, wherein each similarity measure of the first plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of a plurality of clusters of documents, wherein a first likelihood of the first document similarity function to yield a false negative result exceeds a second likelihood of the first document similarity function to yield a false positive result;

based on the plurality of similarity measures, determining that the input document belongs to a subset comprising two or more adjacent clusters of the plurality of clusters of documents, wherein a distance between centroids of the two or more adjacent clusters is less than a predefined separation distance;

determining, by evaluating a second document similarity function that is different from the first document similarity function, a second plurality of similarity measures, wherein each similarity measure of the second plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of the subset of the plurality of clusters of documents, and wherein the first document similarity function is based on a first number of attributes of the input document, the second document similarity function is based on a second number of attributes of the input document, the second number exceeding the first number;

associating the input document with a cluster of documents associated with a maximum similarity measure of the second plurality of similarity measures.

2. The method of claim 1 , wherein the first document similarity function is based on one or more attributes of the input document, the one or more attributes comprising at least one of: a grid type attribute, a singular value decomposition (SVD) type attribute, or an image type attribute.

3. The method of claim 1 , wherein the first document similarity function is implemented by a neural network.

4. The method of claim 1 , wherein the input document is a text document.

5. The method of claim 1 , further comprising: responsive to determining that a first cluster of documents of the plurality of clusters of documents is associated with a first document having a first value of a document feature and a second cluster of documents of the plurality of clusters of documents is associated with a second document having the first value of the document feature, merging the first cluster of documents and the second cluster of documents.

6. The method of claim 1 , further comprising: responsive to determining that the maximum similarity measure falls below a similarity measure threshold, creating a new cluster of documents; and associating the input document with the new cluster of documents.

7. The method of claim 1 , wherein the first similarity function is based on a set of attributes, each attribute of the set of attributes computed for a corresponding cell of a grid defined on the input document.

8. The method of claim 1 , wherein each similarity measure of the second plurality of similarity measures reflects a degree of similarity between the input document and one or more randomly selected documents of a corresponding cluster of documents of the subset of the plurality of clusters of documents.

9. A system, comprising:

a memory;

a processor, coupled to the memory, the processor configured to:

receive an input document;

determine, by evaluating a first document similarity function, a first plurality of similarity measures, wherein each similarity measure of the first plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of a plurality of clusters of documents, wherein a first likelihood of the first document similarity function to yield a false negative result exceeds a second likelihood of the first document similarity function to yield a false positive result;

based on the plurality of similarity measures, determine that the input document belongs to a subset comprising two or more adjacent clusters of the plurality of clusters of documents, wherein a distance between centroids of the two or more adjacent clusters is less than a predefined separation distance;

determine, by evaluating a second document similarity function that is different from the first document similarity function, a second plurality of similarity measures, wherein each similarity measure of the second plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of the subset of the plurality of clusters of documents, and wherein the first document similarity function is based on a first number of attributes of the input document, the second document similarity function is based on a second number of attributes of the input document, the second number exceeding the first number; and

associate the input document with a cluster of documents associated with a maximum similarity measure of the second plurality of similarity measures.

10. The system of claim 9 , wherein the first document similarity function is based on one or more attributes of the input document, the one or more attributes comprising at least one of: a grid type attribute, a singular value decomposition (SVD) type attribute, or an image type attribute.

11. The system of claim 9 , wherein the first document similarity function is implemented by a neural network.

12. The system of claim 9 , wherein the input document is a text document.

13. The system of claim 9 , wherein the processor is further configured to:

responsive to determining that a first cluster of documents of the plurality of clusters of documents is associated with a first document having a first value of a document feature and a second cluster of documents of the plurality of clusters of documents is associated with a second document having the first value of the document feature, merge the first cluster of documents and the second cluster of documents.

14. The system of claim 9 , wherein the processor is further configured to:

responsive to determining that the maximum similarity measure falls below a similarity measure threshold, create a new cluster of documents; and

associate the input document with the new cluster of documents.

15. A non-transitory computer-readable storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:

receive an input document;

determine, by evaluating a first document similarity function, a first plurality of similarity measures, wherein each similarity measure of the first plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of a plurality of clusters of documents, wherein a first likelihood of the first document similarity function to yield a false negative result exceeds a second likelihood of the first document similarity function to yield a false positive result;

based on the plurality of similarity measures, determine that the input document belongs to a subset comprising two or more adjacent clusters of the plurality of clusters of documents, wherein a distance between centroids of the two or more adjacent clusters is less than a predefined separation distance;

determine, by evaluating a second document similarity function that is different from the first document similarity function, a second plurality of similarity measures, wherein each similarity measure of the second plurality of similarity measures reflects a degree of similarity between the input document and a corresponding cluster of documents of the subset of the plurality of clusters of documents, and wherein the first document similarity function is based on a first number of attributes of the input document, the second document similarity function is based on a second number of attributes of the input document, the second number exceeding the first number; and

associate the input document with a cluster of documents associated with a maximum similarity measure of the second plurality of similarity measures.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the first document similarity function is based on one or more attributes of the input document, the one or more attributes comprising at least one of: a grid type attribute, a singular value decomposition (SVD) type attribute, or an image type attribute.

17. The non-transitory computer-readable storage medium of claim 15 , wherein the first document similarity function is implemented by a neural network.

18. The non-transitory computer-readable storage medium of claim 15 , wherein the input document is a text document.

19. The non-transitory computer-readable storage medium of claim 15 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:

responsive to determining that a first cluster of documents of the plurality of clusters of documents is associated with a first document having a first value of a document feature and a second cluster of documents of the plurality of clusters of documents is associated with a second document having the first value of the document feature, merge the first cluster of documents and the second cluster of documents.

20. The non-transitory computer-readable storage medium of claim 15 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:

responsive to determining that the maximum similarity measure falls below a similarity measure threshold, create a new cluster of documents; and

associate the input document with the new cluster of documents.

Assignments (3)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2020
From: SEMENOV, STANISLAV; ANTONOVA, ALEXANDRA; MISYUREV, ALEKSEY
To: ABBYY PRODUCTION LLC
Reel/Frame 054410/0088 →