IP Library Granted Patent US 11,074,285
Granted Patent B2
US 11,074,285 · App. 15/972,952 · Granted Jul 27, 2021

Recursive agglomerative clustering of time-structured communications

Inventors: Viacheslav Seledkin (Moscow, RU); David Yan (Portola Valley, CA); Marina Chilingaryan (Menlo Park, CA)
Assignee: YVA.AI, INC.
G06F16/358G06F16/3347G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,074,285
App. No.
15/972,952
Granted
Jul 27, 2021
Kind
B2
Abstract

An example method of document clustering comprises: representing each document of a plurality of documents by a vector comprising a first plurality of real values, wherein each real value of the first plurality of real values reflects a first frequency-based metric of a term comprised by the document; partitioning the plurality of documents into a first set of document clusters based on distances between vectors representing the documents; representing each document cluster of the first set of document clusters by a vector comprising a second plurality of real values, wherein each real value of the second plurality of real values reflects a second frequency-based metric of a term comprised by the document cluster; and partitioning the first set of document clusters into a second set of document clusters based on distances between vectors representing the document clusters of the first set of document clusters.

Claims (37)

1. A method of document clustering by a computer system, the method comprising:

representing each document of a plurality of documents by a vector comprising a first plurality of real values, wherein each real value of the first plurality of real values reflects a first frequency-based metric of a first term comprised by the document;

partitioning the plurality of documents into a first set of document clusters based on distances between vectors representing the documents, wherein a distance between a first vector representing a first document of the plurality of documents and a second vector representing a second document of the plurality of documents is provided by a function of a time-sensitive factor and a content-sensitive factor, wherein the time-sensitive factor is determined based on at least one of: a first time identifier associated with the first document and a second time identifier associated with the second document;

representing each document cluster of the first set of document clusters by a vector comprising a second plurality of real values, wherein each real value of the second plurality of real values reflects a second frequency-based metric of a second term comprised by the document cluster; and

partitioning the first set of document clusters into a second set of document clusters based on distances between vectors representing the document clusters of the first set of document clusters.

2. The method of claim 1 , wherein the first term is provided by at least one of: an identifier of a named entity comprised by the document or a time identifier associated with the document.

3. The method of claim 1 , wherein the plurality of documents is provided by an electronic mailbox comprising a plurality of electronic mail messages.

4. The method of claim 1 , wherein the first frequency-based metric is provided by a term frequency—inverse document frequency (TF-IDF) metric.

5. The method of claim 1 , wherein the second frequency-based metric is provided by a function of a ratio of a number of largest document clusters in the first set of document clusters and a number of the largest clusters which include the second term.

6. The method of claim 1 , further comprising:

representing each document cluster of the second set of document clusters by a vector comprising a third plurality of real values, wherein each real value of the third plurality of real values reflects the second frequency-based metric of a third term comprised by the document cluster; and

partitioning the second set of document clusters into a third set of document clusters based on distances between vectors representing the document clusters of the second set of document clusters.

7. The method of claim 1 , further comprising:

associating each cluster of the second set of document clusters with a textual label.

8. The method of claim 1 , further comprising:

visually representing one or more clusters of the second set of document clusters via a graphical user interface.

9. A non-transitory computer-readable storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:

represent each document of a plurality of documents by a vector comprising a first plurality of real values, wherein each real value of the first plurality of real values reflects a first frequency-based metric of a first term comprised by the document;

partition the plurality of documents into a first set of document clusters based on distances between vectors representing the documents;

represent each document cluster of the first set of document clusters by a vector comprising a second plurality of real values, wherein each real value of the second plurality of real values reflects a second frequency-based metric of a second term comprised by the document cluster, wherein the second frequency-based metric is provided by a function of a ratio of a number of largest document clusters in the first set of document clusters and a number of the largest clusters which include the second term; and

partition the first set of document clusters into a second set of document clusters based on distances between vectors representing the document clusters of the first set of document clusters.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the first term is provided by at least one of: an identifier of a named entity comprised by the document or a time identifier associated with the document.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the plurality of documents is provided by an electronic mailbox comprising a plurality of electronic mail messages.

12. The non-transitory computer-readable storage medium of claim 9 , wherein the first frequency-based metric is provided by a term frequency—inverse document frequency (TF-IDF) metric.

13. The non-transitory computer-readable storage medium of claim 9 , wherein a distance between a first vector representing a first document of the plurality of documents and a second vector representing a second document of the plurality of documents is provided by a function of a time-sensitive factor and a content-sensitive factor, wherein the time-sensitive factor is determined based on at least one of: a first time identifier associated with the first document and a second time identifier associated with the second document.

14. The non-transitory computer-readable storage medium of claim 9 , further comprising executable instructions causing the computer system to:

represent each document cluster of the second set of document clusters by a vector comprising a third plurality of real values, wherein each real value of the third plurality of real values reflects the second frequency-based metric of a third term comprised by the document cluster; and

partition the second set of document clusters into a third set of document clusters based on distances between vectors representing the document clusters of the second set of document clusters.

15. A system, comprising:

a memory; and

a processor coupled to the memory, wherein the processor is configured to:

represent each document of a plurality of documents by a vector comprising a first plurality of real values, wherein each real value of the first plurality of real values reflects a first frequency-based metric of a first term comprised by the document;

partition the plurality of documents into a first set of document clusters based on distances between vectors representing the documents, wherein a distance between a first vector representing a first document of the plurality of documents and a second vector representing a second document of the plurality of documents is provided by a function of a time-sensitive factor and a content-sensitive factor, wherein the time-sensitive factor is determined based on at least one of: a first time identifier associated with the first document and a second time identifier associated with the second document;

represent each document cluster of the first set of document clusters by a vector comprising a second plurality of real values, wherein each real value of the second plurality of real values reflects a second frequency-based metric of a second term comprised by the document cluster; and

partition the first set of document clusters into a second set of document clusters based on distances between vectors representing the document clusters of the first set of document clusters.

16. The system of claim 15 , wherein the first term is provided by at least one of: an identifier of a named entity comprised by the document or a time identifier associated with the document.

17. The system of claim 15 , wherein the second frequency-based metric is provided by a function of a ratio of a number of largest document clusters in the first set of document clusters and a number of the largest clusters which include the second term.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2022
From: YVA.AI, INC.
To: VISIER SOLUTIONS INC.
Reel/Frame 059777/0733 →
CHANGE OF NAME Recorded May 3, 2019
From: FINDO INC.
To: YVA.AI, INC.
Reel/Frame 049086/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2019
From: SELEDKIN, VIACHESLAV; YAN, DAVID; CHILINGARYAN, MARINA
To: FINDO, INC.
Reel/Frame 047987/0114 →
Continuity (2)
Provisional Application 62504390 · May 10, 2017
Related Publication 20180329989A1 · Nov 15, 2018