IP Library Granted Patent US 12,423,345
Granted Patent B2
US 12,423,345 · App. 18/065,803 · Granted Sep 23, 2025

Theme detection within a corpus of information

Inventors: Kasturi Bhattacharjee (Sunnyvale, CA); Rashmi Gangadharaiah (San Jose, CA); Senthil C Chidambaram (Folsom, CA); Ankit Kapoor (Seattle, WA); Sharon Shapira (Sammamish, WA); Tony Chun Tung Ng (San Ramon, CA); Deepak Seetharam Nadig (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06F16/35G06F16/3344G06F16/345
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,345
App. No.
18/065,803
Granted
Sep 23, 2025
Kind
B2
Abstract

Systems and methods are used to detect underlying themes from a collection of documents at an aggregated level. A representative set of documents may be selected from a cluster of documents, with the representative set of documents corresponding to a general theme of the cluster. Candidate theme phrases may then be extracted from the documents and used to generate document embeddings and candidate phrase embeddings, which may be ranked, such as with a diversity-based ranking approach. Certain candidates may be selected from the ranking. Each of the documents forming the representative set may then be concatenated and a query embedding may be generated and ranked against the candidate phrases. In this manner, a collection of phrases associated with both the general underlying theme of the cluster, along with granular topics associated with that theme, may be identified.

Claims (54)

1. A computer-implemented method, comprising:

retrieving a set of documents including unstructured text associated with a cluster, the cluster grouping the set of documents based on one or more central themes;

selecting, from the set of documents, a representative set;

extracting, from portions of the unstructured text in respective documents forming the representative set, one or more candidate theme phrases;

generating, for each document forming the representative set, a document embedding corresponding to individual candidate theme phrases extracted from the document;

generating a candidate phrase embedding corresponding to a set of candidate theme phrases collected from each document of the representative set;

ranking the one or more candidate theme phrases based, at least in part, on the respective document embeddings and the candidate phrase embedding;

selecting, from the ranking, a set of retained candidate theme phrases;

generating, from the representative set, a query embedding corresponding to a concatenated combination of the representative set;

ranking the one or more candidate theme phrases from the query embedding and the set of retained candidate theme phrases;

selecting, based on the ranking, a refined set of retained candidate theme phrases corresponding to a set of semantically similar retained candidate theme phrases associated with the representative set;

determining a retained candidate theme phrase of the refined set of retained candidate theme phrases corresponds to a workflow item; and

changing a priority level for the workflow item.

2. The computer-implemented method of claim 1 , further comprising:

generating a query document corresponding to a concatenated collection of documents in the representative set.

3. The computer-implemented method of claim 1 , wherein at least one of the rankings is based on maximal marginal relevance.

4. The computer-implemented method of claim 1 , further comprising:

mapping the refined set of candidate theme phrases back to respective originating phrases.

5. A method comprising:

determining a representative set of documents from a cluster, based at least in part on one or more metrics associated with a target number of documents;

generating, for each document in the representative set, one or more document embeddings for extracted candidate theme phrases in the form of one or more bi-grams or tri-grams;

generating, for extracted theme phrases from the representative set, one or more embeddings;

identifying, using a diversity-based ranking approach, a set of theme phrases from the one or more document embeddings and the one or more embeddings;

generating, for the representative set of documents, a query embedding including information from each document of the representative set of documents; and

selecting, using the diversity-based ranking approach, a ranked set of theme phrases having a semantic similarity to at least a portion of the cluster and being diverse from other theme phrases in the ranked set of theme phrases, based at least on the query embedding and the set of theme phrases.

6. The method of claim 5 , wherein the query embedding is based on a concatenated collection of each document in the representative set of documents.

7. The method of claim 5 , wherein the diversity-based ranking approach is maximum marginal relevance.

8. The method of claim 5 , wherein the cluster is generated using a centroid-based clustering algorithm.

9. The method of claim 5 , further comprising:

extracting, from the documents forming the representative set, one or more candidate theme phrases; and

removing one or more portions of the candidate theme phrases to generate the theme phrases.

10. The method of claim 9 , wherein the removing includes removing stop words.

11. The method of claim 5 , wherein the extracted theme phrases include non-contiguous tokens.

12. The method of claim 5 , further comprising:

mapping the ranked set of theme phrases back to a respective originating phrase.

13. The method of claim 5 , wherein the theme phrases are extracted using a sentence transformer encoder.

14. The method of claim 5 , further comprising:

generating an action item based on a ranked theme phrase of the ranked set of theme phrases; and

providing a notification to a user including the action item.

15. The method of claim 5 , wherein the representative set includes a number of documents based on a respective distance from a center of the cluster.

16. A system, comprising:

at least one processor; and

memory including instructions that, when executed by the at least one processor, cause the system to:

determine a representative set of documents from a cluster, based at least in part on one or more metrics associated with a target number of documents;

generate, for each document in the representative set, one or more document embeddings for extracted candidate theme phrases in the form of one or more bi-grams or tri-grams;

generate, for extracted theme phrases from the representative set, one or more embeddings;

identify, using a diversity-based ranking approach, a set of theme phrases from the one or more document embeddings and the one or more embeddings;

generate, for the representative set of documents, a query embedding including information from each document of the representative set of documents; and

select, using the diversity-based ranking approach, a ranked set of theme phrases having a semantic similarity to at least a portion of the cluster and being diverse from other theme phrases in the ranked set of theme phrases, based at least on the query embedding and the set of theme phrases.

17. The system of claim 16 , wherein the query embedding is based on a concatenated collection of each document in the representative set of documents.

18. The system of claim 16 , wherein the diversity-based ranking approach is maximum marginal relevance.

19. The system of claim 16 , wherein the cluster is generated using a centroid-based clustering algorithm.

20. The system of claim 16 , wherein the instructions when executed further cause the system to:

map the ranked set of theme phrases back to a respective originating phrase.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2024
From: BHATTACHARJEE, KASTURI; GANGADHARAIAH, RASHMI; CHIDAMBARAM, SENTHIL C; KAPOOR, ANKIT; SHAPIRA, SHARON; NG, TONY CHUN TUNG; NADIG, DEEPAK SEETHARAM
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066814/0865 →
Continuity (2)
Provisional Application 63383577 · Nov 14, 2022
Related Publication 20240160651A1 · May 16, 2024
References Cited (8)
US 12020786B2 · Zhu · 2024 [cited by examiner]
US 20070156732A1 · Surendran · 2007 [cited by examiner]
US 20220261545A1 · Lauber · 2022 [cited by applicant]
US 20220383268A1 · Chen · 2022 [cited by examiner]
US 20230289527A1 · Booth · 2023 [cited by examiner]
US 20240012844A1 · Subraveti · 2024 [cited by examiner]
US 20240104055A1 · McAnallen · 2024 [cited by examiner]
International Search Report and Written Opinion issued Jul. 21, 2023 in PCT Application No. PCT/US2023/067604. [cited by applicant]