IP Library Granted Patent US 10,963,501
Granted Patent B1
US 10,963,501 · App. 15/582,625 · Granted Mar 30, 2021

Systems and methods for generating a topic tree for digital information

Inventors: Naveen Ramachandrappa (San Jose, CA); Ramya Mula (San Jose, CA); Ashwin Kayyoor (Sunnyvale, CA); Bashyam Tca (Saratoga, CA)
Assignee: Veritas Technologies LLC
G06F16/3347G06F7/08G06F16/322
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,963,501
App. No.
15/582,625
Granted
Mar 30, 2021
Kind
B1
Abstract

The disclosed computer-implemented method for generating a topic tree for digital information may include parsing the digital information and extracting a set of keywords. This method may also include comparing the set of keywords to an ontology and extracting hierarchies from the ontology that match the set of keywords. The extracted ontology entries may then be pruned and sorted. Various other methods, systems, and computer-readable media are also disclosed.

Claims (44)

1. A computer-implemented method for generating a topic tree for digital information to classify said digital information into a hierarchy of topics and subtopics, the method comprising:

parsing the digital information to locate at least one sentence;

determining at least one noun in the digital information;

generating at least one word that is similar to the at least one noun;

removing at least one proper noun from the digital information, wherein the proper noun includes at least one of: a person's name, a location's name, or an organization's name;

extracting a set of keywords from the digital information, wherein the set of keywords comprises the at least one noun and the at least one word that is similar to the at least one noun;

comparing the set of keywords to an ontology;

extracting at least one hierarchy from the ontology that matches the set of keywords by comparing the set of keywords to a plurality of hierarchies in the ontology, selecting all hierarchies that include at least one keyword from the set of keywords, and removing all hierarchies from the selected hierarchies that do not include a threshold number of keywords;

generating word embeddings that encode information about a correlation with other keywords using cosine similarity between vectors to find correlated keywords between the set of keywords;

merging all the hierarchies extracted from the ontology; and

sorting the merged hierarchies extracted from the ontology based on relevance scores from the word embeddings for the topic tree for digital information.

2. The method according to claim 1 , further comprising mapping the digital information to weighted vectors, wherein the sorting of the at least one extracted hierarchy is based on the weighted vectors.

3. The method according to claim 2 , further comprising generating a set of similar words based on a cosine similarity between the weighted vectors.

4. The method according to claim 1 , wherein extracting the set of keywords comprises:

locating a plurality of sentences within the digital information; and

applying part-of-speech tagging to the located plurality of sentences.

5. A system for generating a topic tree for digital information to classify said digital information into a hierarchy of topics and subtopics, the system comprising:

a parsing module, stored in memory, that parses the digital information to locate at least one sentence, determines at least one noun in the digital information, generates at least one word that is similar to the at least one noun, removes at least one proper noun from the digital information, wherein the proper noun includes at least one of: a person's name, a location's name, or an organization's name, and extracts a set of keywords, wherein the set of keywords comprises the at least one noun and the at least one word that is similar to the at least one noun;

a comparison module, stored in memory, that compares the set of keywords to an ontology; an extraction module, stored in memory, that extracts a plurality of hierarchies from the ontology that match the set of keywords by comparing the set of keywords to a plurality of hierarchies in the ontology, selects all hierarchies that include at least one keyword that match from the set of keywords, removes all hierarchies from the selected hierarchies that do not include a threshold number of keywords, generates word embeddings that encode information about a correlation with other keywords using cosine similarity between vectors to find correlated keywords between the set of keywords, and merges all the hierarchies extracted from the ontology;

a sorting module, stored in memory, that sorts the merged hierarchies extracted from the ontology based on relevance scores from the word embeddings for the topic tree for digital information; and

at least one processor that executes the parsing module, the comparison module, the extraction module, and the sorting module.

6. The system according to claim 5 , further comprising a mapping module, stored in memory, that maps the digital information to weighted vectors, wherein the sorting module sorts the extracted hierarchies based on the weighted vectors.

7. The system according to claim 6 , wherein the mapping module generates a set of similar words based on a cosine similarity between the weighted vectors.

8. The system according to claim 5 , wherein the parsing module extracting the set of keywords:

locates a plurality of sentences within the digital information; and

applies part-of-speech tagging to the located plurality of sentences.

9. A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device generate a topic tree for digital information to classify said digital information into a hierarchy of topics and subtopics by causing the computing device to:

parse a digital document to locate at least one sentence;

determine at least one noun in the digital document;

generate at least one word that is similar to the at least one noun;

remove at least one proper noun from the digital document, wherein the proper noun includes at least one of: a person's name, a location's name, or an organization's name;

extract a set of keywords from the digital document, wherein the set of keywords comprises the at least one noun and the at least one word that is similar to the at least one noun;

compare the set of keywords to an ontology;

extract at least one hierarchy from the ontology that matches at least one keyword of the set of keywords by comparing the set of keywords to a plurality of hierarchies in the ontology, selecting all hierarchies that include at least one keyword from the set of keywords, and removing all hierarchies from the selected hierarchies that do not include a threshold number of keywords;

generate word embeddings that encode information about a correlation with other keywords using cosine similarity between vectors to find correlated keywords between the set of keywords;

merge all the hierarchies extracted from the ontology; and

sort the merged hierarchies from the ontology based on relevance scores from the word embeddings for the topic tree for digital information.

10. The non-transitory computer-readable medium according to claim 9 , wherein:

the one or more computer-executable instructions cause the computing device to map the digital information to weighted vectors; and

the sorting of the at least one extracted hierarchy is based on the weighted vectors.

11. The non-transitory computer-readable medium according to claim 10 , wherein the one or more computer-executable instructions cause the computing device to generate a set of words that are similar to the keywords based on a cosine similarity between the weighted vectors.

12. The non-transitory computer-readable medium according to claim 9 , wherein extracting the set of keywords comprises:

locating a plurality of sentences within the digital information; and

applying a part-of-speech tagging to the located plurality of sentences.

Assignments (14)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY DATA AND CORRECT THE PATENT NUMBERS PREVIOUSLY RECORDED AT REEL: 69548 FRAME: 468. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 4, 2026
From: VERITAS TECHNOLOGIES LLC
To: ARCTERA US LLC
Reel/Frame 074876/0584 →
SECURITY INTEREST Recorded Dec 12, 2025
From: ARCTERA US LLC
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073951/0470 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 070530/0497 Recorded Dec 1, 2025
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0730 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT AT R/F 069585/0150 Recorded Dec 1, 2025
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: ARCTERA US LLC
Reel/Frame 073833/0848 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069634/0584 →
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 069574/0931 →
SECURITY INTEREST Recorded Dec 10, 2024
From: ARCTERA US LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 069563/0243 →
PATENT SECURITY AGREEMENT Recorded Dec 10, 2024
From: ARCTERA US LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069585/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2024
From: VERITAS TECHNOLOGIES LLC
To: ARCTERA US LLC
Reel/Frame 069548/0468 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 052426/0001 Recorded Nov 30, 2020
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 054535/0565 →
SECURITY INTEREST Recorded Aug 20, 2020
From: VERITAS TECHNOLOGIES LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 054370/0134 →
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Apr 16, 2020
From: VERITAS TECHNOLOGIES, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 052426/0001 →
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Jul 10, 2017
From: VERITAS TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 043141/0403 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2017
From: RAMACHANDRAPPA, NAVEEN; MULA, RAMYA; KAYYOOR, ASHWIN; TCA, BASHYAM
To: VERITAS TECHNOLOGIES LLC
Reel/Frame 042187/0629 →
Cited By (3)
US 12,204,860 US 12,222,974 US 12,242,803