IP Library Granted Patent US 10,089,382
Granted Patent B2
US 10,089,382 · App. 14/887,096 · Granted Oct 2, 2018

Transforming a knowledge base into a machine readable format for an automated system

Inventors: Akhil Arora (New Delhi, IN); Manoj Gupta (Bangalore, IN); Shourya Roy (Bangalore, IN)
Assignee: Conduent Business Services, LLC
G06F17/30598G06F17/30011G06F17/3053G06F17/30327
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,089,382
App. No.
14/887,096
Granted
Oct 2, 2018
Kind
B2
Abstract

A method, non-transitory computer readable medium and apparatus for transforming a knowledge base into a machine readable format for an automated system are disclosed. For example, the method includes clustering two or more documents of a plurality of documents in the knowledge base that are similar based upon a domain specific shingling function, identifying a pattern from each document of the plurality of documents that are clustered, wherein the pattern comprises a sequence of steps, aggregating the pattern of the each document of the plurality of documents that are clustered into a trie data-structure that is machine readable, receiving a request and applying the trie data-structure to provide a solution to the request.

Claims (53)

1. A method for transforming a knowledge base into a machine readable format, comprising:

clustering, by a processor of a knowledge base processing server, two or more documents of a plurality of documents in the knowledge base that are similar based upon a domain specific shingling function;

identifying, by the processor, a pattern from each document of the plurality of documents that are clustered, wherein the pattern comprises a sequence of steps;

aggregating, by the processor, the pattern of the each document of the plurality of documents that are clustered into a trie data-structure that is machine readable;

receiving, by the processor, a request; and

applying, by the processor, the trie data-structure to provide a solution to the request.

2. The method of claim 1 , further comprising:

pre-processing, by the processor, the plurality of documents to extract metadata.

3. The method of claim 2 , wherein the domain specific shingling function comprises identifying a plurality of different domains in each one of the plurality of documents based on the metadata.

4. The method of claim 3 , wherein the clustering comprises:

calculating, by the processor, a similarity score for each one of the plurality of different domains of a document to a corresponding domain of each remaining document of the plurality of documents;

calculating, by the processor, an overall similarity score based on the similarity score for the each one of the plurality of different domains for the document to the each remaining document of the plurality of documents; and

clustering, by the processor, the document and one or more of the each remaining document of the plurality of documents when the overall similarity score is above a threshold.

5. The method of claim 1 , wherein the identifying the pattern, comprises:

building, by the processor, a suffix tree for each one of the plurality of documents that are clustered;

extracting, by the processor, each pattern of a plurality of patterns from the suffix tree for the each one of the plurality of documents that are clustered; and

identifying, by the processor, the pattern from the plurality of patterns that occurs above a predefined number of documents of the plurality of documents that are clustered.

6. The method of claim 5 , wherein the extracting each pattern of the plurality of patterns comprises using a suffix link when traversing the suffix tree.

7. The method of claim 1 , wherein the aggregating comprises:

combining, by the processor, a plurality of patterns into a single pattern represented by the trie data-structure.

8. The method of claim 1 , wherein the clustering, the identifying and the aggregating are repeated periodically to account for new documents added to the knowledge base.

9. The method of claim 1 , wherein the applying the trie data-structure to provide the solution to the request comprises navigating the trie data-structure node by node until the solution is provided.

10. An apparatus, comprising:

a processor; and

a computer readable medium in communication with the processor, the computer readable medium comprising:

a clustering module to cluster two or more documents of a plurality of documents in a knowledge base that are similar based upon a domain specific shingling function;

a pattern identification module to identify a pattern from each document of the plurality of documents that are clustered, wherein the pattern comprises a sequence of steps;

an aggregation module to aggregate the pattern of the each document of the plurality of documents that are clustered into a trie data-structure that is machine readable; and

a customer interaction module to receive a request and apply the trie data-structure to provide a solution to the request.

11. The apparatus of claim 10 , wherein the processor is in communication with the knowledge base to obtain the plurality of documents.

12. The apparatus of claim 10 , wherein the domain specific shingling function comprises identifying a plurality of different domains in each one of the plurality of documents based on metadata extracted from the each one of the plurality of documents.

13. The apparatus of claim 12 , wherein the clustering module is further configured to:

calculate a similarity score for each one of the plurality of different domains of a document to a corresponding domain of each remaining document of the plurality of documents;

calculate an overall similarity score based on the similarity score for the each one of the plurality of different domains for the document to the each remaining document of the plurality of documents; and

cluster the document and one or more of the each remaining document of the plurality of documents when the overall similarity score is above a threshold.

14. The apparatus of claim 10 , wherein the pattern identification module is further configured to:

build a suffix tree for each one of the plurality of documents that are clustered;

extract each pattern of a plurality of patterns from the suffix tree for the each one of the plurality of documents that are clustered; and

identify the pattern from the plurality of patterns that occurs above a predefined number of documents of the plurality of documents that are clustered.

15. The apparatus of claim 14 , wherein the pattern identification module extracts each pattern of the plurality of patterns using a suffix link when traversing the suffix tree.

16. The apparatus of claim 10 , wherein the aggregation module is further configured to:

combine a plurality of patterns into a single pattern represented by the trie data-structure.

17. The apparatus of claim 10 , the cluster module, the pattern identification module and the aggregation module may be periodically activated to account for new documents added to the knowledge base.

18. The apparatus of claim 10 , wherein the customer interaction module applies the trie data-structure to provide the solution to the request by navigating the trie data-structure node by node until the solution is provided.

19. A method for transforming a knowledge base into a machine readable format, comprising:

extracting, by a processor of a knowledge base processing server, HTML tags for each document of a plurality of documents in the knowledge base, wherein the HTML tags define a plurality of domains for each one of the plurality of documents;

calculating, by the processor, an overall similarity score for each pair of documents of the plurality of documents based on a similarity score of each one of the plurality of domains of the each pair of documents;

clustering, by the processor, two or more documents of the plurality of documents into a cluster of documents that have the overall similarity score above a threshold value;

identifying, by the processor, a plurality of patterns comprising a pattern from each document of the cluster of documents using suffix links of a suffix tree associated with the each document of the cluster of documents, wherein the pattern comprises a sequence of steps;

aggregating, by the processor, the plurality of patterns into a trie data-structure that is machine readable;

receiving, by the processor, a customer care request; and

applying, by the processor, the trie data-structure to provide a solution to the customer care request.

20. The method of claim 19 , wherein the extracting, the calculating, the clustering, the identifying and the aggregating are repeated periodically to account for new documents added to the knowledge base.

Assignments (6)
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: U.S. BANK, NATIONAL ASSOCIATION
Reel/Frame 057969/0445 →
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 057970/0001 →
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: CONDUENT BUSINESS SERVICES, LLC; CONDUENT STATE & LOCAL SOLUTIONS, INC.; CONDUENT TRANSPORT SOLUTIONS, INC.; ADVECTIS, INC.; CONDUENT COMMERCIAL SOLUTIONS, LLC; CONDUENT BUSINESS SOLUTIONS, LLC; CONDUENT CASUALTY CLAIMS SOLUTIONS, LLC; CONDUENT HEALTH ASSESSMENTS, LLC
Reel/Frame 057969/0180 →
SECURITY AGREEMENT Recorded Apr 23, 2019
From: CONDUENT BUSINESS SERVICES, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 050326/0511 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2017
From: XEROX CORPORATION
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 041542/0022 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2015
From: ARORA, AKHIL; GUPTA, MANOJ; ROY, SHOURYA
To: XEROX CORPORATION
Reel/Frame 036827/0513 →
Continuity (1)
Related Publication 20170109426A1 · Apr 20, 2017