IP Library Granted Patent US 11,429,876
Granted Patent B2
US 11,429,876 · App. 16/814,830 · Granted Aug 30, 2022

Infusing knowledge into natural language processing tasks using graph structures

Inventors: Pavan Kapanipathi Bangalore (Westchester, NY); Kartik Talamadupula (Port Chester, NY); Veronika Thost (Cambridge, MA); Siva Sankalp Patel (White Plains, NY); Ibrahim Abdelaziz (Tarrytown, NY); Avinash Balakrishnan (Elmsford, NY); Maria Chang (Irvington, NY); Kshitij Fadnis (Astoria, NY); Chulaka Gunasekara (New Hyde Park, NY); Bassem Makni (Bellevue, WA); Nicholas Mattei (New Orleans, LA); Achille Belly Fokoue-Nkoutche (White Plains, NY)
Assignee: International Business Machines Corporation
G06N5/02G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,876
App. No.
16/814,830
Granted
Aug 30, 2022
Kind
B2
Abstract

One embodiment of the invention provides a method for natural language processing (NLP). The method comprises extracting knowledge outside of text content of a NLP instance by extracting a set of subgraphs from a knowledge graph associated with the text content. The set of subgraphs comprises the knowledge. The method further comprises encoding the knowledge with the text content into a fixed size graph representation by filtering and encoding the set of subgraphs. The method further comprises applying a text embedding algorithm to the text content to generate a fixed size text representation, and classifying the text content based on the fixed size graph representation and the fixed size text representation.

Claims (34)

1. A method for natural language processing (NLP), comprising:

extracting knowledge outside of text content of a NLP instance by extracting a set of subgraphs from a knowledge graph associated with the text content, wherein the set of subgraphs comprises the knowledge;

encoding the knowledge with the text content into a fixed size graph representation by filtering and encoding the set of subgraphs, wherein the fixed size graph representation includes at least one filtered subgraph comprising only nodes that satisfy a pre-determined threshold;

applying a text embedding algorithm to the text content to generate a fixed size text representation; and

classifying the text content based on the fixed size graph representation and the fixed size text representation.

2. The method of claim 1 , wherein the knowledge graph is one of a knowledge base, a semantic network, or a social graph.

3. The method of claim 1 , wherein the text content comprises one or more text samples.

4. The method of claim 3 , wherein the one or more text samples include a premise and a hypothesis.

5. The method of claim 3 , wherein the knowledge graph comprises one of a directed graph representation or an undirected graph representation of the one or more text samples.

6. The method of claim 1 , wherein the set of subgraphs is encoded via a Relational Graph Convolutional Network (R-GCN).

7. The method of claim 1 , wherein the set of subgraphs is filtered based on a personalized page rank (PPR) algorithm.

8. The method of claim 1 , wherein the text content is classified via a Feed Forward Network (FFN).

9. The method of claim 1 , wherein the text embedding algorithm comprises one of Bidirectional Encoder Representations from Transformers (BERT) or Global Vectors (GloVe).

10. A system for natural language processing (NLP), comprising:

at least one processor; and

a non-transitory processor-readable memory device storing instructions that when executed by the at least one processor causes the at least one processor to perform operations including:

extracting knowledge outside of text content of a NLP instance by extracting a set of subgraphs from a knowledge graph associated with the text content, wherein the set of subgraphs comprises the knowledge;

encoding the knowledge with the text content into a fixed size graph representation by filtering and encoding the set of subgraphs, wherein the fixed size graph representation includes at least one filtered subgraph comprising only nodes that satisfy a pre-determined threshold;

applying a text embedding algorithm to the text content to generate a fixed size text representation; and

classifying the text content based on the fixed size graph representation and the fixed size text representation.

11. The system of claim 10 , wherein the knowledge graph is one of a knowledge base, a semantic network, or a social graph.

12. The system of claim 10 , wherein the text content comprises one or more text samples.

13. The system of claim 12 , wherein the one or more text samples include a premise and a hypothesis.

14. The system of claim 12 , wherein the knowledge graph comprises one of a directed graph representation or an undirected graph representation of the one or more text samples.

15. The system of claim 10 , wherein the set of subgraphs is encoded via a Relational Graph Convolutional Network (R-GCN).

16. The system of claim 10 , wherein the set of subgraphs is filtered based on a personalized pagerank (PPR) algorithm.

17. The system of claim 10 , wherein the text content is classified via a Feed Forward Network (FFN).

18. The system of claim 10 , wherein the text embedding algorithm comprises one of Bidirectional Encoder Representations from Transformers (BERT) or Global Vectors (GloVe).

19. A computer program product for natural language processing (NLP), the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

extracting knowledge outside of text content of a NLP instance by extracting a set of subgraphs from a knowledge graph associated with the text content, wherein the set of subgraphs comprises the knowledge;

encoding the knowledge with the text content into a fixed size graph representation by filtering and encoding the set of subgraphs, wherein the fixed size graph representation includes at least one filtered subgraph comprising only nodes that satisfy a pre-determined threshold;

applying a text embedding algorithm to the text content to generate a fixed size text representation; and

classifying the text content based on the fixed size graph representation and the fixed size text representation.

20. The computer program product of claim 19 , wherein the knowledge graph is one of a knowledge base, a semantic network, or a social graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2020
From: BANGALORE, PAVAN KAPANIPATHI; TALAMADUPULA, KARTIK; THOST, VERONIKA; PATEL, SIVA SANKALP; ABDELAZIZ, IBRAHIM; BALAKRISHNAN, AVINASH; CHANG, MARIA; FADNIS, KSHITIJ; GUNASEKARA, CHULAKA; MAKNI, BASSEM; MATTEI, NICHOLAS; FOKOUE-NKOUTCHE, ACHILLE BELLY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052073/0894 →
Continuity (1)
Related Publication 20210287103A1 · Sep 16, 2021
Cited By (2)
US 12,511,486 US 12,664,191