IP Library › Granted Patent US 11,869,484
Granted Patent B2
US 11,869,484 · App. 17/458,545 · Granted Jan 9, 2024

Apparatus and method for automatic generation and update of knowledge graph from multi-modal sources

Inventors: Yunzhao Lu (Hong Kong, HK); Wai On Sam Lam (Hong Kong, HK); Man Choi Asa Chan (Hong Kong, HK)
Assignee: Hong Kong Applied Science and Technology Research Institute Company Limited
G10L15/063G10L15/02G10L15/1822G10L15/26G10L17/22G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,869,484
App. No.
17/458,545
Granted
Jan 9, 2024
Kind
B2
Abstract

The present invention provides an apparatus and method for automatic generation and update of a knowledge graph from multi-modal sources. The apparatus comprises a conversation parsing module configured for updating a dynamic information word set V D with labelled words generated from extracted from the multi-modal sources; updating a static information word set V S based on extracted schema of relations extracted from the multi-modal sources; and generating pairs of question and answer based on the dynamic information word set V D , the static information word set V S and the one or more sentence patterns; and a knowledge graph container configured for updating a knowledge graph based on the extracted entities of interest and schema of relations. Therefore, an efficient and cost-effective way for question decomposition, query chain construction and entity association from unstructured data is achieved.

Claims (88)

1. An apparatus for automatic generation and update of a knowledge graph from one or more multi-modal sources, the apparatus comprising:

a speaker diarization module configured for: partitioning an input audio stream into audio segments; classifying speakers of the audio segments as agent or customer; and clustering the audio segments based on speaker classification;

an audio transcription module configured for transcribing the clustered audio segments to transcripts based on an acoustic model;

a speech parsing module configured for:

extracting entities of interest and schema of relations from the transcripts; and

labelling words of the transcripts corresponding to the extracted entities of interest with a plurality of pre-defined tags from a domain-specific language model;

a conversation parsing module configured for:

updating a dynamic information word set V D with the labelled words of the transcripts;

updating a static information word set V S based on the extracted schema of relations from the transcripts;

retrieving one or more sentence patterns from the domain-specific language model; and

generating pairs of question and answer based on the dynamic information word set V D , the static information word set V S and the one or more sentence patterns; and

a knowledge graph container configured for updating a knowledge graph by:

receiving the extracted entities of interest and schema of relations;

representing the extracted entities of interest as nodes in the knowledge graph; and

representing the extracted schema of relations as labels and edges between nodes in the knowledge graph.

2. The apparatus of claim 1 , wherein

the speech parsing module is further configured for: extracting entities of interest and schema of relations from an article; and labelling words of the article corresponding to the extracted entities of interest with a plurality of pre-defined tags from a domain-specific language model; and

the conversation parsing module is further configured for: updating the dynamic information word set V D with the labelled words of the article; and updating the static information word set V S based on the extracted schema of relations from the article.

3. The apparatus of claim 1 , wherein the input audio stream is a soundtrack of a video or audio stream.

4. The apparatus of claim 1 , wherein the domain-specific language model is generated by:

generalizing a table of jargons and corpus with vocabulary lexicon to form a general language model; and

interpolating the general language model with pre-defined domain-specific knowledge based on a heuristic weighting to generate the domain-specific language model.

5. The apparatus of claim 1 , wherein

the conversation parsing module is a machine learning module trained with a region-based attention algorithm for extracting the entities of interest across sentences in the transcripts;

the region-based attention algorithm is formulated by defining a region with intra-sentence information and inter-sentence information; and optimizing an objective function based on the defined region.

6. The apparatus of claim 5 , wherein the intra-sentence information is updated through an intra-sentence attention algorithm given by:

R ia =BLSTM t ( X ),

wherein BLSTM t ( ) is a bidirectional long short-term memory function for intra-sentence attention and X is an input word vector representing a set of words in the labelled transcripts; and R ia is an intra-sentence attention output vector.

7. The apparatus of claim 5 , wherein the inter-sentence information is updated through an inter-sentence attention algorithm given by:

V ir =BLSTM l (Σ L Π T α τ γ τ )

where BLSTM l ( ) is a bidirectional long short-term memory function for inter-sentence attention, α τ is a parametric vector, and γ τ is an intra-sentence attention output vector, and V ir is an inter-sentence attention output vector.

8. The apparatus of claim 5 , wherein the objective function is given by:

Ω=softmax(ωβ l +LinB( t λ )),

wherein Ω is the machine learning objective, ωβ i is maximizing expectation argument, and LinB(t λ ) is linear biased estimation of a heuristic weighting parameter t λ .

9. The apparatus of claim 1 , wherein the knowledge graph container is further configured for:

applying entity classification on the dynamic information word set V D and the static information word set V S to generate one or more classified entities;

calculating relation probabilities for a preset number of classified entities with existing entities in the knowledge graph;

identifying a set of best candidates of entity from the classified entities; and

updating the knowledge graph by incorporating a set of best candidates of entity into the knowledge graph.

10. The apparatus of claim 9 , wherein the relation probabilities are given by:

γ l =foo (λ· S+η·K+φ·t λ )

where γ l is the relation probability, S is a classified entity from the dynamic information word set V D and the static information word set V S , K is an existing entity in the knowledge graph, t λ is a heuristic weighting parameter, λ, η and φ are coefficients for S, K and t λ respectively.

11. A method for automatic generation and update of a knowledge graph from multi-modal sources, the method comprising:

clustering, by a speaker diarization module, an input audio stream by:

partitioning the input audio stream into audio segments;

classifying speakers of the audio segments as agent or customer; and

clustering the audio segments based on speaker classification;

transcribing, by an audio transcription module, the clustered audio segments to transcripts based on an acoustic model;

labelling, by a speech parsing module, the transcripts by:

extracting entities of interest and schema of relations from the transcripts; and

labelling words of the transcripts corresponding to the extracted entities of interest with a plurality of pre-defined tags from a domain-specific language model;

generating, by a conversation parsing module, pairs of question and answer by:

updating a dynamic information word set V D with the labelled words of the transcripts and a static information word set V S based on the extracted schema of relations from the transcripts;

retrieving one or more sentence patterns from the domain-specific language model; and

generating the pairs of question and answer based on the dynamic information word set V D , the static information word set V S and the one or more sentence patterns;

updating, by a knowledge graph container, a knowledge graph by:

receiving the extracted entities of interest and schema of relations;

representing, by a knowledge graph container, the extracted entities of interest as nodes in the knowledge graph; and

representing, by a knowledge graph container, the extracted schema of relations as labels and edges between nodes in the knowledge graph.

12. The method of claim 2 , further comprising:

extracting entities of interest and schema of relations from an article;

labelling words of the article corresponding to the extracted entities of interest with a plurality of pre-defined tags from a domain-specific language model;

updating the dynamic information word set V D with the labelled words of the article; and

updating the static information word set V S based on the extracted schema of relations from the article.

13. The method of claim 11 , wherein the input audio stream is a soundtrack of a video or audio stream.

14. The method of claim 11 , wherein the domain-specific language model is generated by:

generalizing a table of jargons and corpus with vocabulary lexicon to form a general language model; and

interpolating the general language model with pre-defined domain-specific knowledge based on a heuristic weighting to generate the domain-specific language model.

15. The method of claim 11 , further comprising:

training the conversation parsing module with a region-based attention algorithm for extracting the entities of interest across sentences in the transcripts;

the region-based attention algorithm is formulated by defining a region with intra-sentence information and inter-sentence information; and optimizing an objective function based on the defined region.

16. The method of claim 15 , wherein the intra-sentence information is updated through an intra-sentence attention algorithm given by:

R ia =BLSTM t ( X ),

wherein BLSTM t ( ) is a bidirectional long short-term memory function for intra-sentence attention and X is an input word vector representing a set of words in the labelled transcripts; and R ia is an intra-sentence attention output vector.

17. The method of claim 15 , wherein the inter-sentence information is updated through an inter-sentence attention algorithm given by:

V ir =BLSTM l (Σ L Π T α τ γ τ ),

where BLSTM i ( ) is a bidirectional long short-term memory function for inter-sentence attention, α τ is a parametric vector, and γ τ is an intra-sentence attention output vector, and V ir is an inter-sentence attention output vector.

18. The method of claim 15 , wherein the objective function is given by:

Ω=softmax(ωβ l +LinB( t λ )),

wherein Ω is the machine learning objective, ωβ l is maximizing expectation argument, and LinB(t λ ) is linear biased estimation of a heuristic weighting parameter t λ .

19. The method of claim 11 , further comprising:

applying entity classification on the dynamic information word set V D and the static information word set V S to generate one or more classified entities;

calculating relation probabilities for a preset number of classified entities with existing entities in the knowledge graph;

identifying a set of best candidates of entity from the classified entities; and

updating the knowledge graph by incorporating set of best candidates of entity into the knowledge graph.

20. The method of claim 19 , wherein the relation probabilities are given by:

γ l =foo (λ· S+η·K+φ·t λ )

where γ l is the relation probability, S is a classified entity from the dynamic information word set V D and the static information word set V S , K is an existing entity in the knowledge graph, t λ is a heuristic weighting parameter, λ, η and φ are coefficients for S, K and t λ respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2021
From: LU, YUNZHAO; LAM, WAI ON SAM; CHAN, MAN CHOI ASA
To: HONG KONG APPLIED SCIENCE AND TECHNOLOGY RESEARCH INSTITUTE COMPANY LIMITED
Reel/Frame 057320/0617 →
Continuity (1)
Related Publication 20230065468A1 · Mar 2, 2023