IP Library Granted Patent US 11,416,689
Granted Patent B2
US 11,416,689 · App. 16/367,444 · Granted Aug 16, 2022

System and method for natural language processing with a multinominal topic model

Inventors: Florian Büttner (Munich, DE); Yatin Chaudhary (Munich, DE); Pankaj Gupta (Munich, DE)
Assignee: SIEMENS AKTIENGESELLSCHAFT
G06F40/56G06F40/30G06F40/44G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,689
App. No.
16/367,444
Granted
Aug 16, 2022
Kind
B2
Abstract

The invention refers to a natural language processing system configured for receiving an input sequence c i of input words (v 1 , v 2 , . . . v N ) representing a first sequence of words in a natural language of a first text and generating an output sequence of output words ( , , . . . ) representing a second sequence of words in a natural language of a second text and modeled by a multinominal topic model, wherein the multinominal topic model is extended by an incorporation of language structures using a deep contextualized Long-Short-Term Memory model.

Claims (152)

1. A natural language processing system comprising a processor, the processor configured for:

receiving an input sequence c i of input words (v 1 , v 2 , . . . v N ) representing a first sequence of words in a natural language of a first text, and

generating an output sequence of output words ( , , . . . ) representing a second sequence of words in a natural language of a second text using a multinominal topic model,

wherein the multinominal topic model is extended by an incorporation of language structures using a deep contextualized Long-Short-Term Memory model,

wherein the multinominal topic model is a document neural autoregressive topic model, DocNADE, and the extended multinominal topic model is a contextualized document neural autoregressive topic model, ctx-DocNADE, incorporating context information for the first sequence of words, and

wherein the ctx-DocNADE model is extended by incorporation of distributed compositional priors for generating a ctx-DocNADEe model, incorporating external knowledge for each word of the first sequence of words.

2. The natural language processing system of claim 1 , wherein the distributed composition priors are pre-trained word embeddings by LSTM-LM.

3. The natural language processing system of claim 1 , wherein a conditional probability of the word v i in ctx-DocNADE or ctx-DocNADEe is a function of two hidden vectors: h i DN (v <i ) and h i LM (c i ), stemming from the DocNADE-based and LSTM-based components of ctx-DocNADE, respectively:

h i ( v <i )= h i DN ( v <i )+λ h i LM ( c i )

where h i DN (v <j ) is computed as:

h i DN ( v <i )= g ( e+Σ k<i W :,ν k )

and λ is the mixture weight of the LM component, which can be optimized during training and based on the validation set and the second term h i LM is a context-dependent representation and output of an LSTM layer at position i−1 over input sequence c i , trained to predict the next word v i .

4. The natural language processing system of claim 1 , wherein the conditional distribution for each word v i is estimated by:

p

(

v

i

=

w

|

v

<

i

)

=

exp

(

b

w

+

U

w

,

:

h

i

(

v

<

i

)

)

Σ

w

exp

(

b

w

+

U

w

,

:

h

i

(

v

<

i

)

)

.

5. The natural language processing system of claim 1 , wherein the ctx-DocNADE model and the ctx-DocNADEe model are optimized to maximize the pseudo log likelihood,

log p ( v )≈Σ i=1 D log p (ν i |v <i ).

6. The natural language processing system of claim 1 , wherein the ctx-DocNADEe model and the ctx-DocNADE model are extended to a deep, multiple hidden layer architecture by adding new hidden layers as in a regular deep feed-forward neural network.

7. A computer-implemented method for processing an input sequence c i of input words (v 1 , v 2 , . . . v N ) representing a first sequence of words in a natural language of a first text into an output sequence of output words ( , , . . . ) representing a second sequence of words in a natural language of a second text using a multinominal topic model, comprising the steps:

extending the multinominal topic model by an incorporation of language structures, and

using a deep contextualized Long-Short-Term Memory model,

wherein the multinominal topic model is a document neural autoregressive topic model, DocNADE, and the extended multinominal topic model is a contextualized document neural autoregressive topic model, ctx-DocNADE, incorporating context information for the first sequence of words, and

wherein the ctx-DocNADE model is extended by incorporation of distributed compositional priors for generating a ctx-DocNADEe model, incorporating external knowledge for each word of the first sequence of words.

8. The method of claim 7 , wherein the distributed composition priors are pre-trained word embeddings by LSTM-LM.

9. The method of claim 7 , wherein a conditional probability of the word v i in ctx-DocNADE or ctx-DocNADEe is a function of two hidden vectors: h i DN (v <j ) and h i LM (c i ), stemming from the DocNADE-based and LSTM-based components of ctx-DocNADE, respectively:

h i ( v <i )= h i DN ( v <i )+λ h i LM ( c i )

where h i DN (v <j ) is computed as:

h i DN ( v <i )= g ( e+Σ k<i W :,ν k )

and λ is the mixture weight of the LM component, which can be optimized during training and based on the validation set and the second term h i LM is a context-dependent representation and output of an LSTM layer at position i−1 over input sequence c i , trained to predict the next word v i .

10. The method of claim 7 , wherein the conditional distribution for each word v i is estimated by:

p

(

v

i

=

w

|

v

<

i

)

=

exp

(

b

w

+

U

w

,

:

h

i

(

v

<

i

)

)

Σ

w

exp

(

b

w

+

U

w

,

:

h

i

(

v

<

i

)

)

.

11. The method of claim 7 , wherein the ctx-DocNADE model and the ctx-DocNADEe model are optimized to maximize the pseudo log likelihood,

log p ( v )≈Σ i=1 D log p (ν i |v <i ).

12. A non-transitory computer-readable data storage medium comprising executable program code configured to, when executed, perform the method according to claim 7 .

13. The method of claim 7 , wherein the ctx-DocNADEe model and the ctx-DocNADE model are extended to a deep, multiple hidden layer architecture by adding new hidden layers as in a regular deep feed-forward neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2026
From: SIEMENS AKTIENGESELLSCHAFT
To: DRIMCO GMBH
Reel/Frame 073761/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2019
From: BÜTTNER, FLORIAN; CHAUDHARY, YATIN; GUPTA, PANKAJ
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 049515/0472 →
Continuity (1)
Related Publication 20200311213A1 · Oct 1, 2020