IP Library Granted Patent US 11,194,968
Granted Patent B2
US 11,194,968 · App. 15/993,641 · Granted Dec 7, 2021

Automatized text analysis

Inventors: Florian Büttner (Munich, DE); Pankaj Gupta (Munich, DE)
Assignee: SIEMENS AKTIENGESELLSCHAFT
G06F40/30G06N3/0445G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,194,968
App. No.
15/993,641
Granted
Dec 7, 2021
Kind
B2
Abstract

The present invention concerns a text analysis system, the text analysis system being adapted for utilizing a topic model to provide a document representation. The topic model is based on learning performed on a text corpus utilizing hidden layer representations associated to words of the text corpus, wherein each hidden layer representation pertains to a specific word of the text corpus and is based on a word environment including words occurring before and after the specific word in a text of the text corpus.

Claims (217)

1. A text analysis system, comprising:

a processor;

a memory device coupled to the processor; and

a computer readable storage device coupled to the processor, wherein the storage device contains program code executable by the processor via the memory device to implement a method comprising:

utilizing a topic model to provide a document representation, the topic model being based on learning performed on a text corpus utilizing hidden layer representations associated to words of the text corpus, wherein each hidden layer representation pertains to a specific word of the text corpus and is based on a word environment comprising words occurring before and after the specific word in a text of the text corpus, for determining a context of the specific word;

wherein the hidden layer representation is calculated according to a formula, wherein the formula is represented by:

h i ( v n ,v m )= g ( Dc+Σ k<i W :,v k +Σ k>i W :,v k );

wherein the hidden layer representation is a single unified context vector that simultaneously accumulates the context and the uses the context to learn text representation in a single step;

wherein the topic model is a neural autoregressive topic model and an autoregressive conditional is expressed as:

p

(

v

i

=

w

v

n

,

v

m

)

=

exp

(

b

w

+

U

w

,

:

h

i

(

v

n

,

v

m

)

)

w

exp

(

b

w

+

U

w

,

:

h

i

(

v

n

,

v

m

)

)

.

2. The text analysis system according to claim 1 , wherein the document representation represents a word probability distribution.

3. The text analysis system according to claim 1 , the system being adapted to determine a topic of an input text based on the document representation.

4. The text analysis system according to claim 1 , wherein the system utilises a Recurring Neural Network, RNN, for learning hidden layer representations.

5. A method for performing text analysis, the method comprising utilizing a topic model to provide a document representation, the topic model being based on learning performed on a text corpus utilizing hidden layer representations associated to words of the text corpus, wherein each hidden layer representation pertains to a specific word of the text corpus and is based on a word environment comprising words occurring before and after the specific word in a text of the text corpus, for determining a context of the specific word;

wherein the hidden layer representation is calculated according to a formula, wherein the formula is represented by:

h i ( v n ,v m )= g ( Dc+Σ k<i W :,v k +Σ k>i W :,v k );

wherein the hidden layer representation is a single unified context vector that simultaneously accumulates the context and the uses the context to learn text representation in a single step;

wherein the topic model is a neural autoregressive topic model and an autoregressive conditional is expressed as:

p

(

v

i

=

w

v

n

,

v

m

)

=

exp

(

b

w

+

U

w

,

:

h

i

(

v

n

,

v

m

)

)

w

exp

(

b

w

+

U

w

,

:

h

i

(

v

n

,

v

m

)

)

.

6. The method according to claim 5 , wherein the document representation represents a word probability distribution.

7. The method according to claim 5 , the method comprising determining a topic of an input text based on the document representation.

8. The method according to claim 5 , the method comprising utilizing a Recurring Neural Network, RNN, for learning hidden layer representations.

9. A non-transitory computer program comprising instructions causing a computer system to perform and/or control a method according to claim 5 .

10. A non-transitory storage medium storing a computer program according to claim 9 .

11. A text analysis system, comprising:

a processor;

a memory device coupled to the processor; and

a computer readable storage device coupled to the processor, wherein the storage device contains program code executable by the processor via the memory device to implement a method comprising:

utilizing a topic model to provide a document representation, the topic model being based on learning performed on a text corpus utilizing hidden layer representations associated to words of the text corpus, wherein each hidden layer representation pertains to a specific word of the text corpus and is based on a word environment comprising words occurring before and after the specific word in a text of the text corpus, for determining a context of the specific word;

wherein the hidden layer representation is calculated according to a formula, wherein the formula is represented by:

h i ( v )= g ( Dc+Σ k<i W :,V k +Σ k>i W :,v k +W :,v i )

wherein the hidden layer representation is a single unified context vector that simultaneously accumulates the context and the uses the context to learn text representation in a single step;

wherein the topic model is a neural autoregressive topic model and an autoregressive conditional is expressed as:

p

(

v

i

=

w

v

)

=

exp

(

b

w

+

U

w

,

:

h

i

(

v

)

)

w

exp

(

b

w

+

U

w

,

:

h

i

(

v

)

)

.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2026
From: SIEMENS AKTIENGESELLSCHAFT
To: DRIMCO GMBH
Reel/Frame 073761/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2018
From: BÜTTNER, FLORIAN; GUPTA, PANKAJ
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 046216/0646 →
Continuity (1)
Related Publication 20190370331A1 · Dec 5, 2019