IP Library › Granted Patent US 10,943,070
Granted Patent B2
US 10,943,070 · App. 16/265,177 · Granted Mar 9, 2021

Interactively building a topic model employing semantic similarity in a spoken dialog system

Inventors: Akira Koseki (Yokohama, JP); Masaki Ono (Tokyo, JP); Toshiro Takase (Urayasu, JP); Akihiro Kosugi (Machida, JP)
Assignee: International Business Machines Corporation
G06F40/30G06N3/0454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,943,070
App. No.
16/265,177
Granted
Mar 9, 2021
Kind
B2
Abstract

A computer-implemented method is presented for building a topic model to discover topics in a collection of documents generated by a plurality of users. The method includes extracting conversations from the collection of documents, dividing the extracted conversations into a plurality of segments, generating a topic distribution for each of the plurality of segments based on the extracted conversations and a first pre-defined prior probability distribution, and generating continuous value constructs for each of the topic distributions based on an external corpus and a second pre-defined prior probability distribution, wherein similarity is defined between the continuous value constructs.

Claims (37)

1. A computer implemented method executed on a processor for building a topic model to discover topics in a collection of documents generated by a plurality of users, the method comprising steps of:

extracting conversations from the collection of documents;

dividing the extracted conversations into a plurality of segments;

generating a topic distribution for each of the plurality of segments based on the extracted conversations and a first pre-defined prior probability distribution;

generating continuous value constructs for each of the topic distributions based on an external corpus and a second pre-defined prior probability distribution, wherein similarity is defined between the continuous value constructs; and

generating a Gaussian distribution for each topic distribution of each document of the collection of documents.

2. The method of claim 1 , further comprising observing the continuous value constructs in each of the plurality of segments.

3. The method of claim 2 , further comprising estimating parameters of the topic distributions, construct distributions, and hidden topics for the continuous value constructs by using Gibbs sampling.

4. The method of claim 3 , further comprising selecting an appropriate segment of the plurality of segments based on time.

5. The method of claim 4 , further comprising generating a candidate topic and constructs in the candidate topic by using the second pre-defined prior probability distribution.

6. The method of claim 1 , wherein a generation probability is adjusted based on a distance between the continuous value constructs and a mean value of a corresponding topic distribution.

7. The method of claim 1 , wherein constructs distribution for each topic category is a Gaussian distribution with fixed variance and means calculated by corresponding values of constructs within a same topic category whose prior distribution can be a Gaussian distribution with mean μ 0 and variance σ 0 2 I.

8. A non-transitory computer-readable storage medium comprising a computer-readable program executed on a processor in a data processing system for building a topic model to discover topics in a collection of documents generated by a plurality of users, wherein the computer-readable program when executed on the processor causes a computer to perform the steps of:

extracting conversations from the collection of documents;

dividing the extracted conversations into a plurality of segments;

generating a topic distribution for each of the plurality of segments based on the extracted conversations and a first pre-defined prior probability distribution;

generating continuous value constructs for each of the topic distributions based on an external corpus and a second pre-defined prior probability distribution, wherein similarity is defined between the continuous value constructs; and

generating a Gaussian distribution for each topic distribution of each document of the collection of documents.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the continuous value constructs are observed in each of the plurality of segments.

10. The non-transitory computer-readable storage medium of claim 9 , wherein parameters of the topic distributions, construct distributions, and hidden topics for the continuous value constructs are estimated by using Gibbs sampling.

11. The non-transitory computer-readable storage medium of claim 10 , wherein an appropriate segment of the plurality of segments is selected based on time.

12. The non-transitory computer-readable storage medium of claim 11 , wherein a candidate topic and constructs in the candidate topic are generated by using the second pre-defined prior probability distribution.

13. The non-transitory computer-readable storage medium of claim 8 , wherein a generation probability is adjusted based on a distance between the continuous value constructs and a mean value of a corresponding topic distribution.

14. The non-transitory computer-readable storage medium of claim 8 , wherein constructs distribution for each topic category is a Gaussian distribution with fixed variance and means calculated by corresponding values of constructs within a same topic category whose prior distribution can be a Gaussian distribution with mean μ 0 and variance σ 0 2 I.

15. An system for building a topic model to discover topics in a collection of documents generated by a plurality of users, the system comprising:

a memory; and

one or more processors in communication with the memory configured to:

extract conversations from the collection of documents;

divide the extracted conversations into a plurality of segments;

generate a topic distribution for each of the plurality of segments based on the extracted conversations and a first pre-defined prior probability distribution;

generate continuous value constructs for each of the topic distributions based on an external corpus and a second pre-defined prior probability distribution, wherein similarity is defined between the continuous value constructs; and

generate a Gaussian distribution for each topic distribution of each document of the collection of documents.

16. The system of claim 15 , wherein the continuous value constructs are observed in each of the plurality of segments.

17. The system of claim 16 , wherein parameters of the topic distributions, construct distributions, and hidden topics for the continuous value constructs are estimated by using Gibbs sampling.

18. The system of claim 17 , wherein an appropriate segment of the plurality of segments is selected based on time.

19. The system of claim 18 , wherein a candidate topic and constructs in the candidate topic are generated by using the second pre-defined prior probability distribution.

20. The system of claim 15 , wherein a generation probability is adjusted based on a distance between the continuous value constructs and a mean value of a corresponding topic distribution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2019
From: KOSEKI, AKIRA; ONO, MASAKI; TAKASE, TOSHIRO; KOSUGI, AKIHIRO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048219/0727 →
Continuity (1)
Related Publication 20200250269A1 · Aug 6, 2020
Cited By (1)
US 12,579,380