IP Library › Granted Patent US 12,190,061
Granted Patent B2
US 12,190,061 · App. 17/644,856 · Granted Jan 7, 2025

System and methods for neural topic modeling using topic attention networks

Inventors: Shashank Shailabh (Bihar, IN); Madhur Panwar (Uttar Pradesh, IN); Milan Aggarwal (Delhi, IN); Pinkesh Badjatiya (Madhya Pradesh, IN); Simra Shahid (Uttar Pradesh, IN); Nikaash Puri (New Delhi, IN); S Sejal Naidu (Madhya Pradesh, IN); Sharat Chandra Racha (Telangana, IN); Balaji Krishnamurthy (Uttar Pradesh, IN); Ganesh Karbhari Palwe (Uttar Pradesh, IN)
Assignee: ADOBE INC.
G06F40/289G06F40/30G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,061
App. No.
17/644,856
Granted
Jan 7, 2025
Kind
B2
Abstract

Systems and methods for topic modeling are described. The systems and methods include encoding words of a document using an embedding matrix to obtain word embeddings for the document. The words of the document comprise a subset of words in a vocabulary, and the embedding matrix is trained as part of a topic attention network based on a plurality of topics. The systems and methods further include encoding a topic-word distribution matrix using the embedding matrix to obtain a topic embedding matrix. The topic-word distribution matrix represents relationships between the plurality of topics and the words of the vocabulary. The systems and methods further include computing a topic context matrix based on the topic embedding matrix and the word embeddings and identifying a topic for the document based on the topic context matrix.

Claims (67)

1. A method for topic modeling, comprising:

encoding words of a document using an embedding matrix to obtain word embeddings for the document, wherein the words of the document comprise a subset of words in a vocabulary, and wherein the embedding matrix is trained as part of a topic attention network based on a plurality of topics;

encoding a topic-word distribution matrix using the embedding matrix to obtain a topic embedding matrix, wherein the topic-word distribution matrix represents relationships between the plurality of topics and the words of the vocabulary;

computing a topic context matrix based on the topic embedding matrix and the word embeddings; and

identifying a topic for the document based on the topic context matrix.

2. The method of claim 1 , further comprising:

generating a sequence of hidden representations corresponding to the word embeddings using a sequential encoder, wherein the sequence of hidden representations comprises an order based on an order of the words in the document;

computing an attention alignment matrix based on the topic embedding matrix and the sequence of hidden representations, wherein the topic context matrix is based on the attention alignment matrix;

computing a context vector for the document based on a document-topic vector and the topic context matrix, wherein the topic attention network is trained based on the topic context matrix; and

training the topic attention network based on the context vector.

3. The method of claim 2 , further comprising:

generating a first vector and a second vector based on the context vector;

sampling a random noise term;

computing a sum of the first vector and a product of the second vector and the random noise term to obtain a latent vector;

decoding the latent vector to obtain a predicted word vector; and

computing a loss function based on a comparison between the predicted word vector and the words in the document, wherein the training the topic attention network is further based on the loss function.

4. The method of claim 3 , wherein:

the topic-word distribution matrix is based on a linear layer used to decode a document-word representation from the latent vector.

5. The method of claim 2 , further comprising:

selecting a row of the topic context matrix corresponding to a highest value of the document-topic vector as the context vector.

6. The method of claim 2 , further comprising:

computing an average of rows of the topic context matrix weighted by values of the document-topic vector to obtain the context vector.

7. The method of claim 2 , further comprising:

normalizing the word embeddings based on a number of words in the vocabulary to obtain normalized word embeddings, wherein the document-topic vector is based on the normalized word embeddings.

8. The method of claim 1 , further comprising:

generating a key phrase for the document using words associated with the topic.

9. The method of claim 1 , further comprising:

identifying a set of words for each of the plurality of topics based on the topic-word distribution matrix.

10. The method of claim 1 , further comprising:

selecting a topic for each document in a corpus of documents.

11. The method of claim 1 , further comprising:

selecting the words in the vocabulary based on a corpus of documents, wherein the document is selected from the corpus of documents.

12. A method for topic modeling, comprising:

encoding words of a document using an embedding matrix to obtain word embeddings for the document, wherein the words of the document comprise a subset of words in a vocabulary;

generating a sequence of hidden representations corresponding to the word embeddings using a sequential encoder, wherein the sequence of hidden representations comprises an order based on an order of the words in the document;

computing a context vector for the document based on the sequence of hidden representations;

generating a latent vector based on the context vector using an auto-encoder;

computing a loss function based on the latent vector;

updating parameters of the embedding matrix and a topic attention network based on the loss function; and

predicting, using the embedding matrix and the topic attention network, a set of words including a topic for an input document based on a context vector for the input document.

13. The method of claim 12 , further comprising:

encoding a topic-word distribution matrix using the embedding matrix to obtain a topic embedding matrix, wherein the topic-word distribution matrix represents relationships between the topics and the words of the vocabulary;

computing a document embedding based on the words in the document and the embedding matrix;

computing a document-topic vector based on the topic embedding matrix and the document embedding;

computing an attention alignment matrix based on the topic embedding matrix and the sequence of hidden representations; and

computing a topic context matrix based on the attention alignment matrix and the sequence of hidden representations, wherein the context vector is calculated based on the document-topic vector and the topic context matrix.

14. The method of claim 13 , further comprising:

generating a first vector and a second vector based on the context vector;

sampling a random noise term;

computing a sum of the first vector and a product of the second vector and the random noise term to obtain the latent vector;

decoding the latent vector to obtain a predicted word vector; and

computing the loss function based on a comparison between the predicted word vector and the words in the document.

15. The method of claim 13 , further comprising:

selecting a row of the topic context matrix corresponding to a highest value of the document-topic vector as the context vector.

16. The method of claim 13 , further comprising:

computing an average of rows of the topic context matrix weighted by values of the document-topic vector to obtain the context vector.

17. The method of claim 13 , further comprising:

normalizing the word embeddings based on a number of words in the vocabulary to obtain normalized word embeddings, wherein the document-topic vector is based on the normalized word embeddings.

18. An apparatus for topic modeling, comprising:

an embedding component configured to encode words of a document using an embedding matrix to obtain word embeddings for the document, wherein the words in the document comprise a subset of words in a vocabulary from a corpus of documents, and wherein the embedding matrix is trained based on the corpus of documents;

a sequential encoder configured to generate a sequence of hidden representations corresponding to the word embeddings;

a topic attention network configured to compute a topic context matrix based on the embedding matrix and the sequence of hidden representations and to generate a context vector based on the word embeddings and the topic context matrix; and

an auto-encoder configured to predict a set of words for the document based on the context vector.

19. The apparatus of claim 18 , further comprising:

a topic selection component configured to select a topic for each document in a corpus of documents.

20. The apparatus of claim 18 , further comprising:

a vocabulary component configured to select the words in the vocabulary based on a corpus of documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2021
From: SHAILABH, SHASHANK; PANWAR, MADHUR; AGGARWAL, MILAN; BADJATIYA, PINKESH; SHAHID, SIMRA; PURI, NIKAASH; NAIDU, S SEJAL; RACHA, SHARAT CHANDRA; KRISHNAMURTHY, BALAJI; PALWE, GANESH KARBHARI
To: ADOBE INC.
Reel/Frame 058415/0520 →
Continuity (2)
Provisional Application 63264688 · Nov 30, 2021
Related Publication 20230169271A1 · Jun 1, 2023
References Cited (68)
US 11914670B2 · Kumar · 2024 [cited by examiner]
US 20110264443A1 · Takamatsu · 2011 [cited by examiner]
US 20180285348A1 · Shu · 2018 [cited by examiner]
US 20190236139A1 · DeFelice · 2019 [cited by examiner]
US 20210241050A1 · Gunaratna · 2021 [cited by examiner]
US 20230289533A1 · Gupta · 2023 [cited by examiner]
1Aletras, et al., 2013, “Evaluating Topic Coherence Using Distributional Semantics”, In Proceedings of the 10th International Conference on Computational Semantics (IWCS 2013) Long Papers (pp. 13-22), Available at https… [cited by applicant]
2Andrieu, et al., 2003, “An Introduction to MCMC for Machine Learning”, In Machine Learning, 50 (pp. 5-43), Available at https://www.cs.ubc.ca/˜arnaud/andrieu_defreitas_doucet_jordan_intromontecarlomachinelearning.pdf, … [cited by applicant]
3Bahdanau, et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv preprint arXiv:1409.0473v7 [cs.CL] May 19, 2016, 15 pages. [cited by applicant]
4Bianchi, et al., “Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence”, arXiv preprint arXiv:2004.03974v2 [cs.CL] Jun. 17, 2021, 8 pages. [cited by applicant]
5Bird, et al., 2009, “Natural Language Processing with Python”, O'Reilly Media, Inc., Book Information available at https://www.oreilly.com/library/view/natural-language-processing/9780596803346/. [cited by applicant]
6Blei, et al., 2003, “Latent Dirichlet Allocation”, In Journal of Machine Learning Research 3 (pp. 993-1022), Available at https://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf, 30 pages. [cited by applicant]
7Burkhardt, et al., 2019, “Decoupling Sparsity and Smoothness in the Dirichlet Variational Autoencoder Topic Model”, In Journal of Machine Learning Research 20 (pp. 1-27), Available at https://jmlr.org/papers/volume20/1… [cited by applicant]
8Card, et al., “Neural Models for Documents with Metadata”, arXiv preprint arXiv:1705.09296v2 [stat.ML] Oct. 23, 2018, 13 pages. [cited by applicant]
9Chang, et al., 2009, “Reading Tea Leaves: How Humans Interpret Topic Models”, In Advances in Neural Information Processing Systems, vol. 22 (pp. 288-296), Curran Associates, Inc., Available at https://proceedings.neuri… [cited by applicant]
10Chen, et al., “Keyphrase Generation with Correlation Constraints”, arXiv preprint arXiv:1808.07185v1 [cs.CL] Aug. 22, 2018, 10 pages. [cited by applicant]
11Cho, et al., “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation”, arXiv preprint arXiv:1406.1078v3 [cs.CL] Sep. 3, 2014, 15 pages. [cited by applicant]
12Deerwester, et al., 1990, “Indexing by Latent Semantic Analysis”, Journal of the American Society for Information Science 41 (pp. 391-407), Available at http://lsa.colorado.edu/papers/JASIS.lsi.90.pdf, 34 pages. [cited by applicant]
13Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv preprint arXiv:1810.04805v2 [cs.CL] May 24, 2019, 16 pages. [cited by applicant]
14Dieng, et al., “Topic Modeling in Embedding Spaces”, arXiv preprint arXiv:1907.04907v1 [cs.IR] Jul. 8, 2019, 12 pages. [cited by applicant]
15Dieng, et al., “TopicRNN: A Recurrent Neural Network With Long-Range Semantic Dependency”, arXiv preprint arXiv:1611.01702v2 [cs.CL] Feb. 27, 2017, 13 pages. [cited by applicant]
16Ding, et al., “Coherence-Aware Neural Topic Modeling”, arXiv preprint arXiv:1809.02687v1 [cs.CL] Sep. 7, 2018, 7 pages. [cited by applicant]
17Goodfellow, et al., “Generative Adversarial Nets”, arXiv preprint arXiv:1406.2661v1 [stat.ML] Jun. 10, 2014, 9 pages. [cited by applicant]
18Griffiths, et al., 2004, “Finding scientific topics”, In Proceedings of the National academy of Sciences 101 (pp. 5228-5235), Available at https://www.pnas.org/content/pnas/101/suppl_1/5228.full.pdf, 8 pages. [cited by applicant]
19Gupta, et al., “Document Informed Neural Autoregressive Topic Models with Distributional Prior”, arXiv preprint arXiv:1809.06709v2 [cs.CL] Jan. 14, 2019, 13 pages. [cited by applicant]
20Gururangan, et al., “Variational Pretraining for Semi-supervised Text Classification”, arXiv preprint arXiv:1906.02242v1 [cs.CL] Jun. 5, 2019, 15 pages. [cited by applicant]
21Hochreiter, et al., 1997, “Long Short-Term Memory”, In Neural Computation, 9 (pp. 1735-1780), Massachusetts Institute of Technology, Available at https://direct.mit.edu/neco/article/9/8/1735/6109/Long-Short-Term-Memor… [cited by applicant]
22Hoyle, et al., “Improving Neural Topic Models using Knowledge Distillation”, arXiv preprint arXiv:2010.02377v1 [cs. CL] Oct. 5, 2020, 20 pages. [cited by applicant]
23Hoyle, et al., 2020b, “Improving Neural Topic Models using Knowledge Distillation”, In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1752-1771), Association for Co… [cited by applicant]
24Hu, et al., “Neural Topic Modeling with Cycle-Consistent Adversarial Training”, arXiv preprint arXiv:2009.13971v1 [cs.CL] Sep. 29, 2020, 13 pages. [cited by applicant]
25Ioffe, et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, arXiv preprint arXiv:1502.03167v3 [cs.LG] Mar. 2, 2015, 11 pages. [cited by applicant]
26Jin, et al., 2018, “Combining Deep Learning and Topic Modeling for Review Understanding in Context-Aware Recommendation”, In Proceedings of the 2018 Conference of the North American Chapter of the Association for Comp… [cited by applicant]
27Kingma, et al., “ADAM: A Method for Stochastic Optimization”, arXiv arXiv:1412.6980v9 [cs.LG] Jan. 30, 2017, 15 pages. [cited by applicant]
28Kingma, et al., “Auto-Encoding Variational Bayes”, arXiv preprint arXiv:1312.6114v10 [stat.ML] May 1, 2014, 14 pages. [cited by applicant]
29Kumar, et al., “Why Didn't You Listen to Me? Comparing User Control of Human-in-the-Loop Topic Models”, arXiv preprint arXiv:1905.09864v2 [cs.CL] Jun. 4, 2019, 8 pages. [cited by applicant]
30Lang, 1995, “NewsWeeder: Learning to Filter Netnews”, In Machine Learning Proceedings 95 (pp. 331-339), Elsevier, Available at http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.22.6286, 9 pages. [cited by applicant]
31Lau, et al., 2014, “Machine Reading Tea Leaves: Automatically Evaluating Topic Coherence and Topic Model Quality”, In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Lin… [cited by applicant]
32McCallum, 2002, “MALLET: A machine learning for language toolkit”, http://mallet.cs. umass.edu. [cited by applicant]
33Meng, et al., “Deep Keyphrase Generation”, arXiv preprint arXiv:1704.06879v3 [cs.CL] May 31, 2021, 11 pages. [cited by applicant]
34Miao, et al., “Discovering Discrete Latent Topics with Neural Variational Inference”, arXiv preprint arXiv:1706.00359v2 [cs.CL] May 21, 2018, 11 pages. [cited by applicant]
35Miao, et al., “Neural Variational Inference for Text Processing”, arXiv preprint arXiv:1511.06038v4 [cs.CL] Jun. 4, 2016, 12 pages. [cited by applicant]
36Nan, et al., “Topic Modeling with Wasserstein Autoencoders”, arXiv preprint arXiv:1907.12374v2 [cs.IR] Dec. 6, 2019, 37 pages. [cited by applicant]
37Peinelt, et al., 2020, “tBERT: Topic Models and BERT Joining Forces for Semantic Similarity Detection”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7047-7055), Assoc… [cited by applicant]
38Pennington, at al., 2014, “GloVe: Global Vectors for Word Representation”, In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) (pp. 1532-1543), Available at https://nlp.st… [cited by applicant]
39Pergola, et al., 2019, “TDAM: a Topic-Dependent Attention Model for Sentiment Analysis”, In Journal of Information Processing and Management, Available at https://zenodo.org/record/3632937, 26 pages. [cited by applicant]
40Rehurek, et al., 2010, “Software Framework for Topic Modelling with Large Corpora”, In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks (pp. 45-50) ELRA, Available at https://radimrehurek.com… [cited by applicant]
41Rezaee, et al., “A Discrete Variational Recurrent Topic Model without the Reparametrization Trick”, arXiv preprint arXiv:2010.12055v1 [cs.LG] Oct. 22, 2020, 16 pages. [cited by applicant]
42Sia, et al., “Tired of Topic Models? Clusters of Pretrained Word Embeddings Make for Fast and Good Topics too!”, arXiv preprint arXiv:2004.14914v2 [cs.CL] Oct. 6, 2020, 9 pages. [cited by applicant]
43Srivastava, et al., “Autoencoding Variational Inference for Topic Models”, arXiv preprint arXiv:1703.01488v1 [stat.ML] Mar. 4, 2017, 12 pages. [cited by applicant]
44Srivastava, et al., 2014, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, In Journal of Machine Learning Research 15 (pp. 1929-1958), Available at https://jmlr.org/papers/volume15/srivastava 14a/s… [cited by applicant]
45Steyvers, 2007, “Probabilistic topic models”, In “Handbook of latent semantic analysis” (pp. 424-440), Routledge, Book information available at https://www.routledgehandbooks.com/doi/10.4324/9780203936399.ch21. [cited by applicant]
46Thibaux, et al., 2007, “Hierarchical Beta Processes and the Indian Buffet Process”, In Proceedings of Machine Learning Research, vol. 2 (pp. 564-571), PMLR, Available at http://proceedings.mlr.press/v2/thibaux07a/thib… [cited by applicant]
47Tolstikhin, et al., “Wasserstein Auto-Encoders”, arXiv preprint arXiv:1711.01558v4 [stat.ML] Dec. 5, 2019, 20 pages. [cited by applicant]
48Viegas, et al., 2020, “CluHTM—Semantic Hierarchical Topic Modeling based on CluWords”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 8138-8150), Association for Comput… [cited by applicant]
49Wang, et al., “Neural Topic Modeling with Bidirectional Adversarial Training”, arXiv preprint arXiv:2004.12331v1 [cs.CL] Apr. 26, 2020, 11 pages. [cited by applicant]
50Wang, et al., “ATM: Adversarial-neural Topic Model”, arXvi preprint arXiv:1811.00265v2 [cs.AI] Aug. 21, 2019, 15 pages. [cited by applicant]
51Wang, et al., 2020, “Neural Topic Model with Attention for Supervised Learning”, In Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS), PMR vol. 108, Available at http… [cited by applicant]
62Wang, et al., “Topic-Aware Neural Keyphrase Generation for Social Media Language”, arXiv preprint arXiv:1906.03889v1 [cs.CL] Jun. 10, 2019, 11 pages. [cited by applicant]
53Wu, et al., 2020, “Neural Mixed Counting Models for Dispersed Topic Discovery”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6159-6169), Association for Computational… [cited by applicant]
54Zaheer, et al., 2017, “Latent LSTM Allocation Joint Clustering and Non-Linear Dynamic Modeling of Sequential Data”, In Proceedings of the 34th International Conference on Machine Learning, PMLR vol. 70, Available at h… [cited by applicant]
55Zhang, et al., 2020, “Topic Modeling on Document Networks with Adjacent-Encoder”, In Proceedings of the AAAI Conference on Artificial Intelligence 34 (pp. 6737-6745), Available at https://ojs.aaai.org/index.php/AAAl/a… [cited by applicant]
56Zhang, et al., “WHAI: Weibull Hybrid Autoencoding Inference for Deep Topic Modeling”, arXiv preprint arXiv:1803.01328v2 [stat.ML] Apr. 25, 2020, 15 pages. [cited by applicant]
57Zhang, et al., 2016, “Keyphrase Extraction Using Deep Recurrent Neural Networks on Twitter”, In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 836-845), Association for Com… [cited by applicant]
58Zhang, et al., “Character-level Convolutional Networks for Text Classification”, arXix preprint arXiv:1509.01626v3 [cs.LG] Apr. 4, 2016, 9 pages. [cited by applicant]
59Zhao, et al., 2021, “Neural Topic Model via Optimal Transport”, In International Conference on Learning Representations (pp. 1-15), Available at https://openreview.net/forum?id=Oos98K9Lv-k, 15 pages. [cited by applicant]
60Zhou, et al., “Neural Topic Modeling by Incorporating Document Relationship Graph”, arXiv preprint arXiv:2009.13972v1 [cs.CL] Sep. 29, 2020, 7 pages. [cited by applicant]
61Zhu, et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, arXiv preprint arXiv:1703.10593v7 [cs.CV] Aug. 24, 2020, 18 pages. [cited by applicant]
62Zhu, et al., “A Neural Generative Model for Joint Learning Topics and Topic-Specific Word Embeddings”, arXiv preprint arXiv:2008.04702v1 [cs.CL] Aug. 11, 2020, 14 pages. [cited by applicant]