IP Library › Granted Patent US 12,423,599
Granted Patent B2
US 12,423,599 · App. 17/212,163 · Granted Sep 23, 2025

Efficient and accurate regional explanation technique for NLP models

Inventors: Zahra Zohrevand (Vancouver, CA); Tayler Hetherington (Vancouver, CA); Karoon Rashedi Nia (Vancouver, CA); Yasha Pushak (Vancouver, CA); Sanjay Jinturkar (Santa Clara, CA); Nipun Agarwal (Saratoga, CA)
Assignee: Oracle International Corporation
G06N5/04G06F40/20G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,599
App. No.
17/212,163
Granted
Sep 23, 2025
Kind
B2
Abstract

Herein are techniques for topic modeling and content perturbation that provide machine learning (ML) explainability (MLX) for natural language processing (NLP). A computer hosts an ML model that infers an original inference for each of many text documents that contain many distinct terms. To each text document (TD) is assigned, based on terms in the TD, a topic that contains a subset of the distinct terms. In a perturbed copy of each TD, a perturbed subset of the distinct terms is replaced. For the perturbed copy of each TD, the ML model infers a perturbed inference. For TDs of a topic, the computer detects that a difference between original inferences of the TDs of the topic and perturbed inferences of the TDs of the topic exceeds a threshold. Based on terms in the TDs of the topic, the topic is replaced with multiple, finer-grained new topics. After sufficient topic modeling, a regional explanation of the ML model is generated.

Claims (74)

1. A method comprising:

inferring, by a machine learning (ML) model, for each text document in a plurality of text documents that contain a plurality of distinct terms, a respective original inference;

assigning, to each text document in the plurality of text documents, based on terms in the text document, a respective topic of a plurality of topics, wherein the topic contains a respective subset of the plurality of distinct terms;

replacing, in a respective perturbed copy of each text document in the plurality of text documents, a perturbed subset of the plurality of distinct terms without replacing, in said respective perturbed copy of each text document in the plurality of text documents, said subset of the plurality of distinct terms of the topic of the text document;

inferring, by the ML model, for the perturbed copy of each text document in the plurality of text documents, a respective perturbed inference;

detecting, for text documents of a particular topic of the plurality of topics, that a difference between original inferences of the text documents of the particular topic and perturbed inferences of the text documents of the particular topic exceeds a threshold;

generating in response to said detecting, based on terms in the text documents of the particular topic, a plurality of new topics that said plurality of topics does not contain;

discarding the particular topic; and

generating, after said discarding, and displaying a regional explanation for the original inference for a text document of the plurality of text documents, wherein the regional explanation indicates a topic of the plurality of new topics;

wherein the method is performed by one or more computers.

2. The method of claim 1 further comprising reassigning, to each text document of the particular topic, based on terms in the text document, a respective topic of the plurality of new topics.

3. The method of claim 1 further comprising detecting, for text documents of a particular new topic of the plurality of new topics, that an aggregate coherence exceeds a new threshold.

4. The method of claim 1 further comprising selecting the plurality of topics as a smallest plurality of topics that has an aggregate coherence that exceeds an aggregate threshold.

5. The method of claim 1 wherein:

said original inference is an original class of a plurality of classes;

said perturbed inference is a perturbed class of the plurality of classes;

the plurality of topics contains a first topic and a second topic;

the subset of the plurality of distinct terms of the first topic is larger than the subset of the plurality of distinct terms of the second topic.

6. The method of claim 5 wherein before said discarding the particular topic, the plurality of topics is smaller than the plurality of classes.

7. The method of claim 5 further comprising selecting the plurality of topics based on

said plurality of classes.

8. The method of claim 7 wherein said selecting the plurality of topics comprises combining a subset of the plurality of classes into a combined class.

9. The method of claim 5 further comprising:

ranking the plurality of distinct terms based on

said plurality of classes;

selecting the plurality of topics based on said ranking the plurality of distinct terms.

10. The method of claim 1 wherein at least one selected from a group consisting of:

the plurality of topics is not predefined and

a size of the plurality of topics is not predefined.

11. The method of claim 1 wherein:

the method further comprises ranking the plurality of distinct terms;

said replacing in said perturbed copy comprises selecting replacement terms based on said ranking the plurality of distinct terms.

12. The method of claim 1 wherein said replacing in said perturbed copy comprises selecting a same replacement term regardless of the term being replaced.

13. The method of claim 1 wherein said detecting that the difference exceeds the threshold is based on an arithmetic average.

14. The method of claim 1 further comprising generating an explanation that contains at least one selected from a group consisting of:

a respective weight for each term of a subset of the plurality of distinct terms,

multiple weights per topic for at least one topic of the plurality of topics,

a ranking of at least a subset of the plurality of topics,

distances between at least a subset of the plurality of topics, and

visual overlap of two topics of the plurality of topics when said subsets of the plurality of distinct terms of the two topics contain a same term.

15. The method of claim 1 further comprising unsupervised training the ML model to infer, for a text document, a class of a plurality of classes.

16. The method of claim 15 wherein:

said plurality of text documents consists of unclassified text documents;

said original inference is a plurality of original probabilities;

said perturbed inference is a plurality of perturbed probabilities;

said assigning the respective topic to each text document of the plurality of text documents is not based on said plurality of classes.

17. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

inferring, by a machine learning (ML) model, for each text document in a plurality of text documents that contain a plurality of distinct terms, a respective original inference;

assigning, to each text document in the plurality of text documents, based on terms in the text document, a respective topic of a plurality of topics, wherein the topic contains a respective subset of the plurality of distinct terms;

replacing, in a respective perturbed copy of each text document in the plurality of text documents, a perturbed subset of the plurality of distinct terms without replacing, in said respective perturbed copy of each text document in the plurality of text documents, said subset of the plurality of distinct terms of the topic of the text document;

inferring, by the ML model, for the perturbed copy of each text document in the plurality of text documents, a respective perturbed inference;

detecting, for text documents of a particular topic of the plurality of topics, that a difference between original inferences of the text documents of the particular topic and perturbed inferences of the text documents of the particular topic exceeds a threshold;

generating in response to said detecting, based on terms in the text documents of the particular topic, a plurality of new topics that said plurality of topics does not contain;

discarding the particular topic; and

generating, after said discarding, and displaying a regional explanation for the original inference from a text document of the plurality of text documents, wherein the regional explanation indicates a topic of the plurality of new topics.

18. The one or more non-transitory computer-readable media of claim 17 wherein:

said original inference is an original class of a plurality of classes;

said perturbed inference is a perturbed class of the plurality of classes;

the plurality of topics contains a first topic and a second topic;

the subset of the plurality of distinct terms of the first topic is larger than the subset of the plurality of distinct terms of the second topic.

19. The one or more non-transitory computer-readable media of claim 18 wherein before said discarding the particular topic, the plurality of topics is smaller than the plurality of classes.

20. The one or more non-transitory computer-readable media of claim 18 wherein the instructions further cause selecting the plurality of topics based on said plurality of classes.

21. The one or more non-transitory computer-readable media of claim 18 wherein the instructions further cause:

ranking the plurality of distinct terms based on at least one selected from a group ranking the plurality of distinct terms based on

said plurality of classes;

selecting the plurality of topics based on said ranking the plurality of distinct terms.

22. The one or more non-transitory computer-readable media of claim 17 wherein at least one selected from a group consisting of: the plurality of topics is not predefined and a size of the plurality of topics is not predefined.

23. The one or more non-transitory computer-readable media of claim 17 wherein said replacing in said perturbed copy comprises selecting a same replacement term regardless of the term being replaced.

24. The one or more non-transitory computer-readable media of claim 17 wherein the instructions further cause generating an explanation that contains at least one selected from a group consisting of:

a respective weight for each term of a subset of the plurality of distinct terms,

multiple weights per topic for at least one topic of the plurality of topics,

a ranking of at least a subset of the plurality of topics,

distances between at least a subset of the plurality of topics, and

visual overlap of two topics of the plurality of topics when said subsets of the plurality of distinct terms of the two topics contain a same term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2021
From: ZOHREVAND, ZAHRA; HETHERINGTON, TAYLER; NIA, KAROON RASHEDI; PUSHAK, YASHA; JINTURKAR, SANJAY; AGARWAL, NIPUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 055716/0705 →
Continuity (1)
Related Publication 20220309360A1 · Sep 29, 2022
References Cited (113)
US 6789069B1 · Barnhill · 2004 [cited by applicant]
US 8489603B1 · Weissgerber · 2013 [cited by applicant]
US 10705796B1 · Doyle · 2020 [cited by applicant]
US 20060080311A1 · Potok · 2006 [cited by applicant]
US 20080027926A1 · Diao · 2008 [cited by applicant]
US 20090274376A1 · Selvaraj · 2009 [cited by applicant]
US 20100145678A1 · Csomai · 2010 [cited by examiner]
US 20120089621A1 · Liu · 2012 [cited by applicant]
US 20130158982A1 · Zechner · 2013 [cited by applicant]
US 20160180200A1 · Vijayanarasimhan · 2016 [cited by applicant]
US 20160187022A1 · Miwa · 2016 [cited by applicant]
US 20170091692A1 · Guo · 2017 [cited by applicant]
US 20170344822A1 · Popescu · 2017 [cited by applicant]
US 20180018320A1 · Boyer · 2018 [cited by examiner]
US 20180175975A1 · Um · 2018 [cited by applicant]
US 20190130212A1 · Cheng · 2019 [cited by applicant]
US 20190197357A1 · Anderson et al. · 2019 [cited by applicant]
US 20200110982A1 · Gou · 2020 [cited by applicant]
US 20200226212A1 · Tan · 2020 [cited by examiner]
US 20200302318A1 · Hetherington · 2020 [cited by applicant]
US 20200387675A1 · Nugent · 2020 [cited by examiner]
US 20210049503A1 · Nourian · 2021 [cited by examiner]
US 20220229983A1 · Zohrevand et al. · 2022 [cited by applicant]
US 20230259707A1 · Joshi · 2023 [cited by examiner]
WO WO2018006004A1 · 2018 [cited by applicant]
Ushioda, A. (Jun. 1996). Hierarchical clustering of words and application to NLP tasks. In Fourth Workshop on Very Large Corpora. (Year: 1996). [cited by examiner]
Lawrie, D., Croft, W. B., & Rosenberg, A. (Sep. 2001). Finding topic words for hierarchical summarization. In Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information … [cited by examiner]
Menon, A., Narasimhan, H., Agarwal, S., & Chawla, S. (May 2013). On the statistical consistency of algorithms for binary classification under class imbalance. In International Conference on Machine Learning (pp. 603-611… [cited by examiner]
Röder, M., Both, A., & Hinneburg, A. (Feb. 2015). Exploring the space of topic coherence measures. In Proceedings of the eighth ACM international conference on Web search and data mining (pp. 399-408). (Year: 2015). [cited by examiner]
Alvarez-Melis, D., & Jaakkola, T. S. (Nov. 2017). A causal framework for explaining the predictions of black-box sequence-to-sequence models. arXiv preprint arXiv:1707.01943. (Year: 2017). [cited by examiner]
Korenčić, D., Ristov, S., & Šnajder, J. (Jul.2018). Document-based topic coherence measures for news media text. Expert systems with Applications, 114, 357-373. (Year: 2018). [cited by examiner]
Greco, S. (2019). Explaining black-box models in the context of Natural Language Processing (Doctoral dissertation, Politecnico di Torino). (Year: 2019). [cited by examiner]
Ren, S., Deng, Y., He, K., & Che, W. (Jul. 2019). Generating natural language adversarial examples through probability weighted word saliency. In Proceedings of the 57th annual meeting of the association for computation… [cited by examiner]
Carbonero-Ruz, M., Martínez-Estudillo, F. J., Fernández-Navarro, F., Becerra-Alonso, D., & Martínez-Estudillo, A. C. (Dec. 2016). A two dimensional accuracy-based measure for classification performance. Information Scie… [cited by examiner]
Amer, A. A., & Abdalla, H. I. (Sep. 2020). A set theory based similarity measure for text clustering and classification. Journal of Big Data, 7(1), 74. (Year: 2020). [cited by examiner]
Koh et al., Understanding Black-box Predictions via Influence Functions, From the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, dated Dec. 29, 2020, 12 pages. [cited by applicant]
Jin et al., “Data discretization unification”, Regular Paper, Springer-Verlag London Limited 2008, 29 pages. [cited by applicant]
Holte, Robert, “Very Simple Classification Rules Perform Well on Most Commonly Used Datasets”, 1993 Kluwer Academic Publishers, Boston. Manufactured in The Netherlands, 28 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 12 pages. [cited by applicant]
Hand et al., “Data Mining.” Wiley StatsRef: Statistics Reference Online (2014): 1-7. [cited by applicant]
Das et al., “Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey”, dated Jun. 23, 2020, 24 pages. [cited by applicant]
Breiman, Leo. “Random forests.” Machine learning 45.1, dated Jan. 2001, 33 pages. [cited by applicant]
Zintgraf et al., “Visualizing Deep Neural Network Decisions: Prediction difference analysis” dated Feb. 15, 2017, 12 pages. [cited by applicant]
“Decision Trees”, dated Jan. 26, 2005, 10 pages. [cited by applicant]
Agrawal et al., “Fast Discovery of Association Rules”, dated 1996, 22 pages. [cited by applicant]
Akoglu, Haldun, “User's Guide to Correlation Coefficients”, Turkish Journal of Emergency Medicine 18, dated 2018, pp. 91-93. [cited by applicant]
Baehrens et al., “How to Explain Individual Classification Decisions”, Journal of Machine Learning Research 11 dated 2010, 29 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets”, dated 2014, 9 pages. [cited by applicant]
Bloniarz et al., “Supervised Neighborhoods for Distributed Nonparametric Regression”, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics dated 2016, 10 pages. [cited by applicant]
Goodfellow et al., “Explaining and Harnessing Adversarial Examples”, Published as a conference paper at ICLR dated 2015, 11 pages. [cited by applicant]
Dhurandhar et al., “Explanations based on the Missing: Towards Contrastive Explanations with Pertinent Negatives”, dated 2018, 12 pages. [cited by applicant]
Dougherty et al., “Supervised and Unsupervised Discretization of Continous Features”, dated 1995, 9 pages. [cited by applicant]
Dua, D. and Graff, C., “Uci Machine Learning Repository” [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science, dated 2019, 2 pages. [cited by applicant]
Egan et al., “Generalized Latent Variable Recovery for Generative Adversarial Networks”, dated Oct. 19, 2018, 9 pages. [cited by applicant]
Fayyad, “Multi-Interval Discretization of Continuous-Valued Attributed for Classification Learning”, dated 1993, 6 pages. [cited by applicant]
Lundberg et al., “A Unified Approach to Interpreting Model Predictions”, 31st Conference on Neural Information Processing Systems (NIPS 2017), dated 2017, Long Beach, CA, USA, 10 pages. [cited by applicant]
Bakarov, Amir. “A Survey of Word Embeddings Evaluation Methods.” dated Jan. 21, 2018, 26 pages. [cited by applicant]
Vanschoren et al., “OpenML: Networked Science in Machine Learning”, dated Aug. 1, 2014, 12 pages. [cited by applicant]
Lipton et al., “Precise Recovery of Latent Vectors From Generative Adversarial Networks”, Workshop track—ICLR dated Feb. 17, 2017, 4 pages. [cited by applicant]
Tenney et al., “What Do You Learn From Context?”, Published as a conference paper at ICLR 2019, dated May 15, 2019, 17 pages. [cited by applicant]
Turney et al., “From Frequency to Meaning: Vector Space Models of Semantics.”, Journal of artificial intelligence research 37, dated Mar. 4, 2010, 48 pages. [cited by applicant]
UCI Machine Learning Repository, “SMS Spam Collection Data Set”, https://archive.ics.uci.edu/ml/datasets/SMS+Spam+Collection, dated Jun. 22, 2012, 2 pages. [cited by applicant]
UCI Machine Learning Repository, “Twenty Newsgroups Data Set”, https://archive.ics.uci.edu/ml/datasets/Twenty+Newsgroups, dated Sep. 9, 1999, 2 pages. [cited by applicant]
Strumbelj et al., “Explaining Predictions Models and Individual Predictions with Contributions”, Knowledge and Information Systems 41.3, dated Dec. 2014, pp. 647-665. [cited by applicant]
Van Looveren et al., “Interpretable Counterfactual Explanations Guided by Prototypes”, dated Feb. 18, 2020, 17 pages. [cited by applicant]
Strumbelj et al., “An Efficient Explanation of Individual Classifications Using Game Theory”, Journal of Machine Learning Research 11, dated 2010, 18 pages. [cited by applicant]
Wachter et al., “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”, dated 2017, 52 pages. [cited by applicant]
Xu et al., “Modeling Tabular Data using Conditional GAN”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), dated 2019, Vancouver, Canada, 11 pages. [cited by applicant]
Yang et al., “HEIDL: Learning Linguistic Expressions with Deep Learning and Human-in-the-Loop”, dated Jul. 25, 2019, 6 pages. [cited by applicant]
Yang, H., and John Moody. “Feature Selection Based on Joint Mutual Information.” Proceedings of international ICSC symposium on advances in intelligent data analysis, dated 1999, 8 pages. [cited by applicant]
Yu et al. “Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control”. dated Dec. 15, 2019, 13 pages. [cited by applicant]
Zeiler, M. D., & Fergus, R., “Visualizing and Understanding Convolutional Networks”, In European conference on computer vision, Springer, Cham., dated Nov. 28, 2013, 11 pages. [cited by applicant]
Utgoff, Paul, “Incremental Induction of Decision Trees”, 1989 Kluwer Academic Publishers, Boston. Manufactured in The Netherlands, 26 pages. [cited by applicant]
Ribeiro et al., “Why Should I Trust You?” Explaining the Predictions of Any Classifier, KDD dated 2016 San Francisco, CA, USA, 10 pages. [cited by applicant]
Zhao et al., “Data-driven risk-averse stochastic optimization with Wasserstein metric”, Operations Research Letters, dated 2018, 6 pages. [cited by applicant]
Pedregosa et al., “Scikit-learn: Machine Learning in Python”, Journal of Machine Learning Research 12, dated 2011, 6 pages. [cited by applicant]
Plumb et al., “Model Agnostic Supervised Local Explanations”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), dated 2018, Montréal, Canada, 10 pages. [cited by applicant]
Quinlan et al., “Induction of Decision Trees”, dated 1986 Kluwer Academic Publishers, Boston—Manufactured in The Netherlands, 26 pages. [cited by applicant]
Ramírez-Gallego et al., “Data discretization: taxonomy and big data challenge”, WIREs Data Mining Knowl Discov 2015, 17 pages. [cited by applicant]
Sweke et al., “On the Quantum versus Classical Learnability of Discrete Distributions”, dated Jul. 28, 2020, 27 pages. [cited by applicant]
Ribeiro et al., “Anchors: High-precision Model- agnostic Explanations”, In Thirty-Second AAAI Conference on Artificial Intelligence, dated Apr. 2018, 9 pages. [cited by applicant]
Lundberg et al. “A Unified Approach to Interpreting Model Predictions”. dated Nov. 25, 2017, In Advances in neural information processing systems, 10 pages. [cited by applicant]
Ribeiro et al., “Why Should I Trust You?” Explaining the Predictions of Any Classifier, Publication rights licensed to ACM, dated 2016, 10 pages. [cited by applicant]
Robnik-Sikonja et al., “Explaining Classifications for Individual Instances”, IEEE Transactions on Knowledge and Data Engineering, 20:589-600, dated 2008, 24 pages. [cited by applicant]
Roth, Alvin, “The Shapley Value”, Essays in honor of Lloyd S. Shapley, Cambridge University Press, dated 1988, 338 pages. [cited by applicant]
Samek et al., “Explainableartificialintelligence: Understanding, Visualizingand Interpreting Deep Learning Models”, ITU Journal: ICT Discoveries, Special Issue No. 1, Oct. 13, 2017, 10 pages. [cited by applicant]
SciPy.Org, Statistical Functions (scipy.stats), https://scipy.org, SciPy v1.6.1, Feb. 2021, 10 pages. [cited by applicant]
Shapley, Lloyd S. “A Value for N-person Games”, Contributions to the Theory of Games 2.28, dated Aug. 21, 1951, 19 pages. [cited by applicant]
Ribeiro et al., “Why should I trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD, dated Aug. 2016, 10 pages. [cited by applicant]
Laugel. Local post-hoc interpretability for black-box classifiers. Machine Learning [cs.LG]. Sorbonne Universite (Year: 2020). [cited by applicant]
Saito et al., Improving LIME Robustness with Smarter Locality Sampling, arXiv.org, available at https://arxiv.org/pdf/2006.12302v2.pdf (revised Jul. 24, 2020). [cited by applicant]
Molnar et al., Limitations of Interpretable Machine Learning Methods, available at https://slds-lmu.github.io/iml_methods_limitations/ (Oct. 5, 2020). [cited by applicant]
Laugel et al., Defining Locality for Surrogates in Post-hoc Interpretablity, arXiv.org, available at https://arxiv.org/pdf/1806.07498.pdf (Jun. 19, 2018). [cited by applicant]
Lahiri A, Edakunni NU. Accurate and Intuitive Contextual Explanations using Linear Model Trees. arXiv preprint arXiv:2009.05322. Sep. 11, 2020, originally presented at KDD Workshop on ML in Finance 2020, Aug. 24, 2020 (… [cited by applicant]
Guidotti R, Monreale A, Ruggieri S, Pedreschi D, Turini F, Giannotti F. Local rule-based explanations of black box decision systems.arXiv preprint arXiv:1805.10820. May 28, 2018 (Year: 2018). [cited by applicant]
Deep Hyperspherical Learning, arXiv.org, available at https://arxiv.org/pdf/1711.03189.pdf (Jan. 30, 2018). [cited by applicant]
Github.com, “slundberg/shap”, https://github.com/slundberg/shap/blob/master/shap/explainers/_permutation.py, updated on Nov. 17, 2020, 5 pages. [cited by applicant]
Lin et al., “Experiencing SAX: a novel symbolic representation of time series”, Data Min Knowl Disc (2007), 38 pages. [cited by applicant]
Letham et al., “Interpretable classifiers using rules and Bayesian analysis: Building a better stroke prediction model”, vol. 9, No. 3, dated 2015, 2 pages. [cited by applicant]
Letham et al., “Interpretable Classifiers Using Rules and Bayesian Analysis: Building a Better Stroke Prediction Model”, The Annals of Applied Statistics, dated 2015, vol. 9, No. 3, 23 pages. [cited by applicant]
Laugel et al., “Defining Locality for Surrogates in Post-hoc Interpretablity”, dated Jun. 19, 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018), 7 pages. [cited by applicant]
Laugel et al., “Comparison-based Inverse Classification for Interpretability in Machine Learning”., (IPMU 2018), dated Jun. 2018, Cadix, Spain, 12 pages. [cited by applicant]
Lakkaraju et al., “Interpretable Decision Sets: A Joint Framework for Description and Prediction”, KDD, PMC dated Nov. 14, 2016, 24 pages. [cited by applicant]
Lakkaraju et al., “Interpretable Decision Sets: A Joint Framework for Description and Prediction”, KDD '16, Aug. 13-17, 2016, San Francisco, CA, USA, 10 pages. [cited by applicant]
Lakkaraju et al., “Interpretable & Explorable Approximations of Black Box Models”, KDD, dated 2017, 5 pages. [cited by applicant]
Wong et al., “A Vector Space Model for Automatic Indexing” Information Retrieval and Language Processing, Communications of ACM, dated Nov. 1975, vol. 18, No. 11, 8 pages. [cited by applicant]
Partalas et al., “LSHTC: A Benchmark for Large-Scale Text Classification”, dated Mar. 30, 2015, 9 pages. [cited by applicant]
Gabrilovich et al., “Wikipedia-Based Semantic Interpretation for Natural Language Processing”, Journal of Artificial Intelligence Research, vol. 34, dated 2009, 55 pages. [cited by applicant]
Gabrilovich et al., “Overcoming the Brittleness Bottleneck using Wikipedia: Enhancing Text Categorization with Encyclopedic Knowledge”, dated 2006, 6 pages. [cited by applicant]
Gabrilovich et al., “Computing Semantic Relatedness using Wikipedia-based Explicit Semantic Analysis”, dated Jan. 6, 2007, 6 pages. [cited by applicant]
Dong-Hyun Lee, “Multi-Stage Rocchio Classification for Large-scale Multilabeled Text data”, 8 pages. [cited by applicant]
“Rocchio Classification”, https://nlp.stanford.edu/IR-book/html/htmledition/rocchio-classification-1.html, last viewed on Jul. 6, 2017, 6 pages. [cited by applicant]
Mimno et al., “Optimizing Semantic Coherence in Topic Models” Proceedings of the 2011 Conference Empirical Methods in Language Processing, pp. 262-272. [cited by applicant]
Cited By (1)
US 12,670,320