IP Library Granted Patent US 12,333,252
Granted Patent B2
US 12,333,252 · App. 18/458,399 · Granted Jun 17, 2025

Automated system and method to prioritize language model and ontology expansion and pruning

Inventors: Ian Roy Beaver (Spokane, WA); Christopher James Jeffs (Roswell, GA)
Assignee: Verint Americas Inc.
G06F40/295G06F16/3344G06F16/367G10L15/18G10L15/197G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,252
App. No.
18/458,399
Granted
Jun 17, 2025
Kind
B2
Abstract

A system and method for updating computerized language models is provided that automatically adds or deletes terms from the language model to capture trending events or products, while maximizing computer efficiencies by deleting terms that are no longer trending and use of knowledge bases, machine learning model training and evaluation corpora, analysis tools and databases.

Claims (30)

1. A non-transitory computer readable medium comprising instructions that, when executed by a processor of a processing system, cause the processor to perform a method of updating a language model for a language domain for an interactive virtual assistant, the method comprising:

collecting trending terms from textual data monitored across multiple sources over a sliding time window;

comparing a set of the trending terms to a vocabulary in a data model to identify terms that exist in the data model;

for a selected term from the set of trending terms that exist in the data model and appears in a new context, adding the selected term to a training example in the new context for retraining the data model;

for a selected term from the set of trending terms that exist in the data model and does not appear in a new context, ignoring the selected term;

for a selected term from the set of trending terms that does not exist in the data model, checking the selected term for a frequency of use in the textual data and adding the selected term to the training example if the frequency of use has reached a predetermined frequency threshold;

recompiling the language model based on the selected term in a context corresponding to the selected term; and

adopting the recompiled language model as a trained machine learning model used by an interactive virtual assistant to interact with a user.

2. The non-transitory computer readable medium of claim 1 , the method further comprising, for a new term that is not in the data model, determining a frequency of appearance of the new term and, when the frequency reaches a predetermined new term threshold, adding the new term to the training example as the new term appears in a context corresponding to the new term.

3. The non-transitory computer readable medium of claim 2 , the method further comprising ignoring the new term that does not meet the predetermined new term threshold.

4. The non-transitory computer readable medium of claim 1 , the method further comprising deleting a rare term from the vocabulary if the frequency of use of the rare term falls below a deletion threshold.

5. The non-transitory computer readable medium of claim 1 , wherein determining the predetermined frequency threshold for adding a term not in the data model is determined automatically based on predetermined values stored in a digital storage device.

6. The non-transitory computer readable medium of claim 1 , wherein the selected term is passed to a human only when the selected term exists in the data model and appears in the new context meets a predetermined passing threshold based on parameters of an existing ontology.

7. The non-transitory computer readable medium of claim 1 , wherein the textual data comprises social media texts, customer support emails, and/or other business related social media and/or emails.

8. The non-transitory computer readable medium of claim 1 , wherein collecting trending terms from the textual data comprises collecting trending n-grams wherein the n-grams comprise unigrams, bigrams or a contiguous sequence of n items in the textual data, wherein n is a positive integer.

9. A method of updating a language model for a language domain for an interactive virtual assistant, comprising:

collecting trending terms from textual data monitored across multiple sources over a sliding time window;

comparing a set of the trending terms to a vocabulary in a data model to identify terms that exist in the data model and,

for a selected term from the set of trending terms that exist in the data model and appears in a new context, adding the selected term to a training example in the new context for retraining the data model;

for the selected term from the set of trending terms that exist in the data model and does not appears in a new context, ignoring the selected term;

for a selected term from the set of trending terms that does not exist in the data model, checking the selected term for a frequency of use in the textual data and adding the selected term to the training example if the frequency of use has reached a predetermined frequency threshold;

recompiling the language model based on the selected term in a context corresponding to the selected term; and

adopting the recompiled language model as a trained machine learning model used by an interactive virtual assistant to interact with a user.

10. The method of claim 9 , further comprising, for a new term that is not in the data model, determining a frequency of appearance of the new term and, when the frequency reaches a predetermined new term threshold, adding the new term to the training example as the new term appears in a context corresponding to the new term.

11. The method of claim 10 , further comprising ignoring the new terms that does not meet the predetermined new term threshold.

12. The method of claim 9 , further comprising deleting a rare term from the vocabulary if the frequency of use of the rare term falls below a deletion threshold.

13. The method of claim 9 , wherein determining the predetermined frequency threshold for adding a term to the data model is determined automatically based on predetermined values stored in a digital storage device.

14. The method of claim 9 , wherein the selected term is passed to a human only when the selected term exists in the data model and appears in the new context meets a predetermined passing threshold based on parameters of an existing ontology.

15. The method of claim 9 , wherein the textual data comprises social media texts, customer support emails, and/or other business related social media and/or emails.

16. The method of claim 9 , wherein collecting trending terms from the textual data comprises collecting n-grams, wherein the n-grams comprise unigrams, bigrams or a contiguous sequence of n items in the textual data, wherein n is a positive integer.

Assignments (3)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT AMERICAS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2024
From: BEAVER, IAN ROY
To: VERINT AMERICAS INC.
Reel/Frame 066643/0527 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2024
From: JEFFS, CHRISTOPHER JAMES
To: VERINT AMERICAS INC.
Reel/Frame 066643/0585 →
Continuity (3)
Continuation 16829101 · Mar 25, 2020
Provisional Application 62824429 · Mar 27, 2019
Related Publication 20240062012A1 · Feb 22, 2024
References Cited (134)
US 5317673A · Cohen et al. · 1994 [cited by applicant]
US 5737617A · Bernth et al. · 1998 [cited by applicant]
US 6076088A · Paik et al. · 2000 [cited by applicant]
US 6385579B1 · Padmanabhan et al. · 2002 [cited by applicant]
US 6434557B1 · Egilsson et al. · 2002 [cited by applicant]
US 6542866B1 · Jiang et al. · 2003 [cited by applicant]
US 6560590B1 · Shwe et al. · 2003 [cited by applicant]
US 6600821B1 · Chan et al. · 2003 [cited by applicant]
US 6718296B1 · Reynolds et al. · 2004 [cited by applicant]
US 7113958B1 · Lantrip et al. · 2006 [cited by applicant]
US 7149695B1 · Bellegarda · 2006 [cited by applicant]
US 7552053B2 · Gao et al. · 2009 [cited by applicant]
US 7734467B2 · Gao et al. · 2010 [cited by applicant]
US 7844459B2 · Budde et al. · 2010 [cited by applicant]
US 7853544B2 · Scott et al. · 2010 [cited by applicant]
US 7904414B2 · Isaacs · 2011 [cited by applicant]
US 7912701B1 · Gray et al. · 2011 [cited by applicant]
US 8036876B2 · Sanfilippo et al. · 2011 [cited by applicant]
US 8078565B2 · Arseneault et al. · 2011 [cited by applicant]
US 8160232B2 · Isaacs et al. · 2012 [cited by applicant]
US 8190628B1 · Yang et al. · 2012 [cited by applicant]
US 8260809B2 · Platt et al. · 2012 [cited by applicant]
US 8285552B2 · Wang et al. · 2012 [cited by applicant]
US 8447604B1 · Chang · 2013 [cited by applicant]
US 8825488B2 · Scoggins, II et al. · 2014 [cited by applicant]
US 8825489B2 · Scoggins, II et al. · 2014 [cited by applicant]
US 8874432B2 · Qi et al. · 2014 [cited by applicant]
US 9066049B2 · Scoggins, II et al. · 2015 [cited by applicant]
US 9189525B2 · Acharya et al. · 2015 [cited by applicant]
US 9232063B2 · Romano et al. · 2016 [cited by applicant]
US 9355354B2 · Isaacs et al. · 2016 [cited by applicant]
US 9477752B1 · Romano · 2016 [cited by applicant]
US 9569743B2 · Fehr et al. · 2017 [cited by applicant]
US 9575936B2 · Romano et al. · 2017 [cited by applicant]
US 9633009B2 · Alexe · 2017 [cited by applicant]
US 9639520B2 · Yishay · 2017 [cited by applicant]
US 9646605B2 · Biatov et al. · 2017 [cited by applicant]
US 9697246B1 · Romano et al. · 2017 [cited by applicant]
US 9720907B2 · Bangalore et al. · 2017 [cited by applicant]
US 9760546B2 · Galle · 2017 [cited by applicant]
US 9818400B2 · Paulik et al. · 2017 [cited by applicant]
US 20020032564A1 · Ehsani et al. · 2002 [cited by applicant]
US 20020128821A1 · Ehsani et al. · 2002 [cited by applicant]
US 20020188599A1 · McGreevy · 2002 [cited by applicant]
US 20030028512A1 · Stensmo · 2003 [cited by applicant]
US 20030126561A1 · Woehler et al. · 2003 [cited by applicant]
US 20040078190A1 · Fass et al. · 2004 [cited by applicant]
US 20040236737A1 · Weissman et al. · 2004 [cited by applicant]
US 20060149558A1 · Kahn et al. · 2006 [cited by applicant]
US 20060248049A1 · Cao et al. · 2006 [cited by applicant]
US 20070016863A1 · Qu et al. · 2007 [cited by applicant]
US 20070118357A1 · Kasravi et al. · 2007 [cited by applicant]
US 20080021700A1 · Moitra et al. · 2008 [cited by applicant]
US 20080046244A1 · Ohno et al. · 2008 [cited by applicant]
US 20080154578A1 · Xu et al. · 2008 [cited by applicant]
US 20080221882A1 · Bundock et al. · 2008 [cited by applicant]
US 20090012842A1 · Srinivasan et al. · 2009 [cited by applicant]
US 20090043581A1 · Abbott et al. · 2009 [cited by applicant]
US 20090063150A1 · Nasukawa et al. · 2009 [cited by applicant]
US 20090099996A1 · Stefik · 2009 [cited by applicant]
US 20090150139A1 · Jianfeng et al. · 2009 [cited by applicant]
US 20090204609A1 · Labrou et al. · 2009 [cited by applicant]
US 20090254877A1 · Kuriakose et al. · 2009 [cited by applicant]
US 20090306963A1 · Prompt et al. · 2009 [cited by applicant]
US 20090326947A1 · Arnold et al. · 2009 [cited by applicant]
US 20100030552A1 · Chen et al. · 2010 [cited by applicant]
US 20100057688A1 · Anovick et al. · 2010 [cited by applicant]
US 20100131260A1 · Bangalore et al. · 2010 [cited by applicant]
US 20100161604A1 · Mintz et al. · 2010 [cited by applicant]
US 20100275179A1 · Mengusoglu et al. · 2010 [cited by applicant]
US 20110161368A1 · Ishikawa et al. · 2011 [cited by applicant]
US 20110196670A1 · Dang et al. · 2011 [cited by applicant]
US 20120016671A1 · Jaggi et al. · 2012 [cited by applicant]
US 20120131031A1 · Xie et al. · 2012 [cited by applicant]
US 20120303356A1 · Boyle et al. · 2012 [cited by applicant]
US 20130018650A1 · Moore et al. · 2013 [cited by applicant]
US 20130066921A1 · Mark et al. · 2013 [cited by applicant]
US 20130132442A1 · Tsatsou et al. · 2013 [cited by applicant]
US 20130144616A1 · Bangalore · 2013 [cited by applicant]
US 20130166303A1 · Chang et al. · 2013 [cited by applicant]
US 20130260358A1 · Lorge et al. · 2013 [cited by applicant]
US 20140040275A1 · Dang et al. · 2014 [cited by applicant]
US 20140040713A1 · Dzik et al. · 2014 [cited by applicant]
US 20140075004A1 · Van Dusen et al. · 2014 [cited by applicant]
US 20140143157A1 · Jeffs et al. · 2014 [cited by applicant]
US 20140222419A1 · Romano · 2014 [cited by examiner]
US 20140297266A1 · Nielson et al. · 2014 [cited by applicant]
US 20150032746A1 · Lev-Tov et al. · 2015 [cited by applicant]
US 20150066503A1 · Achituv et al. · 2015 [cited by applicant]
US 20150066506A1 · Romano et al. · 2015 [cited by applicant]
US 20150074124A1 · Sexton et al. · 2015 [cited by applicant]
US 20150127652A1 · Romano · 2015 [cited by applicant]
US 20150169746A1 · Hatami-Hanza · 2015 [cited by applicant]
US 20150170040A1 · Berdugo et al. · 2015 [cited by applicant]
US 20150193532A1 · Romano · 2015 [cited by applicant]
US 20150220618A1 · Horesh et al. · 2015 [cited by applicant]
US 20150220626A1 · Carmi et al. · 2015 [cited by applicant]
US 20150220630A1 · Romano et al. · 2015 [cited by applicant]
US 20150220946A1 · Horesh et al. · 2015 [cited by applicant]
US 20160055848A1 · Meruva et al. · 2016 [cited by applicant]
US 20160078016A1 · Ng Tari et al. · 2016 [cited by applicant]
US 20160078860A1 · Paulik et al. · 2016 [cited by applicant]
US 20160117386A1 · Ajmera et al. · 2016 [cited by applicant]
US 20160180437A1 · Boston et al. · 2016 [cited by applicant]
US 20160217127A1 · Segal et al. · 2016 [cited by applicant]
US 20160217128A1 · Baum et al. · 2016 [cited by applicant]
US 20180197531A1 · Baughman · 2018 [cited by examiner]
WO 2000026795A1 · 2000 [cited by applicant]
Chung, Grace, et al. “A dynamic vocabulary spoken dialogue interface.” Proc. ICSLP. 2004. (Year: 2004). [cited by examiner]
Li, Jiwei, et al. “Dialogue learning with human-in-the-loop.” arXiv preprint arXiv:1611.09823 (2017). (Year: 2017). [cited by examiner]
Applicant-Initiated Interview Summary Received in U.S. Appl. No. 16/829,101, dated Apr. 20, 2023, 2 pages. [cited by applicant]
Chen, S.F., et al, “An Empirical Study of Smoothing Techniques for Language Modeling,” Computer Speech and Language, vol. 13, 1998, 62 pages. [cited by applicant]
Coursey, K., et al., “Topic identification using Wikipedia graph centrality,” Proceedings of the 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Human Language Tech… [cited by applicant]
Extended European Search Report issued in Application No. 19204698.5, dated Mar. 11, 2020, 7 pages. [cited by applicant]
Extended European Search Report, dated Feb. 18, 2015, received in connection with corresponding European Application No. 14182714.7, 9 pages. [cited by applicant]
Federico, M., et al., “IRSTLM: an Open Source Toolkit for Handling Large Scale Language Models,” Ninth Annual Conference of the International Speech Communication Association, 2008, pp. 1618-1621. [cited by applicant]
Final Office Action Received in U.S. Appl. No. 17/838,461, dated Aug. 8, 2023, 13 pages. [cited by applicant]
Final Office Action received in U.S. Appl. No. 17/838,459, dated Aug. 8, 2023, 19 pages. [cited by applicant]
Galescu et al, “Bi-directional conversion between graphemes and phonemes using a joint n-gram model”, 2001, In 4th ISCA Tutorial and Research Workshop (ITRW) on Speech Synthesis 2001, pp. 1-6. [cited by applicant]
International Search Report and Written Opinion, dated Jun. 24, 2020, received in connection with International Patent Application No. PCT/US2020/025134. [cited by applicant]
Král, P., et al., “Dialogue Act Recognition Approaches,” Computing and Informatics, vol. 29, 2010, pp. 227-250. [cited by applicant]
Kumar, N., et al., “Automatic Keyphrase Extraction from Scientific Documents Using N-gram Filtration Technique,” Proceedings of the 8th ACM symposium on Document engineering, 2008, pp. 199-208. [cited by applicant]
Mikolov, T., et al., “Efficient Estimation of Word Representations in Vector Space,” Proceedings of the International Conference on Learning Representations (ICLR), arXiv:1301.3781v3, 2013, 12 pages. [cited by applicant]
Non-Final Office Action Received in U.S. Appl. No. 17/838,461, dated Apr. 10, 2023, 20 pages. [cited by applicant]
Notice of Allowance received in U.S. Appl. No. 17/225,589, dated Jan. 12, 2023, 10 pages. [cited by applicant]
Office Action issued in Application No. 19204698.5, dated Oct. 19, 2021, 6 pages. [cited by applicant]
Ponte, J.M., et al., “Text Segmentation by Topic,” Computer Science Department, University of Massachusetts, Amherst, 1997, 13 pages. [cited by applicant]
Ramos, J., “Using TF-IDF to Determine Word Relevance in Document Queries,” Proceedings of the First Instructional Conference on Machine Learning, 2003, 4 pages. [cited by applicant]
Rosenfeld, R. “The CMU Statistical Language Modeling Toolkit and its use in the 1994 ARPA CSR Evaluation,” Proceedings of the Spoken Language Systems Technology Workshop, 1995, pp. 47-50. [cited by applicant]
Saon et al, “Data-driven approach to designing compound words for continuous speech recognition.”, 2001, IEEE Transactions on Speech and audio processing. May 2001;9(4):327-32. [cited by applicant]
Stolcke, A., “SRILM—An Extensible Language Modeling Toolkit,” Seventh International Conference on Spoken Language Processing, 2002, 4 pages. [cited by applicant]
Stolcke, A., et al., “Automatic Linguistic Segmentation of Conversational Speech,” IEEE, vol. 2, 1996, pp. 1005-1008. [cited by applicant]
Summons to Attend Oral Proceedings in European Application No. 19 204 698.5, dated Mar. 22, 2023, 9 pages. [cited by applicant]
Zimmerman, M., et al., “Joint Segmentation and Classification of Dialog Acts in Multiparty Meetings,” Acoustics, Speech and Signal Processing, 2006, pp. 581-584. [cited by applicant]