IP Library › Granted Patent US 12,694,208
Granted Patent B2
US 12,694,208 · App. 18/187,875 · Granted Jul 28, 2026

Domain-specificity prediction for natural language processing

Inventors: Diego Matteo Antognini (Ruvigliana, CH); Francesco Fusco (Zurich, CH)
Assignee: International Business Machines Corporation
G06F40/284G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,208
App. No.
18/187,875
Filed
Mar 22, 2023
Granted
Jul 28, 2026
Kind
B2
Art Unit
2658
USPC
704/9
Abstract

A method, computer-program product and computer system are provided to determine domain-specificity of a text term. A processor receives a plurality of domain-specific text corpora, wherein each of the plurality of domain-specific text corpora comprises a plurality of text documents of a respective domain. A processor trains a set of subword-unit tokenizers with at least two different vocabulary sizes of the respective domain-specific text corpus. A processor receives the text-term. A processor determines a domain-specificity fingerprint of the text-term, wherein the domain-specificity fingerprint comprises for each subword-unit tokenizer a number of subword-units required to represent the text-term. A processor provides the domain-specificity fingerprint for determining the domain-specificity of the text term.

Claims (41)

1 . A computer-implemented method for determining domain-specificity of a text term to improve precision of pre-trained language models in the computerized field of natural language processing, the method comprising:

receiving a plurality of domain-specific text corpora, wherein each of the plurality of domain-specific text corpora comprises a plurality of text documents associated with a respective domain;

training a set of subword-unit tokenizers with at least two different vocabulary sizes of the respective domain-specific text corpus;

receiving the text-term;

determining a domain-specificity fingerprint of the text-term, wherein the domain-specificity fingerprint comprises for each subword-unit tokenizer a number of subword-units required to represent the text-term; and

providing the domain-specificity fingerprint for determining the domain-specificity of the text term, wherein the domain-specificity of the text term is a measure of whether a term is used predominantly in one or more specific domains or whether it is used in a plurality of domains, wherein the domain-specificity of the text term is determined at least in part by analyzing a contour of the domain-specificity fingerprint and the contour comprises a variable corresponding to a respective subword-unit tokenizer of the domain at which a minimum number of subword units required to encode the text term across all domains is reached.

2 . The computer-implemented method of claim 1 , wherein the text term comprises words or multi-words.

3 . The computer-implemented method of claim 1 , wherein determining the domain-specificity fingerprint further comprises:

assigning the text term to one or more domains of the respective set of subword-unit tokenizers that compute a minimum number of subword-units required to represent the text-term with the smallest vocabulary size.

4 . The computer-implemented method of claim 1 , wherein determining the domain specificity by analyzing the contour further comprises:

computing the maximum delta between any pair of values in the contour; and

assigning the maximum delta as a score for the domain-specificity of the text term.

5 . The computer-implemented method of claim 2 , wherein the domain-specificity fingerprint of the multi-words are based on a domain provenance of the text-term using a k-nearest neighbor algorithm (k-NN) classifier.

6 . A computer program product for determining domain-specificity of a text term to improve precision of pre-trained language models in the computerized field of natural language processing, the computer program product comprising:

one or more non-transitory computer-readable storage media; and

program instructions stored on the one or more non-transitory computer-readable storage media, the program instructions to perform-operations comprising:

receiving a plurality of domain-specific text corpora, wherein each of the plurality of domain-specific text corpora comprises a plurality of text documents associated with a respective domain;

training a set of subword-unit tokenizers with at least two different vocabulary sizes of the respective domain-specific text corpus;

receiving the text-term;

determining a domain-specificity fingerprint of the text-term, wherein the domain-specificity fingerprint comprises for each subword-unit tokenizer a number of subword-units required to represent the text-term; and

providing the domain-specificity fingerprint for determining the domain-specificity of the text term, wherein the domain-specificity of the text term is a measure of whether a term is used predominantly in one or more specific domains or whether it is used in a plurality of domains and the domain specificity is determined by analyzing a contour of the domain-specificity fingerprint by computing a maximum delta between any pair of values in the contour and assigning the maximum delta as a score for the domain-specificity of the text term.

7 . The computer program product of claim 6 , wherein the text term comprises words or multi-words.

8 . The computer program product of claim 6 , wherein determining the domain-specificity fingerprint further comprises:

assigning the text term to one or more domains of the respective set of subword-unit tokenizers that compute a minimum number of subword-units required to represent the text-term with the smallest vocabulary size.

9 . The computer program product of claim 6 , wherein the contour comprises a value corresponding to a respective subword-unit tokenizer of the domain at which a minimum number of subword units required to encode the text term across all domains and vocabulary sizes is reached.

10 . The computer program product of claim 7 , wherein the domain-specificity fingerprint of the multi-words are based on a domain provenance of the text-term using a k-nearest neighbor algorithm (k-NN) classifier.

11 . A computer system for determining domain-specificity of a text term to improve precision of pre-trained language models in the computerized field of natural language processing, the computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the computer readable storage media for execution by at least one of the one or more processors, the program instructions comprising:

program instructions to receive a plurality of domain-specific text corpora, wherein each of the plurality of domain-specific text corpora comprises a plurality of text documents associated with a respective domain;

program instructions to train a set of subword-unit tokenizers with at least two different vocabulary sizes of the respective domain-specific text corpus;

program instructions to receive the text-term;

program instructions to determine a domain-specificity fingerprint of the text-term, wherein the domain-specificity fingerprint comprises for each subword-unit tokenizer a number of subword-units required to represent the text-term; and

program instructions to provide the domain-specificity fingerprint for determining the domain-specificity of the text term, wherein the domain-specificity of the text term is a measure of whether a term is used predominantly in one or more specific domains or whether it is used in a plurality of domains, wherein the domain-specificity of the text term is determined at least in part by analyzing a contour of the domain-specificity fingerprint and the contour comprises a variable corresponding to a respective subword-unit tokenizer of the domain.

12 . The computer system of claim 11 , wherein the text term comprises words or multi-words.

13 . The computer system of claim 11 , wherein determining the domain-specificity fingerprint further comprises:

program instructions to assign the text term to one or more domains of the respective set of subword-unit tokenizers that compute a minimum number of subword-units required to represent the text-term with the smallest vocabulary size.

14 . The computer system of claim 11 , wherein determining the domain specificity by analyzing the contour further comprises:

program instructions to compute the maximum delta between any pair of values in the contour; and

program instructions to assign the maximum delta as a score for the domain-specificity of the text term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2023
From: ANTOGNINI, DIEGO MATTEO; FUSCO, FRANCESCO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063058/0089 →
Continuity (1)
Related Publication 20240320429A1 · Sep 26, 2024
References Cited (58)
US 8250092B2 · Gollapudi · 2012 [cited by applicant]
US 8370319B1 · Krynski · 2013 [cited by applicant]
US 8583640B2 · Zhang · 2013 [cited by applicant]
US 10157223B2 · Misra · 2018 [cited by applicant]
US 11030999B1 · Yu · 2021 [cited by examiner]
US 11361571B1 · Fusco · 2022 [cited by applicant]
US 20090177463A1 · Gallagher et al. · 2009 [cited by applicant]
US 20130086509A1 · Satyanarayana · 2013 [cited by applicant]
US 20160117386A1 · Ajmera · 2016 [cited by examiner]
US 20170060842A1 · Dwarakanath et al. · 2017 [cited by applicant]
US 20180144744A1 · Badarinath et al. · 2018 [cited by applicant]
US 20180260383A1 · Beller · 2018 [cited by applicant]
US 20180260472A1 · Kelsey et al. · 2018 [cited by applicant]
US 20210141798A1 · Steedman Henderson · 2021 [cited by examiner]
US 20220101113A1 · Tam · 2022 [cited by examiner]
US 20220179906A1 · Desai · 2022 [cited by examiner]
US 20220279014A1 · Stokes, III · 2022 [cited by examiner]
US 20220382972A1 · El-Kurdi · 2022 [cited by examiner]
US 20220383096A1 · Zhu et al. · 2022 [cited by applicant]
US 20230017396A1 · Weerasinghe et al. · 2023 [cited by applicant]
US 20230055769A1 · Fusco · 2023 [cited by applicant]
US 20240095268A1 · Dhar · 2024 [cited by examiner]
US 20240184818A1 · Fume · 2024 [cited by examiner]
US 20240241902A1 · Shalmashi · 2024 [cited by examiner]
US 20240289551A1 · Agarwal · 2024 [cited by examiner]
US 20240320249A1 · Fusco et al. · 2024 [cited by applicant]
CN 107679244A · 2018 [cited by examiner]
Eigenmann et al, “Evaluating Text Classification Models on Multilingual Documents”, 2021, Master Thesis, Department of Informatics, University of Fribourg, pp. 1-42 (Year: 2021). [cited by examiner]
Lee et al, “Optimizing Domain Specificity of Transformer-based Language Models for Extractive Summarization of Financial News Articles in Korean”, 2021, InProceedings of the 35th Pacific Asia Conference on Language, Inf… [cited by examiner]
Bordea, Georgeta, “Domain adaptive extraction of topical hierarchies for Expertise Mining”, NUI Galway, Sep. 11, 2013, 191 pages. [cited by applicant]
Chirkova et al., “CodeBPE: Investigating Subtokenization Options for Large Language Model Pretraining on Source Code”, ICLR Workshop on Deep Learning for Code, 2022, 13 pages. [cited by applicant]
Constant et al., “Multiword Expression Processing: A Survey”, Computational Linguistics, vol. 43, No. 4, © 2017 Association for Computational Linguistics, 56 pages. [cited by applicant]
Dowlagar et al., “Unsupervised Technical Domain Terms Extraction using Term Extractor”, arXiv:2101.09015v1 [cs.CL] Jan. 22, 2021, 4 pages. [cited by applicant]
He, Tiantian, “Specificity Prediction for Sentences in Press Releases”, Uppsala University, Jun. 17, 2020,31 pages. [cited by applicant]
Kim et al., “An Unsupervised Approach to Domain-Specific Term Extraction”, printed on Dec. 20, 2022, 5 pages, <https://aclanthology.org/U09-1013.pdf>. [cited by applicant]
Korkontzelos, Ioannis, “Unsupervised Learning of Multiword Expressions”, Ph.D, Thesis, The University of York, Sep. 20, 2010, 258 pages. [cited by applicant]
Qi et al., “Deep Learning for Character-based Information Extraction”, printed on Dec. 20, 2022, 9 pages, <http://www.cs.cmu.edu/~qyj/zhSenna/2014_ecir2014_full.pdf>. [cited by applicant]
Riedl et al., “A Single Word is not Enough: Ranking Multiword Expressions Using Distributional Semantics”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal, Sep. 1… [cited by applicant]
Ryu et al., “Determining the Specificity of Terms based on Information Theoretic Measures”, CompuTerm 2004 Poster Session—3rd International Workshop on Computational Terminology, 5 pages. [cited by applicant]
Sachidananda et al., “Efficient Domain Adaptation of Language Models via Adaptive Tokenization”, Proceedings of the 2nd Workshop on Simple and Efficient Natural Language Processing, Nov. 10, 2021, © 201 Association for … [cited by applicant]
Zhang et al., “Improving Domain-specific Entity Recognition with Automatic Term Recognition and Feature Extraction”, 2010, 8 pages., <http:/lrec.elra.info/proceedings/lrec2010/pdf/214_Paper.pdf>. [cited by applicant]
Fusco et al., Determining Specificity of Text Terms in Application Contexts, U.S. Appl. No. 18/187,862, filed Mar. 22, 2023, 41 pages. [cited by applicant]
IBM Appendix P, list of patents and patent applications treated as related, Filed Herewith, 2 pages. [cited by applicant]
“SimGet is a service to train and query word embeddings (word2vec and Glove)”, zrl-cogsys/simget, Jan. 15, 2025, 6 pages. [cited by applicant]
Baeza-Yates, “Semantic query understanding,” in Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, ser. SIGIR '17. New York, NY, USA: Association for Computi… [cited by applicant]
Beltagy et al., “Scibert: Pretrained contextualized embeddings for scientific text,” arXiv:1903.10676, Sep. 2019, 6 pages. [cited by applicant]
Bostrom et al., “Byte pair encoding is suboptimal for language model pretraining,” in Findings of the Association for Computational Linguistics: EMNLP 2020. Online: Association for Computational Linguistics, Nov. 2020, … [cited by applicant]
Fix et al., “Discriminatory analysis. nonparametric discrimination: Consistency properties,” International Statistical Review, vol. 57, No. 3, Dec. 1989, 138 pages. [cited by applicant]
Kudo, “Subword regularization: Improving neural network translation models with multiple subword candidates,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Pape… [cited by applicant]
Lo et al., “S2ORC: The Semantic Scholar Open Research Corpus”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020, pp. 4969-4983. [cited by applicant]
Schuster et al., “Japanese and korean voice search,” in International Conference on Acoustics, Speech and Signal Processing, 2012, pp. 5149-5152. [cited by applicant]
Sennrich et al., “Neural Machine Translation of Rare Words with Subword Units”, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Aug. 2016, pp. 1715-1725. [cited by applicant]
Staar et al., “IBM Research's open-source toolkit for Deep Search”, Retrieved from: https://research.ibm.com/blog/deep-search-toolkit, Jul. 2022, 6 pages. [cited by applicant]
Unknown, “Deep Search”, Retrieved from: https://ds4sd.github.io/, 2023, 11 pages. [cited by applicant]
Unknown, “Expanding Concept Understanding in Microsoft Academic Graph”, Retrieved from: https://www.microsoft.com/en-us/research/articles/expanding-concept-understanding-in-microsoft-academic-graph/, Feb. 26, 2020, 6 pa… [cited by applicant]
Unknown, “Inside—outside—beginning (tagging)”, Retrieved from: https://en.wikipedia.org/wiki/Inside%E2%80%93outside%E2%80%93beginning_(tagging), Sep. 2013, 3 pages. [cited by applicant]
Zheng et al., “A survey of query result diversification,” Knowl. Inf. Syst., vol. 51, No. 1, Apr. 2017, 46 pages. [cited by applicant]
Tran, Hanh Thi Hong, et al. “The recent advances in automatic term extraction: A survey.” arXiv preprint arXiv:2301.06767, 2023, 25 pages. [cited by applicant]