IP Library › Granted Patent US 12,387,043
Granted Patent B2
US 12,387,043 · App. 18/472,746 · Granted Aug 12, 2025

Generating an improved named entity recognition model using noisy data with a self-cleaning discriminator model

Inventors: Ruiyi Zhang (San Jose, CA); Zhendong Chu (Charlottesville, VA); Vlad Morariu (Potomac, MD); Tong Yu (San Jose, CA); Rajiv Jain (Falls Church, VA); Nedim Lipka (Santa Clara, CA); Jiuxiang Gu (College Park, MD)
Assignee: Adobe Inc.
G06F40/295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,043
App. No.
18/472,746
Granted
Aug 12, 2025
Kind
B2
Abstract

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that train a named entity recognition (NER) model with noisy training data through a self-cleaning discriminator model. For example, the disclosed systems utilize a self-cleaning guided denoising framework to improve NER learning on noisy training data via a guidance training set. In one or more implementations, the disclosed systems utilize, within the denoising framework, an auxiliary discriminator model to correct noise in the noisy training data while training an NER model through the noisy training data. For example, while training the NER model to predict labels from the noisy training data, the disclosed systems utilize a discriminator model to detect noisy NER labels and reweight the noisy NER labels provided for training in the NER model.

Claims (73)

1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, utilizing a named entity recognition model, a predicted label from a first training sentence of a first set of training data, wherein the first set of training data comprises a noisy training data set;

generating a loss for the predicted label by comparing the predicted label to a ground truth label for the first training sentence;

determining, utilizing a discriminator model, a discriminator weight for the predicted label, wherein the discriminator model is trained on a second set of training data, wherein:

the second set of training data comprises a guidance training data set, and

the noisy training data set comprises a greater error rate than the guidance training data set; and

modifying parameters of the named entity recognition model utilizing the loss for the predicted label and the discriminator weight for the predicted label.

2. The non-transitory computer-readable medium of claim 1 , wherein generating, utilizing the named entity recognition model, the predicted label from the first training sentence comprises generating an entity label or a class label .

3. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

generating, utilizing the named entity recognition model, an additional predicted label from a second training sentence of the first set of training data;

determining, utilizing the discriminator model, an additional discriminator weight for the additional predicted label; and

modifying the parameters of the named entity recognition model based on an additional measure of loss based on the additional predicted label and the additional discriminator weight.

4. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise generating the predicted label by generating a predicted entity label or a predicted class label for a word within the first training sentence.

5. The non-transitory computer-readable medium of claim 4 , wherein the operations further comprise:

determining, utilizing the discriminator model, the discriminator weight for the predicted label by determining an entity label discriminator weight for the predicted entity label and a class label discriminator weight for the predicted class label; and

modifying the parameters of the named entity recognition model based on the entity label discriminator weight and the class label discriminator weight.

6. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

identifying a second training sentence from the second set of training data;

identifying, for the second training sentence, a prompt sentence indicating an entity and a class for the entity described within the second training sentence;

generating, utilizing the discriminator model, an authenticity prediction for the prompt sentence; and

modifying parameters of the discriminator model based on a comparison of the authenticity prediction and a ground truth authenticity prediction for the prompt sentence from the second set of training data.

7. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

identifying a second training sentence from the second set of training data based on a comparison between the first training sentence and the second training sentence; and

generating, utilizing the named entity recognition model, the predicted label for the first training sentence based on a demonstration from the second training sentence.

8. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise adding the first training sentence and the corresponding predicted label to the second set of training data based on an accuracy metric for the predicted label and the discriminator weight for the predicted label.

9. The non-transitory computer-readable medium of claim 8 , wherein the operations further comprise modifying the parameters of the discriminator model based on an authenticity prediction by the discriminator model on the first training sentence and the corresponding predicted label.

10. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise determining the discriminator weight by determining, utilizing the discriminator model, an authenticity score of the predicted label indicating an authenticity of the predicted label in context of the first training sentence.

11. A system comprising:

a memory component comprising a named entity recognition model and a discriminator model; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

generating, utilizing the named entity recognition model, a predicted entity label from a first training sentence and a predicted class label from the first training sentence, wherein the first training sentence is from a noisy training data set;

generating, utilizing a discriminator model, a first discriminator weight for the predicted entity label and a second discriminator weight for the predicted class label, wherein:

the discriminator model is trained on a guidance training data set, and the noisy training data set comprises a greater error rate than the guidance training data set;

determining a first weighted measure of loss from the first discriminator weight and a second weighted measure of loss from the second discriminator weight; and

modifying parameters of the named entity recognition model based on the first weighted measure of loss and the second weighted measure of loss.

12. The system of claim 11 , wherein the operations further comprise:

generating an entity label loss for the predicted entity label by comparing the predicted entity label to a ground truth entity label for the first training sentence;

determining the first weighted measure of loss based on the entity label loss and the first discriminator weight;

generating a class label loss for the predicted class label by comparing the predicted class label to a ground truth class label for the first training sentence; and

determining the second weighted measure of loss based on the class label loss and the second discriminator weight.

13. The system of claim 11 , wherein the operations further comprise:

generating, utilizing the named entity recognition model, an additional predicted label from a second training sentence;

determining, utilizing the discriminator model, a third discriminator weight for the additional predicted label; and

modifying the parameters of the named entity recognition model based on a third measure of loss based on the additional predicted label and the third discriminator weight.

14. The system of claim 11 , wherein the operations further comprise generating, utilizing the discriminator model, the first discriminator weight for the predicted entity label by:

generating a prompt sentence for the predicted entity label indicating a word from the first training sentence as an entity; and

generating the first discriminator weight by generating, utilizing the discriminator model, an authenticity prediction score for the prompt sentence.

15. The system of claim 11 , wherein the operations further comprise:

generating the predicted entity label to indicate a word from the first training sentence as an entity; and

generating the predicted class label to indicate a class for the word, wherein the class classifies the word as a place, a person, or an object.

16. A computer-implemented method comprising:

generating, utilizing a named entity recognition model, a predicted label from a training sentence, wherein the training sentence is from a noisy training data set;

determining, utilizing a discriminator model, a discriminator weight for the predicted label wherein:

the discriminator model is trained on a guidance training data set, and

the noisy training data set comprises a greater error rate than the guidance training data set;

generating a weighted loss for the predicted label based on the predicted label, a ground truth label for the training sentence, and the discriminator weight; and

modifying parameters of the named entity recognition model utilizing the weighted loss for the predicted label.

17. The computer-implemented method of claim 16 , further comprising generating the weighted loss for the predicted label by:

generating a loss for the predicted label by comparing the predicted label to the ground truth label for the training sentence; and

generating the weighted loss based on a combination of the loss and the discriminator weight.

18. The computer-implemented method of claim 16 , further comprising:

generating, utilizing the named entity recognition model, an additional predicted label from an additional training sentence;

determining, utilizing the discriminator model, an additional discriminator weight for the additional predicted label; and

modifying the parameters of the named entity recognition model utilizing an additional weighted loss based on the additional predicted label, an additional ground truth label for the additional training sentence, and the additional discriminator weight.

19. The computer-implemented method of claim 16 , further comprising:

generating the predicted label by generating a predicted entity label and a predicted class label for a word within the training sentence;

determining, utilizing the discriminator model, the discriminator weight for the predicted label by determining an entity label discriminator weight for the predicted entity label and a class label discriminator weight for the predicted class label; and

generating the weighted loss by:

generating an entity label weighted loss based on the predicted entity label, a ground truth entity label for the training sentence, and the entity label discriminator weight; and

generating a class label weighted loss based on the predicted class label, a ground truth class label for the training sentence, and the class label discriminator weight.

20. The computer-implemented method of claim 19 , further comprising:

generating the predicted entity label to indicate a word from the training sentence as an entity; and

generating the predicted class label to indicate a class for the word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: ZHANG, RUIYI; CHU, ZHENDONG; MORARIU, VLAD; YU, TONG; JAIN, RAJIV; LIPKA, NEDIM; GU, JIUXIANG
To: ADOBE INC.
Reel/Frame 064997/0872 →
Continuity (1)
Related Publication 20250103813A1 · Mar 27, 2025
References Cited (39)
US 10628483B1 · Rao · 2020 [cited by examiner]
US 20220269939A1 · Zhao · 2022 [cited by examiner]
US 20230325599A1 · Nezami · 2023 [cited by examiner]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9. [cited by applicant]
Antti Tarvainen and Harri Valpola. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing systems, 30. [cited by applicant]
Bianca Zadrozny and Charles Elkan. 2001. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml, vol. 1, pp. 609-616. Citeseer. [cited by applicant]
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. 2018. Co-teaching: Robust training of deep neural networks with extremely noisy labels. Advances in neural information pr… [cited by applicant]
Chen Liang, Yue Yu, Haoming Jiang, Siawpeng Er, Ruijia Wang, Tuo Zhao, and Chao Zhang. 2020. Bond: Bert-assisted open-domain named entity recognition with distant supervision. In Proceedings of the 26th ACM SIGKDD Inter… [cited by applicant]
Christopher Schröder, Andreas Niekler, and Martin Potthast. 2021. Revisiting uncertainty-based query strategies for active learning with transformers. arXiv preprint arXiv:2107.05687. [cited by applicant]
Dan Hendrycks, Mantas Mazeika, Duncan Wilson, and Kevin Gimpel. 2018. Using trusted data to train deep networks on labels corrupted by severe noise. Advances in neural information processing systems, 31. [cited by applicant]
Dominic Balasuriya, Nicky Ringland, Joel Nothman, Tara Murphy, and James R Curran. 2009. Named entity recognition in wikipedia. In Proceedings of the 2009 workshop on the people's web meets NLP: Collaboratively construc… [cited by applicant]
Dong-Ho Lee, Mahak Agarwal, Akshen Kadakia, Jay Pujara, and Xiang Ren. 2021. Good examples make a faster learner: Simple demonstration-based learning for low-resource ner. arXiv preprint arXiv:2110.08454. [cited by applicant]
Eric Arazo, Diego Ortego, Paul Albert, Noel E O'Connor, and Kevin McGuinness. 2020. Pseudo-labeling and confirmation bias in deep semi-supervised learning. In 2020 International Joint Conference on Neural Networks (IJCN… [cited by applicant]
Erik F Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050. [cited by applicant]
Filipe Rodrigues and Francisco Pereira. 2018. Deep learning from crowds. In Proceedings of the AAAI conference on artificial intelligence, vol. 32. [cited by applicant]
Filipe Rodrigues, Francisco Pereira, and Bernardete Ribeiro. 2014. Sequence labeling with multiple annotators. Machine learning, 95:165-181. [cited by applicant]
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. arXiv preprint arXiv:1603.01360. [cited by applicant]
Haoming Jiang, Danqing Zhang, Tianyu Cao, Bing Yin, and Tuo Zhao. 2021. Named entity recognition with small strongly labeled and large weakly labeled data. arXiv preprint arXiv:2106.08977. [cited by applicant]
Hongxin Zhang, Yanzhe Zhang, Ruiyi Zhang, and Diyi Yang. 2022. Robustness of demonstration-based learning under limited data scenario. arXiv preprint arXiv:2210.10693. [cited by applicant]
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:22… [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. [cited by applicant]
Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2020. A survey on deep learning for named entity recognition. IEEE Transactions on Knowledge and Data Engineering, 34(1):50-70. [cited by applicant]
John Lafferty, Andrew McCallum, and Fernando CN Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. [cited by applicant]
Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. 2019. Meta-weight-net: Learning an explicit mapping for sample weighting. Advances in neural information processing systems, 32. [cited by applicant]
Kun Liu, Yao Fu, Chuanqi Tan, Mosha Chen, Ningyu Zhang, Songfang Huang, and Sheng Gao. 2021a. Noisy-labeled ner with confidence estimation. In Proceedings of the 2021 Conference of the North American Chapter of the Asso… [cited by applicant]
Lance A Ramshaw and Mitchell P. Marcus. 1999. Text chunking using transformation-based learning. In Natural language processing using very large corpora, pp. 157-176. Springer. [cited by applicant]
Linzhi Wu, Pengjun Xie, Jie Zhou, Meishan Zhang, Chunping Ma, Guangwei Xu, and Min Zhang. 2022. Self-augmentation for named entity recognition with meta reweighting. arXiv preprint arXiv:2204.11406. [cited by applicant]
Michael A Hedderich, Dawei Zhu, and Dietrich Klakow. 2021. Analysing the noise model error for realistic noisy label data. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 7675-7684. [cited by applicant]
Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084. [cited by applicant]
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXi… [cited by applicant]
Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 Idc2013t19. Linguistic Data Con… [cited by applicant]
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat… [cited by applicant]
Tim Finin, Will Murnane, Anand Karandikar, Nicholas Keller, Justin Martineau, Mark Dredze, et al. 2010. Annotating named entities in twitter data with crowd-sourcing. In Proceedings of the NAACL Workshop on Creating Spe… [cited by applicant]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in n… [cited by applicant]
Yazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen, Qi Wu, Jian Zhang, and Zhenmin Tang. 2021. Jo-src: A contrastive approach for combating noisy labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa… [cited by applicant]
Mnhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining ap-proach. arXiv preprint arXiv… [cited by applicant]
Yu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang, Yu Zhang, Heng Ji, and Jiawei Han. 2021. Distantly-supervised named entity recognition with noise-robust learning and language model augmented self-training. In Proceedings… [cited by applicant]
Zhendong Chu and Hongning Wang. 2021. Improve learning from crowds via generative augmentation. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 167-175. [cited by applicant]
Zhendong Chu, Jing Ma, and Hongning Wang. 2021. Learning from crowds by modeling common confusions. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 5832-5840. [cited by applicant]