IP Library › Granted Patent US 12,353,998
Granted Patent B2
US 12,353,998 · App. 17/534,421 · Granted Jul 8, 2025

Multi-task sequence tagging with injection of supplemental information

Inventors: Luis Gerardo Mojica De La Vega (Redmond, WA); Qiang Lou (Sammamish, WA); Jian Jiao (Bellevue, WA); Ruofei Zhang (Mountain View, CA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/08G06F16/93G06N3/045G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,998
App. No.
17/534,421
Granted
Jul 8, 2025
Kind
B2
Abstract

A tagging system appends supplemental information to an original sequence of items, to produce a supplemented sequence of items. The tagging system includes a transformer-based encoder neural network that maps the supplemented sequence into hidden state information. The tagging system includes a post-processing neural network that transform the hidden state information into a tagged output sequence of items. That is, each item in the tagged output sequence includes a tag that identifies its entity class or some other characteristic. The tagging system can increase the accuracy of the tags it produces by virtue of the inclusion of the supplemental information added to each original sequence. A training system trains the tagging system to perform plural tasks, which further increases the accuracy of the tags it produces. The training system may commence training of the tagging system using a pre-trained model for the encoder neural network.

Claims (59)

1. A computer-implemented method for tagging sequences of items, comprising:

obtaining an original sequence of items from at least one source of original information;

obtaining supplemental information pertaining to the original sequence of items from a search system, the search system including matching logic that maps the original sequence of items to the supplemental information;

appending the supplemental information to the original sequence of items, with a separator token therebetween, to produce a supplemented sequence of items;

mapping the supplemented sequence of items into hidden state information using an encoder machine-trained model;

processing the hidden state information with a particular post-processing machine-trained model, to produce a tagged output sequence of items; and

providing output information that is based on the output sequence of items,

the encoder machine-trained model and the particular post-processing machine-trained model having been trained in a prior training process,

the prior training process including:

obtaining plural sets of training examples, the plural sets of training examples being generated based on plural respective data sets;

selecting a training example from a chosen set of training examples, the training example including: a supplemented sequence of items that includes an original sequence of items having text combined with supplemental information obtained from the search system; and labels that identify respective entity classes of the items in the original sequence of items of the training example,

the supplemental information associated with the training example being obtained by: obtaining search results generated by the search system for the original sequence of items having text in the training example, the search results including a set of matching-document digests that describe documents that match the original sequence of items having text, as determined by the search system; and selecting one or more supplemental items from the search results;

mapping the supplemented sequence of items of the training example into hidden state information using the encoder machine-trained model;

processing the hidden state information of the training example with a selected post-processing machine-trained model, to produce a tagged output sequence of items for the training example, each particular item in the tagged output sequence of items for the training example having a tag that identifies a class of entity to which the particular item pertains, the selected post-processing machine-trained model being selected from among plural post-processing machine-trained models, the plural post-processing machine-trained models being trained using the plural respective sets of training examples;

adjusting weights of the encoder machine-trained model and the selected post-processing machine-trained model based on a comparison between tags in the tagged output sequence of items associated with the training example and the labels of the training example; and

repeating said training process until a training objective is achieved.

2. The computer-implemented method of claim 1 , wherein one supplemental item appended to the original sequence of items of the training example is a portion of a document address extracted from one of the matching-document digests.

3. The computer-implemented method of claim 1 , wherein one supplemental item appended to the original sequence of items of the training example is a portion of a document title extracted from one of the matching-document digests.

4. The computer-implemented method of claim 1 , wherein one supplemental item appended to the original sequence of items of the training example is a portion of a document summary extracted from one of the matching document digests.

5. The computer-implemented method of claim 1 , wherein said appending also comprises placing separator tokens between each neighboring pair of supplemental items that make up the supplemental information that is appended.

6. A computing system for performing a training process, comprising:

hardware logic circuitry, the hardware logic circuitry corresponding to: (a) one or more hardware processors that perform operations by executing machine-readable instructions stored in a memory, and/or (b) one or more other hardware logic units that perform the operations using a collection of configured logic gates, the operations including:

obtaining plural sets of training examples, the plural sets of training examples being generated based on plural respective data sets;

selecting a training example from a chosen set of training examples, the training example including: a supplemented sequence of items that includes an original sequence of items having text combined with supplemental information obtained from at least one source, said at least one source including matching logic that maps the original sequence of items to the supplemental information; and labels that identify respective entity classes of the items in the original sequence of items,

the supplemental information being obtained by: obtaining search results generated by a search system for the original sequence of items having text, the search results including a set of matching-document digests that describe documents that match the original sequence of items having text, as determined by the search system; and selecting one or more supplemental items from the search results;

mapping the supplemented sequence of items into hidden state information using an encoder machine-trained model;

processing the hidden state information with a post-processing machine-trained model, to produce a tagged output sequence of items, each particular item in the tagged output sequence of items having a tag that identifies a class of entity to which the particular item pertains,

the post-processing machine-trained model being selected from among plural post-processing machine-trained models, the plural post-processing machine-trained models being trained using the plural respective sets of training examples;

adjusting weights of the encoder machine-trained model and the post-processing machine-trained model based on a comparison between tags in the tagged output sequence of items and the labels of the training example; and

repeating said selecting, mapping, processing, and adjusting plural times until a training objective is achieved.

7. The computing system of claim 6 , wherein the training example does not assign respective entity-specific labels to the supplemental items.

8. The computing system of claim 6 , wherein the training example assigns a same default label to each of the plural supplemental items.

9. The computing system of claim 6 , wherein one supplemental item is a portion of a document address extracted from one of the matching-document digests.

10. The computing system of claim 6 , wherein one supplemental item is a portion of a document title extracted from one of the matching-document digests.

11. The computing system of claim 6 , wherein one supplemental item is a portion of a document summary extracted from one of the matching document digests.

12. The computing system of claim 6 , wherein the encoder machine-trained model is a transformer-based encoder machine-trained model that is pre-trained, prior to the training process, based on a multilingual set of training examples.

13. The computing system of claim 6 , wherein the training examples in the plural sets of training examples include text expressed in a single particular natural language, the transformer-based encoder machine-trained model and the post-processing machine-trained model, once trained, also being capable of producing tagged output sequences of items for natural languages other than the particular natural language.

14. The computing system of claim 6 , wherein the plural post-processing machine-trained models use different respective label vocabularies.

15. A non-transitory computer-readable storage medium for storing computer-readable instructions, the computer-readable instructions, when executed by one or more hardware processors, performing a method that comprises:

obtaining an original sequence of items from at least one source of original information;

obtaining supplemental information pertaining to the original sequence of items from a search system, the search system including matching logic that maps the original sequence of items to the supplemental information;

appending the supplemental information to the original sequence of items, with a separator token therebetween, to produce a supplemented sequence of items;

mapping the supplemented sequence of items into hidden state information using an encoder machine-trained model;

processing the hidden state information with a particular post-processing machine-trained model, to produce a tagged output sequence of items; and

providing output information that is based on the output sequence of items,

the encoder machine-trained model and the particular post-processing machine-trained model having been trained in a prior training process,

the prior training process including:

obtaining plural sets of training examples, the plural sets of training examples being generated based on plural respective data sets;

selecting a training example from a chosen set of training examples, the training example including: a supplemented sequence of items that includes an original sequence of items having text combined with supplemental information obtained from the search system; and labels that identify respective entity classes of the items in the original sequence of items of the training example,

the supplemental information associated with the training example being obtained by: obtaining search results generated by the search system for the original sequence of items having text in the training example, the search results including a set of matching-document digests that describe documents that match the original sequence of items having text, as determined by the search system; and selecting one or more supplemental items from the search results;

mapping the supplemented sequence of items of the training example into hidden state information using the encoder machine-trained model;

processing the hidden state information of the training example with a selected post-processing machine-trained model, to produce a tagged output sequence of items for the training example, each particular item in the tagged output sequence of items for the training example having a tag that identifies a class of entity to which the particular item pertains, the selected post-processing machine-trained model being selected from among plural post-processing machine-trained models, the plural post-processing machine-trained models being trained using the plural respective sets of training examples;

adjusting weights of the encoder machine-trained model and the selected post-processing machine-trained model based on a comparison between tags in the tagged output sequence of items associated with the training example and the labels of the training example; and

repeating said training process until a training objective is achieved.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the training examples include original sequences of items that are given entity-specific labels and instances of supplemental information that lack entity-specific labels.

17. The computer-implemented method of claim 1 , wherein the original sequence of items is obtained from a query submitted via a user computing device.

18. The computer-implemented method of claim 1 , wherein the providing output information comprises:

identifying, using the search system, a target item that matches the tagged output sequence; and

providing information regarding the target item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2021
From: MOJICA DE LA VEGA, LUIS GERARDO; LOU, QIANG; JIAO, JIAN; ZHANG, RUOFEI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 058201/0103 →
Continuity (1)
Related Publication 20230162020A1 · May 25, 2023
References Cited (37)
US 9600764B1 · Rastrow · 2017 [cited by examiner]
US 11042700B1 · Walters · 2021 [cited by examiner]
US 11361151B1 · Guberman · 2022 [cited by examiner]
US 12141732B1 · Solmer · 2024 [cited by examiner]
US 20100256969A1 · Li et al. · 2010 [cited by applicant]
US 20160180247A1 · Li et al. · 2016 [cited by applicant]
US 20170337202A1 · Arya et al. · 2017 [cited by applicant]
US 20200273449A1 · Kumar · 2020 [cited by examiner]
US 20200285916A1 · Wang · 2020 [cited by examiner]
US 20200320988A1 · Rastogi · 2020 [cited by examiner]
US 20200349180A1 · Kempf et al. · 2020 [cited by applicant]
US 20200356592A1 · Yada · 2020 [cited by examiner]
US 20200394455A1 · Lee · 2020 [cited by examiner]
US 20210042366A1 · Hicklin · 2021 [cited by examiner]
US 20210286989A1 · Zhong · 2021 [cited by examiner]
US 20210390392A1 · Lagos · 2021 [cited by examiner]
US 20220012296A1 · Marey · 2022 [cited by examiner]
US 20220138402A1 · Kraus · 2022 [cited by examiner]
US 20220284261A1 · Lillo · 2022 [cited by examiner]
CA 3172730A1 · 2021 [cited by examiner]
CN 111191001A · 2020 [cited by examiner]
CN 109446338B · 2020 [cited by examiner]
CN 113408721A · 2021 [cited by examiner]
CN 109933682B · 2022 [cited by examiner]
CN 113971743A · 2022 [cited by examiner]
CN 109558966B · 2022 [cited by examiner]
CN 113901228B · 2022 [cited by examiner]
CN 116547474A · 2023 [cited by examiner]
CN 111461229B · 2023 [cited by examiner]
WO WO2017124116A1 · 2017 [cited by examiner]
Search Report and Written Opinion for PCT/US2022/041607, mailed Dec. 7, 2022, 18 pages. [cited by applicant]
Alammar, Jay, “The Illustrated Transformer,” available at http://jalammar.github.io/illustrated-transformer/, Github, Jun. 27, 2018, 23 pages. [cited by applicant]
Wu, et al., “Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” in arXiv e-print, arXiv: 1609.08144v2 [cs.CL], Oct. 8, 2016, 23 pages. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in arXiv e-print, arXiv: 1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Liu, et al., “ROBERTa: A Robustly Optimized BERT Pretraining Approach,” in arXiv e-print, arXiv:1907.11692v1 [cs.CL], Jul. 26, 2019, 13 pages. [cited by applicant]
Viswani, et al., “Attention Is All You Need,” in arXiv e-print, arXiv:1706.03762v5 [cs.CL], Dec. 6, 2017, 15 pages. [cited by applicant]
Yadav, e al., “A Survey on Recent Advances in Named Entity Recognition from Deep Learning Models,” in Proceedings of the 27th International Conference on Computational Linguistics, Aug. 2018, pp. 2145-2158. [cited by applicant]