IP Library › Granted Patent US 12,323,262
Granted Patent B2
US 12,323,262 · App. 17/806,751 · Granted Jun 3, 2025

Systems and methods for coreference resolution

Inventors: Tuan Manh Lai (Urbana, IL); Trung Huu Bui (San Jose, CA); Doo Soon Kim (San Jose, CA)
Assignee: ADOBE INC.
H04L12/1831G06F40/284G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,262
App. No.
17/806,751
Filed
Jun 14, 2022
Granted
Jun 3, 2025
Kind
B2
Art Unit
2442
USPC
709/204
Abstract

Systems and methods for coreference resolution are provided. One aspect of the systems and methods includes inserting a speaker tag into a transcript, wherein the speaker tag indicates that a name in the transcript corresponds to a speaker of a portion of the transcript; encoding a plurality of candidate spans from the transcript based at least in part on the speaker tag to obtain a plurality of span vectors; extracting a plurality of entity mentions from the transcript based on the plurality of span vectors, wherein each of the plurality of entity mentions corresponds to one of the plurality of candidate spans; and generating coreference information for the transcript based on the plurality of entity mentions, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity.

Claims (64)

1. A method for coreference resolution, comprising:

obtaining a transcript including a first span of text separated from a second span of text;

determining that the first span of text comprises a name of a speaker of the second span of text;

inserting a speaker tag into the obtained transcript based on the determination, wherein the speaker tag comprises an opening tag inserted before the first span of text of the obtained transcript and a closing tag inserted after the first span of text of the obtained transcript, and wherein the speaker tag indicates that the first span of text comprises the name of the speaker;

encoding a plurality of candidate spans from the transcript using an encoder network of a machine learning model to obtain a plurality of span vectors, wherein the speaker tag is included in at least one candidate span of the plurality of candidate spans;

extracting a plurality of entity mentions from the transcript based on the plurality of span vectors using a mention extractor network of the machine learning model, wherein each of the plurality of entity mentions corresponds to one of the plurality of candidate spans; and

generating coreference information for the transcript based on the plurality of entity mentions using a mention linker network of the machine learning model, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity.

2. The method of claim 1 , further comprising:

identifying a threshold span length; and

selecting each span in the transcript that is less than the threshold span length to obtain the plurality of candidate spans.

3. The method of claim 1 , further comprising:

encoding individual tokens of the transcript including the speaker tag to obtain a plurality of encoded tokens; and

identifying a starting token and an end token for each of the plurality of candidate spans, wherein a span vector corresponding to each of the plurality of candidate spans includes the starting token and the end token.

4. The method of claim 3 , further comprising:

generating an attention vector based on a subset of the encoded tokens corresponding to each of the plurality of candidate spans, wherein the span vector includes the attention vector.

5. The method of claim 1 , further comprising:

generating a mention score for each of the plurality of candidate spans based on a corresponding span vector from the plurality of span vectors;

identifying a mention score threshold; and

determining that each of the plurality of entity mentions has a mention score that exceeds the mention score threshold, wherein the plurality of entity mentions are extracted based on the determination.

6. The method of claim 1 , further comprising:

identifying the pair of candidate spans from the plurality of candidate spans;

combining a pair of span vectors of the plurality of span vectors corresponding to the pair of candidate spans to obtain a span pair vector; and

applying a mention linker network to the span pair vector to obtain a similarity score for the pair of candidate spans, wherein the coreference information is based on the similarity score.

7. The method of claim 6 , further comprising:

combining the similarity score with mention scores for each of the pair of candidate spans to obtain a coreference score, wherein the coreference information includes the coreference score.

8. The method of claim 6 , further comprising:

computing a product of the pair of span vectors, wherein the span pair vector includes the pair of span vectors and the product of the pair of span vectors.

9. A method for coreference resolution, comprising:

obtaining training data comprising training text, mention annotation data, and coreference annotation data, wherein the training text includes a first span of text separated from a second span of text;

determining that the first span of text comprises a name of a speaker of the second span of text;

inserting a speaker tag into the obtained training text based on the determination, wherein the speaker tag comprises an opening tag inserted before the first span of text of the obtained training text and a closing text inserted after the first span of text of the obtained training text, and wherein the speaker tag indicates that the first span of text comprises the name of the speaker;

encoding a plurality of candidate spans from the training text using an encoder network of a machine learning model to obtain a plurality of span vectors, wherein the speaker tag is included in at least one candidate span of the plurality of candidate spans;

extracting a plurality of entity mentions from the training text based on the plurality of span vectors using a mention extractor network of the machine learning model, wherein each of the plurality of entity mentions corresponds to one of the plurality of candidate spans;

updating parameters of the mention extractor network in a first training phase based on the plurality of entity mentions and the mention annotation data;

extracting an updated plurality of entity mentions from the training text based on the plurality of span vectors using the mention extractor network with the updated parameters;

generating coreference information based on the updated plurality of entity mentions using a mention linker network of the machine learning model, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity; and

updating the mention linker network in a second training phase based on the coreference information and the coreference annotation data.

10. The method of claim 9 , further comprising:

updating the parameters of the mention extractor network in the second training phase based on the coreference information and the coreference annotation data.

11. The method of claim 9 , further comprising:

generating a mention score for each of the plurality of candidate spans based on a corresponding span vector from the plurality of span vectors using the mention extractor network;

computing a detection score for each of the plurality of candidate spans based on the mention score and a binary value indicating whether the candidate span is included in the mention annotation data; and

computing a detection loss based on the detection score, wherein the parameters of the mention extractor network are updated based on the detection loss in the first training phase.

12. The method of claim 9 , further comprising:

identifying an antecedent for an entity mention of the plurality of entity mentions based on the coreference annotation data;

identifying a probability of the antecedent for the entity mention based on the coreference information; and

computing an objective function based on the probability, wherein the parameters of the mention linker network are updated to optimize the objective function.

13. A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device configured to perform operations comprising:

obtaining a transcript including a first span of text separated from a second span of text;

determining that the first span of text comprises a name of a speaker of the second span of text;

inserting a speaker tag into the obtained transcript based on the determination, wherein the speaker tag comprises an opening tag inserted before the first span of text of the obtained transcript and a closing tag inserted after the first span of text of the obtained transcript, and wherein the speaker tag indicates that the first span of text comprises the name of the speaker;

encoding a plurality of candidate spans from the transcript using an encoder network of a machine learning model to obtain a plurality of span vectors, wherein the speaker tag is included in at least one candidate span of the plurality of candidate spans;

extracting a plurality of entity mentions from the transcript based on the plurality of span vectors using a mention extractor network of the machine learning model, wherein the mention extractor network is trained based on mention annotation data in a first training phase and based on coreference annotation data in a second training phase; and

generating coreference information for the transcript based on the plurality of entity mentions using a mention linker network of the machine learning model, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity, and wherein the mention linker network is trained jointly with the mention extractor network on the coreference annotation data in the second training phase.

14. The system of claim 13 , wherein:

the encoder network comprises a transformer network.

15. The system of claim 13 , wherein:

the mention extractor network comprises a feed-forward neural network.

16. The system of claim 13 , wherein:

the mention linker network comprises a feed-forward neural network.

17. The system of claim 13 , further comprising:

a training component configured to update parameters of the mention extractor network and the mention linker network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2022
From: LAI, TUAN MANH; BUI, TRUNG HUU; KIM, DOO SOON
To: ADOBE INC.
Reel/Frame 060189/0048 →
Continuity (1)
Related Publication 20230403175A1 · Dec 14, 2023
References Cited (34)
US 11366966B1 · Ramsey · 2022 [cited by examiner]
US 20150127340A1 · Epshteyn · 2015 [cited by examiner]
US 20210034701A1 · Fei · 2021 [cited by examiner]
US 20210142785A1 · Wang · 2021 [cited by examiner]
US 20220261555A1 · Lebanoff · 2022 [cited by examiner]
US 20220414130A1 · Master Ben-Dor · 2022 [cited by examiner]
1Ng, “Supervised Noun Phrase Coreference Research: The First Fifteen Years”, Association for Computational Linguistics, Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, Jul. 2010,… [cited by applicant]
2Raghunathan, et al., “A Multi-Pass Sieve for Coreference Resolution”, Association for Computational Linguistics, Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, Oct. 2010, pp. 49… [cited by applicant]
3Durrett, et al., “Easy Victories and Uphill Battles in Coreference Resolution”, Association for Computational Linguistics, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 20… [cited by applicant]
4Clark, et al., “Entity-Centric Coreference Resolution with Model Stacking”, Association for Computational Linguistics, Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th… [cited by applicant]
5Wiseman, et al., “Learning Global Features for Coreference Resolution”, arXiv preprint arXiv:1604.03035v1 [cs.CL] Apr. 11, 2016, 11 pages. [cited by applicant]
6Clark, et al., “Deep Reinforcement Learning for Mention-Ranking Coreference Models”, arXiv preprint arXiv:1609.08667v3 [cs.CL] Oct. 31, 2016, 7 pages. [cited by applicant]
7Lee, et al., “End-to-end Neural Coreference Resolution”, arXiv preprint arXiv:1707.07045v2 [cs.CL] Dec. 15, 2017, 10 pages. [cited by applicant]
8Zhang, et al., “Neural Coreference Resolution with Deep Biaffine Attention by Joint Mention Detection and Mention Clustering”, arXiv preprint arXiv:1805.04893v1 [cs.CL] May 13, 2018, 6 pages. [cited by applicant]
9Lee, et al., “Higher-order Coreference Resolution with Coarse-to-fine Inference”, arXiv:1804.05392v1 [cs.CL] Apr. 15, 2018, 6 pages. [cited by applicant]
10Gu, et al, “A Study on Improving End-to-End Neural Coreference Resolution”, in Sun, et al. (eds), Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data, CCL NLP-NABD 2… [cited by applicant]
11Fei, et al., “End-to-end Deep Reinforcement Learning Based Coreference Resolution”, in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, pp. 660-665, 6 pages. [cited by applicant]
12Kantor, et al., “Coreference Resolution with Entity Equalization”, in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, pp. 673-677, 5 pages. [cited by applicant]
13Joshi, et al., “BERT for Coreference Resolution: Baselines and Analysis”, arXiv preprint arXiv:1908.09091v4 [cs.CL] Dec. 22, 2019, 6 pages. [cited by applicant]
14Fang, et al., “Incorporating Structural Information for Better Coreference Resolution”, in Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-19), Aug. 10-16, 2019, pp. 5039-5045,… [cited by applicant]
15Joshi, et al., “SpanBERT: Improving Pre-training by Representing and Predicting Spans”, in Transactions of the Association for Computational Linguistics, 2020, vol. 8, pp. 64-77, Jan. 1, 2020, 14 pages. [cited by applicant]
16Xu, et al., “Revealing the Myth of Higher-Order Inference in Coreference Resolution”, arXiv preprint arXiv:2009.12013v2 [cs.CL] Sep. 28, 2020, 7 pages. [cited by applicant]
17Kirstain, et al., “Coreference Resolution without Span Representations”, arXiv preprint arXiv:2101.00434v2 [cs.CL] May 31, 2021, 6 pages. [cited by applicant]
18Wu, et al., “CorefQA: Coreference Resolution as Query-based Span Prediction”, arXiv preprint arXiv:1911.01746v4 [cs.CL] Jul. 18, 2020, 12 pages. [cited by applicant]
19Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv preprint arXiv:1810.04805v2 [cs.CL] May 24, 2019, 16 pages. [cited by applicant]
20Turian, et al., “Word representations: A simple and general method for semi-supervised learning”, in Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, Jul. 2010, pp. 384-394, 11 … [cited by applicant]
21Pennington, et al., “GloVe: Global Vectors for Word Representation”, in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Oct. 25-29, 2014, pp. 1532-1543, 12 pages. [cited by applicant]
22Bahdanau, et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv preprint arXiv:1409.0473v7 [cs.CL] May 19, 2016, 15 pages. [cited by applicant]
23Pradhan, et al., “CoNLL-2012 Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes”, in Proceedings of the Joint Conference on EMNLP and CoNLL: Shared Task, Jul. 13, 2012, pp. 1-40, 40 pages. [cited by applicant]
24Sanh, et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, arXiv preprint arXiv:1910.01108v4 [cs.CL] Mar. 1, 2020, 5 pages. [cited by applicant]
25Sun, et al., “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices”, arXiv preprint arXiv:2004.02984v2 [cs.CL] Apr. 14, 2020, 13 pages. [cited by applicant]
26Lai, et al., “A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems”, arXiv preprint arXiv:1910.12995v3 [cs.CL] Feb. 9, 2020, 5 pages. [cited by applicant]
27Lai, et al., “A Context-Dependent Gated Module for Incorporating Symbolic Semantics into Event Coreference Resolution”, arXiv preprint arXiv:2104.01697v1 [cs.CL] Apr. 4, 2021, 9 pages. [cited by applicant]
28Wen, et al., “RESIN: A Dockerized Schema-Guided Cross-document Cross-lingual Cross-media Information Extraction and Event Tracking System”, in Proceedings of the 2021 Conference of the North American Chapter of the As… [cited by applicant]