IP Library › Granted Patent US 12,001,795
Granted Patent B2
US 12,001,795 · App. 17/399,574 · Granted Jun 4, 2024

Extractive method for speaker identification in texts with self-training

Inventors: Dian Yu (Bellevue, WA); Dong Yu (Bothell, WA)
Assignee: TENCENT AMERICA LLC
G06F40/284G06F40/169G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,001,795
App. No.
17/399,574
Granted
Jun 4, 2024
Kind
B2
Abstract

A method, computer program, and computer system is provided for identifying a speaker in at text based work. Labeled and unlabeled instances corresponding to one or more speakers are extracted. Pseudo-labels are inferred for the extracted unlabeled instances based on the labeled instances. One or more of the unlabeled instances are labeled based on the inferred pseudo-labels.

Claims (35)

1. A method of identifying a speaker in a text-based work, executable by a processor, comprising:

extracting labeled instances and unlabeled instances corresponding to one or more speakers;

inferring pseudo-labels for the extracted unlabeled instances based on the labeled instances; and

labeling one or more of the unlabeled instances based on the inferred pseudo-labels,

wherein the labeled and unlabeled instances correspond to a class token, tokens in a first piece of text containing an utterance, a separator token, and tokens in a second piece of text that covers the first piece of text,

wherein constructing an input sequence comprises concatenating the class token, the tokens in the first piece of text containing the utterance, the separator token, and the tokens in the second piece of text that covers the first piece of text, and

wherein two vectors correspond to estimated probabilities of each of the tokens being a starting token or an ending token of an answer span that appears in the second piece of text.

2. The method of claim 1 , further comprising training a first model based on the labeled instances.

3. The method of claim 2 , further comprising training a second model based on the inferred pseudo-labels and the labeled instances.

4. The method of claim 3 , further comprising replacing the first model with the second model.

5. The method of claim 1 , wherein the labeled and unlabeled instances correspond to a speaker or a quotation.

6. A computer system for identifying a speaker in a text-based work, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

extracting code configured to cause the one or more computer processors to extract labeled instances and unlabeled instances corresponding to one or more speakers;

inferring code configured to cause the one or more computer processors to infer pseudo-labels for the extracted unlabeled instances based on the labeled instances; and

labeling code configured to cause the one or more computer processors to label one or more of the unlabeled instances based on the inferred pseudo-labels,

wherein the labeled and unlabeled instances correspond to a class token, tokens in a first piece of text containing an utterance, a separator token, and tokens in a second piece of text that covers the first piece of text,

wherein constructing an input sequence comprises concatenating the class token, the tokens in the first piece of text containing the utterance, the separator token, and the tokens in the second piece of text that covers the first piece of text, and

wherein two vectors correspond to estimated probabilities of each of the tokens being a starting token or an ending token of an answer span that appears in the second piece of text.

7. The computer system of claim 6 , further comprising training code configured to cause the one or more computer processors to train a first model based on the labeled instances.

8. The computer system of claim 7 , further comprising training code configured to cause the one or more computer processors to train a second model based on the inferred pseudo-labels and the labeled instances.

9. The computer system of claim 8 , further comprising replacing code configured to cause the one or more computer processors to replace the first model with the second model.

10. The computer system of claim 6 , wherein the labeled and unlabeled instances correspond to a speaker or a quotation.

11. A non-transitory computer readable medium having stored thereon a computer program for identifying a speaker in a text-based work, the computer program configured to cause one or more computer processors to:

extract labeled instances and unlabeled instances corresponding to one or more speakers;

infer pseudo-labels for the extracted unlabeled instances based on the labeled instances; and

label one or more of the unlabeled instances based on the inferred pseudo-labels,

wherein the labeled and unlabeled instances correspond to a class token, tokens in a first piece of text containing an utterance, a separator token, and tokens in a second piece of text that covers the first piece of text,

wherein constructing an input sequence comprises concatenating the class token, the tokens in the first piece of text containing the utterance, the separator token, and the tokens in the second piece of text that covers the first piece of text, and

wherein two vectors correspond to estimated probabilities of each of the tokens being a starting token or an ending token of an answer span that appears in the second piece of text.

12. The non-transitory computer readable medium of claim 11 , wherein the computer program is further configured to cause one or more computer processors to train a first model based on the labeled instances.

13. The non-transitory computer readable medium of claim 12 , wherein the computer program is further configured to cause one or more computer processors to train a second model based on the inferred pseudo-labels and the labeled instances.

14. The non-transitory computer readable medium of claim 13 , wherein the computer program is further configured to cause one or more computer processors to replace the first model with the second model.

15. The non-transitory computer readable medium of claim 11 , wherein the labeled and unlabeled instances correspond to a speaker or a quotation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2021
From: YU, DIAN; YU, DONG
To: TENCENT AMERICA LLC
Reel/Frame 057149/0605 →
Continuity (1)
Related Publication 20230053148A1 · Feb 16, 2023