IP Library › Granted Patent US 12,619,872
Granted Patent B2
US 12,619,872 · App. 17/808,242 · Granted May 5, 2026

System and method for filtering datasets using conditional-likelihood filtration

Inventors: Helen Ngo (Toronto, CA); Nicholas Frosst (Toronto, CA)
Assignee: Cohere Inc.
G06N3/08G06F40/289G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,872
App. No.
17/808,242
Granted
May 5, 2026
Kind
B2
Abstract

A system and method are provided for generating a trained model to filter data sets for filtering hate speech. The method includes obtaining an unfiltered corpus of data, obtaining a set of trigger phrases, and using the set of trigger phrases to generate a trained model which comprises at least one conditional likelihood of the trigger phrases conditioned on documents in the corpus of data. A system and method are also provided for filtering data sets for hate speech using pre-trained models. The method includes obtaining a pretrained model generated using a set of trigger phrases and which comprises at least one conditional likelihood of the trigger phrases conditioned on document in a corpus of data used to generate the pretrained model; using the pretrained model to filter an unfiltered dataset and generate a filtered dataset; and outputting the filtered dataset.

Claims (53)

1 . A method of generating a trained natural language processing model having a lower propensity of generating hate speech, the method comprising:

obtaining an unfiltered dataset;

obtaining a set of trigger phrases;

determining, for each trigger phrase, a conditional likelihood of portions of the unfiltered dataset satisfying a threshold under a probability distribution for at least one trigger axis, wherein the conditional likelihood for each trigger phrase is determined based on a context provided by the respective portions of the unfiltered dataset;

filtering the unfiltered dataset to remove the one or more portions having at least one determined conditional likelihood that satisfies the threshold to generate a filtered dataset;

generating the trained natural language processing model using the filtered dataset;

providing the trained natural language processing model to an application; and

utilizing the trained natural language processing model in the application thereby reducing the conditional likelihood that the trained natural language processing model generates subject matter defined by the set of trigger phrases when generating a result from the utilizing.

2 . The method of claim 1 , further comprising obtaining a baseline model trained on the unfiltered dataset, wherein the probability distribution is represented by the baseline model and wherein the natural language processing model is generated by training the baseline model by finetuning the baseline model using the filtered dataset.

3 . The method of claim 1 , further comprising:

generating a text phrase with the trained natural language processing model, the trained natural language processing model generating the text phrase at least in part based on received input.

4 . The method of claim 1 , further comprising:

updating the set of trigger phrases;

filtering the filtered dataset based on the updated trigger phrases to generate a further filtered dataset; and

further training the trained natural language processing model with the further filtered dataset.

5 . The method of claim 1 , wherein one or more of the set of trigger phrases is appended to each portion of the unfiltered data.

6 . The method of claim 1 , wherein the set of trigger phrases are responsive to a first type of hate speech, and the threshold is responsive to a second type of hate speech.

7 . The method of claim 1 , wherein the threshold is based on a comparison of determined conditional likelihoods of entries of the dataset to properties of the unfiltered dataset.

8 . The method of claim 1 , wherein the at least one conditional likelihood is generated based on the relationship p(t|d) where t is a trigger phrase and d is an extract from the beginning of a document.

9 . The method of claim 1 , further comprising removing portions of the unfiltered dataset which include one or more blocked phrases, and the set of trigger terms do not include the one or more blocked phrases.

10 . The method of claim 1 , wherein the threshold includes a plurality of thresholds associated with different types of hate speech.

11 . A system for generating a trained natural language processing model having a lower propensity of generating hate speech, the system comprising:

a processor;

a memory in communication with the processor, the memory comprising computer executable instructions that when executed by the processor cause the processor to:

obtain an unfiltered dataset;

obtain a set of trigger phrases;

determine, for each trigger phrase, a conditional likelihood of portions of the unfiltered dataset satisfying a threshold under a probability distribution for at least one trigger axis, wherein the conditional likelihood for each trigger phrase is determined based on a context provided by the respective portions of the unfiltered dataset;

filter the unfiltered dataset to remove the one or more portions having at least one determined conditional likelihood that satisfies the threshold to generate a filtered dataset;

generate the trained natural language processing model using the filtered dataset;

provide the trained natural language processing model to an application; and

utilize the trained natural language processing model in the application thereby reducing the conditional likelihood that the trained natural language processing model generates subject matter defined by the set of trigger phrases when generating a result from the utilizing.

12 . The system of claim 11 , further comprising instructions to obtain a baseline model trained on the unfiltered dataset, wherein the probability distribution is represented by the baseline model and wherein the natural language processing model is generated by training the baseline model by finetuning the baseline model using the filtered dataset.

13 . The system of claim 11 , wherein the instructions cause the processor to:

generate a text phrase with the trained natural language processing model, the trained natural language processing model generating the text phrase at least in part based on received input.

14 . The system of claim 11 , wherein the instructions cause the processor to:

update the set of trigger phrases;

filter the filtered dataset based on the updated trigger phrases to generate a further filtered dataset; and

further train the trained natural language processing model with the further filtered dataset.

15 . The system of claim 11 , wherein one or more of the set of trigger phrases is appended to each portion of the unfiltered data.

16 . The system of claim 11 , wherein the set of trigger phrases are responsive to a first type of hate speech, and the threshold is responsive to a second type of hate speech.

17 . The system of claim 11 , wherein the threshold is based on a comparison of determined conditional likelihoods of entries of the dataset to properties of the unfiltered dataset.

18 . The system of claim 11 , wherein the instructions cause the processor to:

remove portions of the unfiltered dataset which include one or more blocked phrases, and

wherein the set of trigger terms do not include the one or more blocked phrases.

19 . The system of claim 11 , wherein the threshold includes a plurality of thresholds associated with different types of hate speech.

20 . A non-transitory computer readable medium for training a neural network model including a first plurality of nodes, the computer readable medium comprising computer executable instructions to:

obtain an unfiltered dataset;

obtain a set of trigger phrases;

determine, for each trigger phrase, a conditional likelihood of portions of the unfiltered dataset satisfying a threshold under a probability distribution for at least one trigger axis, wherein the conditional likelihood for each trigger phrase is determined based on a context provided by the respective portions of the unfiltered dataset;

filter the unfiltered dataset to remove the one or more portions having at least one determined conditional likelihood that satisfies the threshold to generate a filtered dataset;

generate the trained natural language processing model using the filtered dataset;

provide the trained natural language processing model to an application; and

utilize the trained natural language processing model in the application thereby reducing the conditional likelihood that the trained natural language processing model generates subject matter defined by the set of trigger phrases when generating a result from the utilizing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2022
From: NGO, HELEN; FROSST, NICHOLAS
To: COHERE INC.
Reel/Frame 060278/0896 →
Continuity (2)
Provisional Application 63202785 · Jun 24, 2021
Related Publication 20220414467A1 · Dec 29, 2022
References Cited (17)
US 10885279B2 · Barachha · 2021 [cited by examiner]
US 20150163184A1 · Kanter · 2015 [cited by examiner]
US 20160336006A1 · Levit · 2016 [cited by examiner]
US 20200125928A1 · Doyle · 2020 [cited by examiner]
US 20210097405A1 · McNeil et al. · 2021 [cited by applicant]
US 20220229991A1 · Duong · 2022 [cited by examiner]
US 20220253447A1 · Boytsov · 2022 [cited by examiner]
US 20220284884A1 · Tongya · 2022 [cited by examiner]
Combined Search Report issued in related GB application No. 2209258.9, dated Apr. 8, 2024. [cited by applicant]
Vidgen B; Derczynski, L.; Directions in abusive language training data, a systematic review: Garbage in, garbage out. PLOS ONE 15(12): e0243300; https://doi.org/10.1371/journal.pone.0243300 (2020). [cited by applicant]
Raffel, Colin; Shazeer, Noam; Roberts, Adam; Narang, Katherine Lee, Sharan; Matena, Michael; Zhou, Yanqi; Li, Wei; and Liu, Peter J.; Exploring the limits of transfer learning with a unified text-totext transformer. Jou… [cited by applicant]
Dodge, Jesse; Sap, Maarten, Marasovic, Ana; Agnew, William; Ilharco, Gabriel; Groeneveld, Dirk; Gardner, Matt; Documenting the English Colossal Clean Crawled Corpus; https://doi.org/10.48550/arXiv.2104.08758; (Apr. 18, … [cited by applicant]
Einstein, Albert; Zur Elektrodynamik bewegter Korper. (German) [On the electrodynamics of moving bodies]. Annalen der Physik, 322(10):891-921, 1905; Translation: http://hermes.ffn.ub.es/luisnavarro/nuevo_maletin/Einstei… [cited by applicant]
Kennedy, B., Atari, M., Davani, A. M., Yeh, L., Omrani, A., Kim, Y.; Dehghani, M.; The Gab Hate Corpus: A collection of 27k posts annotated for hate speech. (Jul. 18, 2018) https://doi.org/10.31234/osf.io/hqixn (Jul. 18… [cited by applicant]
Gehman, Sam; Gururangan, Suchin; Sap, Maarten; Choi, Yejin & Smith, Noah A.; RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. Findings of EMNLP; https://doi.org/10.48550/arXiv.2009.11462 (20… [cited by applicant]
Dodge, Jesse et al.; Documenting the English Colossal Clean Crawled Corpus. https://arxiv.org/abs/2104.08758 (Apr. 18, 2021). [cited by applicant]
Hookr, Sara; https://twitter.com/sarahookr/status/1361373527861915648 (Feb. 15, 2021). [cited by applicant]