IP Library › Granted Patent US 12,493,751
Granted Patent B2
US 12,493,751 · App. 18/179,862 · Granted Dec 9, 2025

Natural language processing for bias identification

Inventors: Scott Friedman (Minneapolis, MN); Vasanth Sarathy (Chelmsford, MA); Sara Friedman (Minneapolis, MN)
Assignee: Smart Information Flow Technologies, LLC
G06F40/40G06F40/166G06F40/284G06F40/30G06F40/58G06N3/0442G06N3/08G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,751
App. No.
18/179,862
Granted
Dec 9, 2025
Kind
B2
Abstract

A computing machine accesses text from a record. The computing machine identifies, using a natural language processing engine, an entity mapped to a first span of the text. The first span includes a contiguous sequence of one or more words or subwords in the text. The computing machine determines a bias category for the entity. The bias category is selected from a predefined list of bias categories. The determined bias category for the entity depends on a second span of the text. The second span includes a contiguous sequence of one or more words or subwords in the text. The second span is different from the first span.

Claims (36)

1 . A system comprising:

processing circuitry; and

a memory storing instructions which, when executed by the processing circuitry, cause the processing circuitry to perform operations comprising:

receiving, via a graphical user interface, an entry of text;

identifying, using an entity classifier sub-engine of a natural language processing engine, a first span of the text including a reference to a subject, wherein the first span comprises a subject span;

identifying, using the natural language processing engine, a second span of the text including an attribute of the subject;

determining, based on the second span of the text and using a bias determination engine, a bias category in the text, wherein the second span comprises a bias span; and

providing for display, via the graphical user interface, of an indication of the determined bias category and the second span of the text, wherein determining the bias category comprises computing a vector embedding representative of the first span, wherein the vector embedding depends on the second span, wherein the subject is identified based on the computed vector embedding, wherein the text is in a first natural language, wherein the natural language processing engine is trained in a second natural language, different from the first natural language, wherein training the natural language processing engine leverages zero-shot cross-lingual model transfer from the second natural language to the first natural language, wherein the indication of the determined bias category and second span of the text are displayed in real-time after the entry of the text is received, wherein the indication of the determined bias category comprises an emphasizing of a portion of the text comprising the second span of the text used to determine the bias category and a displaying of text identifying the determined bias category within a sidebar of the graphical user interface.

2 . The system of claim 1 , the operations further comprising:

remove the bias category, the prompt comprising a proposed modification of the text lacking the bias category.

3 . The system of claim 1 , wherein the second span of the text is different from the first span of the text, wherein the second span of the text is contiguous, wherein the first span of the text is contiguous.

4 . The system of claim 1 , wherein the bias determination engine comprises at least one artificial neural network, wherein the bias determination engine leverages a feature vector comprising at least the first span of the text and the second span of the text.

5 . The system of claim 1 , wherein the bias determination engine determines the bias category in the text based on the second span of the text being used in a stigmatizing context.

6 . The system of claim 1 , wherein the entry of the text is entered into a healthcare record, and wherein the subject is a patient.

7 . A method comprising:

receiving, via a graphical user interface, an entry of text;

identifying, using an entity classifier sub-engine of a natural language processing engine, a first span of the text including a reference to a subject, wherein the first span comprises a subject span;

identifying, using the natural language processing engine, a second span of the text including an attribute of the subject;

determining, based on the second span of the text and using a bias determination engine, a bias category in the text, wherein the second span comprises a bias span; and

providing for display, via the graphical user interface, of an indication of the determined bias category and the second span of the text, wherein determining the bias category comprises computing a vector embedding representative of the first span, wherein the vector embedding depends on the second span, wherein the subject is identified based on the computed vector embedding, wherein the text is in a first natural language, wherein the natural language processing engine is trained in a second natural language, different from the first natural language, wherein training the natural language processing engine leverages zero-shot cross-lingual model transfer from the second natural language to the first natural language, wherein the indication of the determined bias category and second span of the text are displayed in real-time after the entry of the text is received, wherein the indication of the determined bias category comprises an emphasizing of a portion of the text comprising the second span of the text used to determine the bias category and a displaying of text identifying the determined bias category within a sidebar of the graphical user interface.

8 . The method of claim 7 , further comprising:

providing for display, via the graphical user interface, of a prompt to modify the text to remove the bias category, the prompt comprising a proposed modification of the text lacking the bias category.

9 . The method of claim 7 , wherein the second span of the text is different from the first span of the text, wherein the second span of the text is contiguous, wherein the first span of the text is contiguous.

10 . The method of claim 7 , wherein the bias determination engine comprises at least one artificial neural network, wherein the bias determination engine leverages a feature vector comprising at least the first span of the text and the second span of the text.

11 . The method of claim 7 , wherein the bias determination engine determines the bias category in the text based on the second span of the text being used in a stigmatizing context.

12 . The method of claim 7 , wherein the entry of the text is entered into a healthcare record, and wherein the subject is a patient.

13 . A non-transitory computer-readable medium storing instructions which, when executed by processing circuitry, cause the processing circuitry to perform operations comprising:

receiving, via a graphical user interface, an entry of text;

identifying, using an entity classifier sub-engine of a natural language processing engine, a first span of the text including a reference to a subject, wherein the first span comprises a subject span;

identifying, using the natural language processing engine, a second span of the text including an attribute of the subject;

determining, based on the second span of the text and using a bias determination engine, a bias category in the text, wherein the second span comprises a bias span; and

providing for display, via the graphical user interface, of an indication of the determined bias category and the second span of the text, wherein determining the bias category comprises computing a vector embedding representative of the first span, wherein the vector embedding depends on the second span, wherein the subject is identified based on the computed vector embedding, wherein the text is in a first natural language, wherein the natural language processing engine is trained in a second natural language, different from the first natural language, wherein training the natural language processing engine leverages zero-shot cross-lingual model transfer from the second natural language to the first natural language, wherein the indication of the determined bias category and second span of the text are displayed in real-time after the entry of the text is received, wherein the indication of the determined bias category comprises an emphasizing of a portion of the text comprising the second span of the text used to determine the bias category and a displaying of text identifying the determined bias category within a sidebar of the graphical user interface.

14 . The non-transitory computer-readable medium of claim 13 , the operations further comprising:

providing for display, via the graphical user interface, of a prompt to modify the text to remove the bias category, the prompt comprising a proposed modification of the text lacking the bias category.

15 . The non-transitory computer-readable medium of claim 13 , wherein the second span of the text is different from the first span of the text, wherein the second span of the text is contiguous, wherein the first span of the text is contiguous.

16 . The non-transitory computer-readable medium of claim 13 , wherein the bias determination engine comprises at least one artificial neural network, wherein the bias determination engine leverages a feature vector comprising at least the first span of the text and the second span of the text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: FRIEDMAN, SCOTT, DR.; SARATHY, VASANTH, DR.; FRIEDMAN, SARA, DR.
To: SMART INFORMATION FLOW TECHNOLOGIES, LLC
Reel/Frame 062909/0661 →
Continuity (2)
Provisional Application 63325914 · Mar 31, 2022
Related Publication 20230315994A1 · Oct 5, 2023
References Cited (61)
US 10671942B1 · McGovern et al. · 2020 [cited by applicant]
US 11048741B2 · Cintas et al. · 2021 [cited by applicant]
US 20180246873A1 · Latapie · 2018 [cited by examiner]
US 20200143247A1 · Jonnalagadda et al. · 2020 [cited by applicant]
US 20200210695A1 · Walters et al. · 2020 [cited by applicant]
US 20200250264A1 · Bhide et al. · 2020 [cited by applicant]
US 20210056263A1 · Xia · 2021 [cited by examiner]
US 20210365773A1 · Subramanian et al. · 2021 [cited by applicant]
US 20210406706A1 · Hasan et al. · 2021 [cited by applicant]
US 20220075945A1 · Zhang et al. · 2022 [cited by applicant]
US 20220180068A1 · Sahayaraj et al. · 2022 [cited by applicant]
US 20220180991A1 · Peters · 2022 [cited by applicant]
US 20230067628A1 · Carter et al. · 2023 [cited by applicant]
US 20230153687A1 · Vu et al. · 2023 [cited by applicant]
US 20230342558A1 · Wang et al. · 2023 [cited by applicant]
US 20240126995A1 · Ahmed · 2024 [cited by examiner]
CN 114218947A · 2022 [cited by applicant]
Alipourfard, N., et al. (2021). Systematizing Confidence in Open Research and Evidence (SCORE). SocArXiv. [cited by applicant]
Allen, J., de Beaumont, W., Galescu, L., & Teng, C. M. (2015). Complex event extraction using drum. Technical report, Florida Institute for Human and Machine Cognition Pensacola United States. [cited by applicant]
Aziato, L., Odai, P. N., & Omenyo, C. N. (2016). Religious beliefs and practices in pregnancy and labour: an inductive qualitative study among post-partum women in ghana. BMC pregnancy and childbirth, 16, 1-10. [cited by applicant]
Beltagy, I., Lo, K., & Cohan, A. (2019). Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676. [cited by applicant]
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., & Kalai, A. T. (2016). Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29, 43… [cited by applicant]
Das, D., Schneider, N., Chen, D., & Smith, N. A. (2010). Probabilistic frame-semantic parsing. Human language technologies: The 2010 annual conference of the North American chapter of the association for computational i… [cited by applicant]
Dennett, D. C. (1989). The intentional stance. MIT press. [cited by applicant]
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019 (pp. 4171-4186). Minneapolis, Minnesota: Associa… [cited by applicant]
Eberts, M., & Ulges, A. (2020). Span-based joint entity and relation extraction with transformer pre-training. 24th European Conference on Artificial Intelligence. [cited by applicant]
Fellbaum, C. (2010). Wordnet. In Theory and applications of ontology: computer applications, 231-243. Springer. [cited by applicant]
Forbus, K. D. (1984). Qualitative process theory. Artificial Intelligence, 24, 85-168. [cited by applicant]
Forbus, K. D. (2019). Qualitative representations: How people reason and learn about the continuous world. MIT Press. [cited by applicant]
Friedman, S., Schmer-Galunder, S., Chen, A., & Rye, J. (2019). Relating word embedding gender biases to gender gaps: a cross-cultural analysis. Proceedings of the First Workshop on Gender Bias in Natural Language Proces… [cited by applicant]
Friedman, S. E., Magnusson, I. H., Schmer-Galunder, S. M., Wheelock, R., Gottlieb, J., Patel, P., & Miller, C. (2021). Toward Transformer-Based NLP for Extracting Psychosocial Indicators of Moral Disengagement. Annual M… [cited by applicant]
Garg, N., Schiebinger, L., Jurafsky, D., & Zou, J. (2018). Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences, 115, E3635-E3644. [cited by applicant]
Gelman, B., Clark, C., Friedman, S. E., Kuter, U., & Gentile, J. E. (2021). Toward a robust method for understanding the replicability of research. AAAI Workshop on Scientific Document Understanding. [cited by applicant]
Kuipers, B. (1986). Qualitative simulation. Artificial Intelligence, 29, 289-338. [cited by applicant]
Lee, K., He, L., Lewis, M., & Zettlemoyer, L. (2017). End-to-end neural coreference resolution. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (pp. 188-197). Copenhagen, Denmark: … [cited by applicant]
Lombrozo, T., & Carey, S. (2006). Functional explanation and the function of explanation. Cognition, 99, 167-204. [cited by applicant]
Loureiro, D., & Jorge, A. (2019). Language modelling makes sense: Propagating representations through WordNet for full-coverage word sense disambiguation. Proceedings of the 57th Annual Meeting of the Association for Co… [cited by applicant]
Magnusson, I. H., & Friedman, S. E. (2021). Extracting fine-grained knowledge graphs of scientific claims: Dataset and transformer-based results. Proceedings of the 2021 Conference on Empirical Methods in Natural Langua… [cited by applicant]
Mueller, R., & Abdullaev, S. (2019). Deepcause: Hypothesis extraction from information systems papers with deep earning for theory ontology learning. Proceedings of the 52nd Hawaii Inter-national Conference on System Sc… [cited by applicant]
Pustejovsky, J. (1991). The syntax of event structure. Cognition, 41, 47-81. [cited by applicant]
Wang, L. L., et al. (2020). CORD-19: The COVID-19 Open Research Dataset. [cited by applicant]
Astor, M. (2019). How the politically unthinkable can become mainstream. New York Times, 26. [cited by applicant]
Bandura, A. (1999). Moral disengagement in the perpetration of inhumanities. Personality and social psychology review, 3(3), 193-209. [cited by applicant]
Bandura, A. (2002). Selective moral disengagement in the exercise of moral agency. Journal of moral education, 31(2), 101-119. [cited by applicant]
Bandura, A. (2016). Moral disengagement: How people do harm and live with themselves. Worth publishers. [cited by applicant]
Ging, D. (2019). Alphas, betas, and incels: Theorizing the masculinities of the manosphere. Men and Masculinities, 22(4), 638-657. [cited by applicant]
Hoffman, B., Ware, J., & Shapiro, E. (2020). Assessing the threat of incel violence. Studies in Conflict & Terrorism, 43(7), 565-587. [cited by applicant]
Hoover, J., Portillo-Wightman, G., Yeh, L., Havaldar, S., Davani, A. M., Lin, Y., . . . others (2020). Moral foundations twitter corpus: a collection of 35k tweets annotated for moral sentiment. Social Psychological and… [cited by applicant]
Kennedy, B., Atari, M., Davani, A. M., Yeh, L., Omrani, A., Kim, Y., . . . others (2018). The gab hate corpus: a collection of 27k posts annotated for hate speech. [cited by applicant]
Magnusson, I. H., & Friedman, S. E. (2021). Graph knowledge extraction of causal, comparative, predictive, and proportional associations in scientific claims with a transformer-based model. In AAAI Workshop on Scientifi… [cited by applicant]
Mendelsohn, J., Tsvetkov, Y., & Jurafsky, D. (2020). A framework for the computational linguistic analysis of dehumanization. Frontiers in Artificial Intelligence, 3, 55. [cited by applicant]
Olteanu, A., Castillo, C., Boy, J., & Varshney, K. (2018). The effect of extremist violence on hateful speech online. In Proceedings of the International AAAI Conference on Web and Social Media (vol. 12). [cited by applicant]
Peters, J., Grynbaum, M., Collins, K., Harris, R., & Taylor, R. (2019). How the El Paso Killer Echoed the Incendiary Words of Conservative Media Stars. The New York Times. [cited by applicant]
Saidon, I., Galbreath, J., & Whiteley, A. (2010). Antecedents of moral disengagement: Preliminary empirical study in malaysia. In Proceedings of the 24th annual australian and new zealand academy of management conferenc… [cited by applicant]
Soral, W., Bilewicz, M., & Winiewski, M. (2018). Exposure to hate speech increases prejudice through desensitization. Aggressive behavior, 44(2), 136-146. [cited by applicant]
Turner, R. M. (2008). Moral disengagement as a predictor of bullying and aggression: Are there gender differences? The University of Nebraska-Lincoln. [cited by applicant]
Van Bavel, J. J., & Pereira, A. (2018). The partisan brain: an identity-based model of political belief. Trends in cognitive sciences, 22(3), 213-224. [cited by applicant]
Voigt, R., Camp, N. P., Prabhakaran, V., Hamilton, W. L., Hetey, R. C., Griffiths, C. M., . . . Eberhardt, J. L. (2017). Language from police body camera footage shows racial disparities in officer respect. Proceedings … [cited by applicant]
Non-Final Office Action mailed on Feb. 20, 2025 for U.S. Appl. No. 18/179,842, 39 pp. [cited by applicant]
Final Office Action mailed on Jun. 10, 2025 for U.S. Appl. No. 18/179,842, 38 pp. [cited by applicant]
Non-Final Office Action mailed on Oct. 21, 2025 for U.S. Appl. No. 18/179,842, 64 pp. [cited by applicant]