IP Library Granted Patent US 12,505,293
Granted Patent B2
US 12,505,293 · App. 18/180,428 · Granted Dec 23, 2025

Systems and methods for unsupervised neologism normalization of electronic content using embedding space mapping

Inventors: Aasish Pappu (New York, NY); Kapil Thadani (New York, NY); Nasser Zalmout (Abu Dhabi, AE)
Assignee: Yahoo Assets LLC
G06F40/284G06F16/31G06F16/3344G06F40/232G06F40/295G06F40/30G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,293
App. No.
18/180,428
Granted
Dec 23, 2025
Kind
B2
Abstract

Systems and methods are disclosed for utilizing a comment moderation bot for detecting and normalizing neologisms in social media. One method comprises transmitting, by a neologism normalization system, a comment moderation bot for detecting neologisms on an online platform maintained by one or more publisher systems. The comment moderation bot may aggregate data related to user comments and transmit the aggregated data to the neologism normalization system for further processing. The neologism normalization system implements unsupervised machine learning models for detecting neologisms in the aggregated data through tokenization and filtering; and normalizing the neologisms through similarity analysis and lattice decoding.

Claims (43)

1 . A system for neologism normalization, the system comprising:

a data storage device storing instructions and a processor configured to execute the instructions for:

detecting user generated comments on an electronic platform;

removing known lexemes of the user generated comments from further analysis;

normalizing, using a machine learning model, the remaining lexemes through lattice decoding by iteratively replacing frequently used pairs of bytes in a sequence with single unused bytes; and

storing the normalized lexemes in a database.

2 . The system of claim 1 further comprising tokenizing the user generated comments, wherein tokenizing includes executing code for implementing one or more of: creating white space splits, identifying Uniform Resource Locators, or identifying specific punctuation patterns.

3 . The system of claim 1 , wherein detecting one or more user generated comments on the electronic platform further comprises, scanning content retrieving lexical data for analysis.

4 . The system of claim 1 further comprising: determining a type of content associated with the user generated comments.

5 . The system of claim 1 , wherein removing known lexemes of the user generated comments from further analysis further comprises identifying one or more of social media jargon, named entities, spelling errors, and abbreviations, in the one or more user generated comments.

6 . The system of claim 1 , wherein normalizing the remaining lexemes through lattice decoding further comprises:

retrieving a list of known lexemes from a corpus and comparing the list of known lexemes to the remaining lexemes and assigning a score to the lexemes in the list of known lexemes based on similarity metrics.

7 . The system of claim 6 , further comprising:

wherein comparing the list of known lexemes to the list of remaining lexemes further comprises comparing a type of lexemes in the list of known lexemes to the type of lexemes in the list of lexemes; and

wherein assigning a score to the lexemes in the list of known lexemes comprises analyzing the lexemes in the list of known lexemes for semantic similarity, lexical similarity and phonetic similarity, to lexemes in the list of lexemes.

8 . A computer-implemented method for neologism normalization, the computer-implemented method comprising:

detecting user generated comments on an electronic platform;

removing known lexemes of the user generated comments from further analysis;

normalizing, using a machine learning model, the remaining lexemes through lattice decoding by iteratively replacing frequently used pairs of bytes in a sequence with single unused bytes; and

storing the normalized lexemes in a database.

9 . The computer-implemented method of claim 8 further comprising tokenizing the user generated comments, wherein tokenizing includes executing code for implementing one or more of: creating white space splits, identifying Uniform Resource Locators, or identifying specific punctuation patterns.

10 . The computer-implemented method of claim 8 , wherein detecting one or more user generated comments on the electronic platform further comprises, scanning content retrieving lexical data for analysis.

11 . The computer-implemented method of claim 8 , further comprising:

determining a type of content associated with the user generated comments.

12 . The computer-implemented method of claim 8 , wherein removing known lexemes of the user generated comments from further analysis further comprises identifying one or more of social media jargon, named entities, spelling errors, and abbreviations, in the one or more user generated comments.

13 . The computer-implemented method of claim 8 , wherein normalizing the remaining lexemes through lattice decoding further comprises:

retrieving a list of known lexemes from a corpus and comparing the list of known lexemes to the remaining lexemes and assigning a score to the lexemes in the list of known lexemes based on similarity metrics.

14 . The computer-implemented method of claim 13 , further comprising:

wherein comparing the list of known lexemes to the list of remaining lexemes further comprises comparing a type of lexemes in the list of known lexemes to the type of lexemes in the list of lexemes; and

wherein assigning a score to the lexemes in the list of known lexemes comprises analyzing the lexemes in the list of known lexemes for semantic similarity, lexical similarity and phonetic similarity, to lexemes in the list of lexemes.

15 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for neologism normalization, comprising:

detecting user generated comments on an electronic platform;

removing known lexemes of the user generated comments from further analysis;

normalizing, using a machine learning model, the remaining lexemes through lattice decoding by iteratively replacing frequently used pairs of bytes in a sequence with single unused bytes; and

storing the normalized lexemes in a database.

16 . The non-transitory computer readable medium of claim 15 further comprising tokenizing the user generated comments, wherein tokenizing includes executing code for implementing one or more of: creating white space splits, identifying Uniform Resource Locators, or identifying specific punctuation patterns.

17 . The non-transitory computer readable medium of claim 15 , wherein detecting one or more user generated comments on the electronic platform further comprises, scanning content retrieving lexical data for analysis.

18 . The non-transitory computer readable medium of claim 15 , further comprising: determining a type of content associated with the user generated comments.

19 . The non-transitory computer readable medium of claim 15 , wherein normalizing the remaining lexemes through lattice decoding further comprises:

retrieving a list of known lexemes from a corpus and comparing the list of known lexemes to the remaining lexemes and assigning a score to the lexemes in the list of known lexemes based on similarity metrics.

20 . The non-transitory computer readable medium of claim 19 , further comprising:

wherein comparing the list of known lexemes to the list of remaining lexemes further comprises comparing a type of lexemes in the list of known lexemes to the type of lexemes in the list of lexemes; and

wherein assigning a score to the lexemes in the list of known lexemes comprises analyzing the lexemes in the list of known lexemes for semantic similarity, lexical similarity and phonetic similarity, to lexemes in the list of lexemes.

Assignments (4)
SUPPLEMENTAL PATENT SECURITY AGREEMENT Recorded Sep 17, 2025
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 072915/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2023
From: PAPPU, AASISH; THADANI, KAPIL; ZALMOUT, NASSER
To: OATH INC.
Reel/Frame 062922/0022 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2023
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 063009/0717 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2023
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 063010/0001 →
Continuity (3)
Continuation 17021824 · Sep 15, 2020
Continuation 16175533 · Oct 30, 2018
Related Publication 20230229862A1 · Jul 20, 2023
References Cited (29)
US 6278973B1 · Chung · 2001 [cited by examiner]
US 7788084B2 · Brun et al. · 2010 [cited by applicant]
US 8788263B1 · Richfield · 2014 [cited by applicant]
US 9275041B2 · Ghosh et al. · 2016 [cited by applicant]
US 9563693B2 · Zhang et al. · 2017 [cited by applicant]
US 9715493B2 · Papadopoullos et al. · 2017 [cited by applicant]
US 9805018B1 · Richfield · 2017 [cited by applicant]
US 10002543B2 · Telep · 2018 [cited by examiner]
US 10275407B2 · Ohazulike · 2019 [cited by examiner]
US 10382367B2 · Pappu et al. · 2019 [cited by applicant]
US 10628737B2 · Cohen et al. · 2020 [cited by applicant]
US 10810373B1 · Pappu · 2020 [cited by examiner]
US 11636266B2 · Pappu · 2023 [cited by examiner]
US 20020165716A1 · Mangu · 2002 [cited by examiner]
US 20070299664A1 · Peters · 2007 [cited by examiner]
US 20120089400A1 · Henton · 2012 [cited by applicant]
US 20160307114A1 · Ghosh et al. · 2016 [cited by applicant]
US 20180004718A1 · Pappu et al. · 2018 [cited by applicant]
US 20180285459A1 · Soni et al. · 2018 [cited by applicant]
US 20190027155A1 · Liu · 2019 [cited by examiner]
US 20200334093A1 · Dubey · 2020 [cited by examiner]
US 20220365837A1 · Dubey · 2022 [cited by examiner]
Cabré, Maria Teresa, et al., “Stratégie Pour La Semi-Automatique des Néologismes de Presse” , TTR: Traduction, Terminologie, Redaction, www.erudit.org, vol. 8, No. 2, (1995), pp. 89-100. [cited by applicant]
Cartier, Emmanuel, “Neoveille, A Web Platform for Neologism Tracking.” Proceedings of the Software Demonstrations of the 15th Conference of the European Chapter of the Association for Computational Linguistics., Classiq… [cited by applicant]
Falk, Ingrid, et al., “From Non Word to New Word: Automatically Identifying Neologisms in French Newspapers.” LREC—The 9th Edition of the Language Resources and Evaluation Conference, (2014). 10 pages. [cited by applicant]
Gérard, Christophe et al., “Traitement Automatisé De La Néologie: Pourquoi Et Comment Intégrer L'analyse Thématique.” SHS Web of Conferences, EDP Sciences, vol. 8, (2014), pp. 2627-.2646. [cited by applicant]
Kerremans, Daphné, et al., “The NeoCrawler: Identifying and Retrieving Neologisms From the Internet and Monitoring Ongoing Change.” Current Methods in Historical Semantics, Berlin etc.: De Gruyter Mouton (2012), pp. 59-… [cited by applicant]
Renouf, Antoinette, “Sticking to the Text: A Corpus Linguist's View of Language.” Aslib Proceedings. vol. 45. No. 5. MCB UP Ltd, 1993. [cited by applicant]
Stenetorp, Pontus, Automated Extraction of Swedish Neologisms Using a Temporally Annotated Corpus, Royal Institute of Technology, www.kth.se/csc, (2010), 69 pages. [cited by applicant]