IP Library › Granted Patent US 12,591,610
Granted Patent B2
US 12,591,610 · App. 18/667,680 · Granted Mar 31, 2026

Systems and methods for removing non-conforming web text

Inventor: Nitin Kishore Sai Samala (Milpitas, CA)
Assignee: Walmart Apollo, LLC
G06F16/338G06F16/353
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,610
App. No.
18/667,680
Granted
Mar 31, 2026
Kind
B2
Abstract

Systems including one or more processors and one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform operations: determining web text sentiment scores for web texts; creating a ranked list of one or more match words; scoring the one or more match words in the ranked list; labeling, using a generative model, the one or more match words to create labeled training data; training, using the labeled training data, a word-based classifier to identify non-conforming web text; receiving first web text for display on the website; using the word-based classifier, as trained, to identify the first web text as non-conforming web text; and automatically not displaying the first web text on the website when at least the classifier score exceeds the predetermined threshold. Other embodiments are described.

Claims (63)

1 . A system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform operations comprising:

determining web text sentiment scores for web texts;

creating a ranked list of one or more match words in the web texts;

periodically scoring, for a specified period of time, the one or more match words in the ranked list of the one or more match words based on the web text sentiment scores and mention weights associated with the one or more match words for the specified period of time;

labeling, using a generative model, the one or more match words to create labeled training data;

training, using the labeled training data, a word-based classifier to identify non-conforming web text submitted for display on a website, wherein the training of the word-based classifier comprises using embedded images to train an image-based classifier to classify an image into at least one classification by applying a natural language processing algorithm to a vector representation of the image, wherein the at least one classification comprises a non-conforming image;

receiving first web text for display on the website;

using the word-based classifier, as trained, to identify the first web text as non-conforming web text when the word-based classifier determines at least a classifier score for the first web text to exceed a predetermined threshold; and

automatically not displaying the first web text on the website when at least the classifier score exceeds the predetermined threshold.

2 . The system of claim 1 , wherein creating the ranked list of the one or more match words comprises ranking the one or more match words by frequency of use.

3 . The system of claim 1 , wherein creating the ranked list of the one or more match words comprises filtering out at least one web text of the web texts when the respective web text sentiment score of at least one web text is approximately zero.

4 . The system of claim 1 , wherein scoring the one or more match words in the ranked list of the one or more match words comprises:

determining a respective score for each of the one or more match words, wherein the respective score comprises a compound sentiment score weighted by a number of mentions in the one or more web texts.

5 . The system of claim 1 , wherein the operations further comprise:

creating a report covering a predetermined period of time using the one or more match words, as scored, in the ranked list; and

extracting one or more topics from the report covering the predetermined period of time.

6 . The system of claim 5 , wherein extracting the one or more topics comprises:

using a term frequency-inverse document frequency algorithm to extract the one or more topics.

7 . The system of claim 5 , wherein extracting the one or more topics comprises at least one of:

using a non-negative matrix factorization algorithm to extract the one or more topics; or

using a latent dirichlet allocation algorithm to extract the one or more topics.

8 . The system of claim 1 , wherein using the word-based classifier further comprises:

determining a word-based classifier score and an image-based classifier score for the first web text; and

creating the classifier score for the first web text based on the word-based classifier score and the image-based classifier score.

9 . The system of claim 1 , wherein the operations further comprise:

parsing one or more of the web texts, wherein parsing the one or more web texts comprises removing one or more of:

one or more stop words from the web texts;

one or more HTML tags from the web texts;

whitespace from the web texts; or

one or more punctuation marks from the web texts.

10 . The system of claim 1 , wherein the one or more non-transitory computer-readable media stores further computing instructions that, when executed on the one or more processors, cause the one or more processors to create a rule to identify the non-conforming web text, wherein the rule comprises a search query configured to return the non-conforming web text.

11 . A method implemented via execution of computing instructions configured to run at one or more processors and configured to be stored at non-transitory computer-readable media, the method comprising:

determining web text sentiment scores for web texts;

creating a ranked list of one or more match words in the web texts;

periodically scoring, for a specified period of time, the one or more match words in the ranked list of the one or more match words based on the web text sentiment scores and mention weights associated with the one or more match words for the specified period of time;

labeling, using a generative model, the one or more match words to create labeled training data;

training, using the labeled training data, a word-based classifier to identify non-conforming web text submitted for display on a website, wherein the training of the word-based classifier comprises using embedded images to train an image-based classifier to classify an image into at least one classification by applying a natural language processing algorithm to a vector representation of the image, wherein the at least one classification comprises a non-conforming image;

receiving first web text for display on the website;

using the word-based classifier, as trained, to identify the first web text as non-conforming web text when the word-based classifier determines at least a classifier score for the first web text to exceed a predetermined threshold; and

automatically not displaying the first web text on the website when at least the classifier score exceeds the predetermined threshold.

12 . The method of claim 11 , wherein creating the ranked list of the one or more match words comprises ranking the one or more match words by frequency of use.

13 . The method of claim 11 , wherein creating the ranked list of the one or more match words comprises filtering out at least one web text of the web texts when the respective web text sentiment score of at least one web text is approximately zero.

14 . The method of claim 11 , wherein scoring the one or more match words in the ranked list of the one or more match words comprises:

determining a respective score for each of the one or more match words, wherein the respective score comprises a compound sentiment score weighted by a number of mentions in the one or more web texts.

15 . The method of claim 11 further comprising:

creating a report covering a predetermined period of time using the one or more match words, as scored, in the ranked list; and

extracting one or more topics from the report covering the predetermined period of time.

16 . The method of claim 15 , wherein extracting the one or more topics comprises:

using a term frequency-inverse document frequency algorithm to extract the one or more topics.

17 . The method of claim 15 , wherein extracting the one or more topics comprises at least one of:

using a non-negative matrix factorization algorithm to extract the one or more topics; or using a latent dirichlet allocation algorithm to extract the one or more topics.

18 . The method of claim 11 , wherein using the word-based classifier further comprises:

determining a word-based classifier score and an image-based classifier score for the first web text; and

creating the classifier score for the first web text based on the word-based classifier score and the image-based classifier score.

19 . The method of claim 11 further comprising:

parsing one or more of the web texts, wherein parsing the one or more web texts comprises removing one or more of:

one or more stop words from the web texts;

one or more HTML tags from the web texts;

whitespace from the web texts; or

one or more punctuation marks from the web texts.

20 . The method of claim 11 further comprising, creating a rule to identify the non-conforming text, wherein the rule comprises a search query configured to return the non-conforming text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2024
From: SAMALA, NITIN KISHORE SAI
To: WALMART APOLLO, LLC
Reel/Frame 067755/0444 →
Continuity (2)
Continuation 17479993 · Sep 20, 2021
Related Publication 20240303264A1 · Sep 12, 2024
References Cited (19)
US 9105008B2 · Popescu et al. · 2015 [cited by applicant]
US 10176500B1 · Mohan · 2019 [cited by applicant]
US 10489393B1 · Mittal et al. · 2019 [cited by applicant]
US 20100057721A1 · Kasano · 2010 [cited by examiner]
US 20110026829A1 · Bandou · 2011 [cited by examiner]
US 20120136985A1 · Popescu et al. · 2012 [cited by applicant]
US 20120330958A1 · Xu · 2012 [cited by examiner]
US 20140101119A1 · Li et al. · 2014 [cited by applicant]
US 20160162576A1 · Arino · 2016 [cited by applicant]
US 20170249389A1 · Brovinsky et al. · 2017 [cited by applicant]
US 20180060338A1 · DeLuca et al. · 2018 [cited by applicant]
US 20190005020A1 · Gregory et al. · 2019 [cited by applicant]
US 20190065589A1 · Wen · 2019 [cited by examiner]
US 20200134095A1 · Weldemariam et al. · 2020 [cited by applicant]
US 20200242750A1 · Kokkula et al. · 2020 [cited by applicant]
US 20210011961A1 · Guan et al. · 2021 [cited by applicant]
EP 0752676 · 2002 [cited by applicant]
WO 2016035072A2 · 2013 [cited by applicant]
Cambridge Consutlants, “Use of AI in Online Content Moderation,” 2019 Report produced on behalf of Ofcom, 84 pgs. 2019. [cited by applicant]