IP Library › Granted Patent US 12,530,626
Granted Patent B2
US 12,530,626 · App. 18/459,030 · Granted Jan 20, 2026

Pre-publication assessment of digital content

Inventor: Mustafa Behan (Berlin, DE)
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,626
App. No.
18/459,030
Granted
Jan 20, 2026
Kind
B2
Abstract

Disclosed is a digital content pre-publication assessment approach using online and centralized features to train a machine learning assessment model implemented to assess the validity of certain digital content prior to its publication. The machine learning assessment model generating pre-publication identification of fraudulent digital content based on a relational technology that identifies, compares, and assesses connections between digital content elements (such as time patterns, sentence structure, proximity, characterization similarities, etc.) to build an intelligent assessment tool that identifies singularities and patterns within sets of digital content elements to identify fraudulent sources of digital content. The model creating an assessment of the authenticity of digital content by screening digital content against peer data records, the screening multiplied according to a plurality of iteration steps. Singularities corresponding to linkages may be weighted according to various known fraudulent characteristics. Digital content may comprise reviews generated on consumer sites and/or publicly accessible commentary sites.

Claims (68)

1 . A computer-implemented machine learning method for pre-publication assessment of fraudulent digital consumer reviews in an electronic content publishing system, the method comprising:

receiving, by a computing system comprising one or more processors and a memory storing executable instructions, pre-publication digital consumer reviews submitted for publication to a network-accessible content platform;

verifying the digital consumer reviews by:

(i) capturing a plurality of relevant dimensions of the digital consumer reviews according to a predefined protocol for digital content capture;

(ii) converting the captured digital content into structured data;

(iii) generating an initial authenticity assessment of the structured data;

(iv) translating the structured data into a plurality of characteristic parameters;

(v) performing an intrinsic authenticity analysis of the digital content based on the characteristic parameters;

(vi) performing a metrical analysis using distance operators of distances between the digital content and previously stored digital content in a database to simulate a network of content after adding the digital content to the network by checking the distances between all datapoints in the network that would be established after adding the new datapoint;

(vii) applying one or more distance operators to the computed distances;

(viii) generating an authenticity decision for the digital content; and

(ix) updating a machine learning system with the authenticity decision for use in subsequent analyses of future digital content;

training, by the computing system, a machine learning assessment model using known peer group data points as training data, the machine learning assessment model configured to perform pre-publication identification of fraudulent consumer reviews;

generating, by the computing system, based on a plurality of data records accessed from an electronic database, a fraud-detection model comprising a set of patterns and relationships among the known peer group data points that are indicative of fraudulent digital content;

retrieving, by the computing system, from the electronic database, a plurality of the known peer group data points;

generating, by the computing system, an assessment analysis of the digital content by screening the digital content against the retrieved peer group data points, the screening being performed over a plurality of iteration steps;

executing, by the computing system, the machine learning assessment model using the known peer group data points as input to produce an assessment of whether the consumer review is fraudulent, wherein the machine learning assessment model comprises a density-based clustering technique defined by a density parameter; and

generating, by the computing system, an indicator marking the digital content as fraudulent in response to the assessment indicating fraudulent content,

wherein the indicator is automatically applied by the content publishing system to prevent publication of the fraudulent digital consumer reviews and to reduce false-positive consumer review blocking, thereby improving accuracy and efficiency of automated consumer review moderation on the network.

2 . The method of claim 1 , wherein the predefined protocol for digital content capture comprises extracting at least one of: textual features, metadata attributes, embedded file signatures, pixel intensity histograms, or compression artifacts.

3 . The method of claim 1 , wherein translating the structured data into the plurality of characteristic parameters comprises computing a feature vector including at least one of: temporal posting patterns, author identity confidence scores, semantic similarity metrics, or cryptographic hash variance.

4 . The method of claim 1 , wherein the one or more distance operators comprise normalization functions that scale computed distances into a probabilistic authenticity score.

5 . The method of claim 4 , wherein the density parameter of the density-based clustering technique is dynamically adjusted based on at least one of: average distance between cluster points, variance in content similarity scores, or number of connected components in the network.

6 . The method of claim 1 , wherein training the machine learning assessment model comprises performing distributed training across a plurality of GPU-enabled computing nodes.

7 . The method of claim 1 , wherein generating the indicator marking the digital content as fraudulent further comprises embedding the indicator as a machine-readable tag in the metadata of the content.

8 . A computing system for pre-publication assessment of fraudulent digital consumer reviews in an electronic publishing system, the system comprising:

one or more processors; and

a memory storing executable instructions that, when executed by the one or more processors, verifies the digital consumer reviews by:

(i) capturing a plurality of relevant dimensions of digital content of the consumer review according to a predefined protocol for digital capture;

(ii) converting the captured digital content into structured data;

(iii) generating an initial authenticity assessment of the structured data;

(iv) translating the structured data into a plurality of characteristic parameters;

(v) performing an intrinsic authenticity analysis of the digital content based on the characteristic parameters;

(vi) performing a metrical analysis using distance operators of distances between the digital content and previously stored digital content in a database to simulate a network of content after adding the digital content to the network by checking the distances between all datapoints in the network that would be established after adding the new datapoint;

(vii) applying one or more distance operators to the computed distances;

(viii) generating an authenticity decision for the digital content; and

(ix) updating a machine learning system with the authenticity decision for use in subsequent analyses of future digital content;

training, by the computing system, a machine learning assessment model using known peer group data points as training data, the machine learning assessment model configured to perform pre-publication identification of fraudulent digital content;

generating, by the computing system, based on a plurality of data records accessed from an electronic database, a fraud-detection model comprising a set of patterns and relationships among the known peer group data points that are indicative of fraudulent digital content;

retrieving, by the computing system, from the electronic database, a plurality of the known peer group data points;

generating, by the computing system, an assessment analysis of the digital content by screening the digital content against the retrieved peer group data points, the screening being performed over a plurality of iteration steps;

executing, by the computing system, the machine learning assessment model using the known peer group data points as input to produce an assessment of whether the digital content is fraudulent, wherein the machine learning assessment model comprises a density-based clustering technique defined by a density parameter; and

generating, by the computing system, an indicator marking the digital content as fraudulent in response to the assessment indicating fraudulent content,

wherein the indicator is automatically applied by the content publishing system to prevent publication of the fraudulent digital consumer reviews and to reduce false-positive content blocking, thereby improving accuracy and efficiency of automated content moderation on the network.

9 . The system of claim 8 , wherein the predefined protocol for digital content capture comprises extracting at least one of: textual features, metadata attributes, embedded file signatures, pixel intensity histograms, or compression artifacts.

10 . The system of claim 9 , wherein the density parameter of the density-based clustering technique is dynamically adjusted based on at least one of: average distance between cluster points, variance in content similarity scores, or number of connected components in the network.

11 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to verify digital consumer reviews by:

(i) capturing a plurality of relevant dimensions of the digital consumer reviews according to a predefined protocol for digital content capture;

(ii) converting the captured digital content into structured data;

(iii) generating an initial authenticity assessment of the structured data;

(iv) translating the structured data into a plurality of characteristic parameters;

(v) performing an intrinsic authenticity analysis of the digital content based on the characteristic parameters;

(vi) performing a metrical analysis using distance operators of distances between the digital content and previously stored digital content in a database to simulate a network of content after adding the digital content to the network by checking the distances between all datapoints in the network that would be established after adding the new datapoint;

(vii) applying one or more distance operators to the computed distances;

(viii) generating an authenticity decision for the digital content; and

(ix) updating a machine learning system with the authenticity decision for use in subsequent analyses of future digital content;

training, by the computing system, a machine learning assessment model using known peer group data points as training data, the machine learning assessment model configured to perform pre-publication identification of fraudulent digital consumer reviews;

generating, by the computing system, based on a plurality of data records accessed from an electronic database, a fraud-detection model comprising a set of patterns and relationships among the known peer group data points that are indicative of fraudulent digital consumer reviews;

retrieving, by the computing system, from the electronic database, a plurality of the known peer group data points;

generating, by the computing system, an assessment analysis of the digital consumer review by screening the digital content against the retrieved peer group data points, the screening being performed over a plurality of iteration steps;

executing, by the computing system, the machine learning assessment model using the known peer group data points as input to produce an assessment of whether the digital consumer review is fraudulent, wherein the machine learning assessment model comprises a density-based clustering technique defined by a density parameter; and

generating, by the computing system, an indicator marking the digital content as fraudulent in response to the assessment indicating fraudulent content,

wherein execution of the instructions further causes the computing system to automatically apply the indicator to prevent publication of fraudulent digital consumer reviews and to reduce false-positive content blocking, thereby improving accuracy and efficiency of automated content moderation on the network.

12 . The non-transitory computer-readable medium of claim 11 , wherein the instructions cause the computing system to compute a feature vector including at least one of: temporal posting patterns, author identity confidence scores, semantic similarity metrics, or cryptographic hash variance.

13 . The non-transitory computer-readable medium of claim 12 , wherein the instructions cause the computing system to dynamically adjust a density parameter based on at least one of: average distance between cluster points, variance in content similarity scores, or number of connected components in the network.

14 . The method of claim 11 , wherein preprocessing of the digital content prior to verification comprises applying at least one of: natural language processing tokenization, language translation into a canonical form, image resolution normalization, or audio frequency spectrum smoothing.

15 . The method of claim 11 , wherein the computing system further comprises a fraud detection dashboard configured to display the authenticity decision and supporting metrics.

16 . The method of claim 11 , wherein the plurality of iteration steps in screening the digital content comprises performing progressive filtering using increasingly strict similarity thresholds.

Continuity (1)
Related Publication 20250077952A1 · Mar 6, 2025
References Cited (15)
US 8095441B2 · Wu et al. · 2012 [cited by applicant]
US 11425608B2 · Wang et al. · 2022 [cited by applicant]
US 11640609B1 · Shoumaker et al. · 2023 [cited by applicant]
US 20120131316A1 · Mitola, III · 2012 [cited by examiner]
US 20190182252A1 · Dandu · 2019 [cited by examiner]
US 20210319240A1 · Demir · 2021 [cited by applicant]
US 20220232029A1 · Wei · 2022 [cited by applicant]
Chandola, Varun, Arindam Banerjee, and Vipin Kumar. “Anomaly detection: A survey.” ACM computing surveys (CSUR) 41, No. 3 (2009): 1-58. (Year: 2009). [cited by examiner]
Checco (Checco, Alessandro, Lorenzo Bracciale, Pierpaolo Loreti, Stephen Pinfield, and Giuseppe Bianchi. “AI-assisted peer review.” Humanities and social sciences communications 8, No. 1 (2021): 1-11.) (Year: 2021). [cited by examiner]
Kołcz, A., and Abdur Chowdhury. “Lexicon randomization for near-duplicate detection with I-Match.” The Journal of Supercomputing 45 (2008): 255-276. (Year: 2008). [cited by examiner]
Jia, Zhi-juan, Liang Zhao, and Na Zhou. “The research on fraud group mining which based on social network analysis.” In 2017 29th Chinese Control and Decision Conference (CCDC), pp. 2387-2391. IEEE, 2017. (Year: 2017). [cited by examiner]
Marchal, Samuel, and Sebastian Szyller. “Detecting organized ecommerce fraud using scalable categorical clustering.” In Proceedings of the 35th annual computer security applications conference, pp. 215-228. 2019. (Year:… [cited by examiner]
Akoglu, Leman, Rishi Chandy, and Christos Faloutsos. “Opinion fraud detection in online reviews by network effects.” In Proceedings of the international AAAI conference on web and social media, vol. 7, No. 1, pp. 2-11. … [cited by examiner]
Heydari, Atefeh, Mohammad ali Tavakoli, Naomie Salim, and Zahra Heydari. “Detection of review spam: A survey.” Expert Systems with Applications 42, No. 7 (2015): 3634-3642. (Year: 2015). [cited by examiner]
Santamaria Ruiz and Guzman, fraud detection model based on the discovery symbolic classification rules extracted from a neural network, MICAI 2010, Part ii, LNAI 6438, 2010, pp. 290-302, Springer-Verlag Berlin Heidelber… [cited by applicant]