Methods, systems, and media for providing digital advertisers with improved context for dynamic webpages
Methods, systems, and media for providing contextual information associated with webpages are provided. In some embodiments, the method comprises: receiving a plurality of risk tolerance values from a first advertiser; accessing a first webpage through a first universal resource locator (URL), wherein the first webpage contains at least one dynamic advertising region; determining, using a machine learning model, (i) that the first URL is a cybersquatting attempt of a second URL based on a comparison of a first domain name associated with the first URL and a second domain name associated with the second URL; (ii) a first plurality of sentiments associated with content items within the first webpage, a plurality of keywords associated with the first webpage, and a similarity score between the first plurality of sentiments and the plurality of keywords; and (iii) a plurality of sentiment risk scores, wherein a sentiment risk score for at least one sentiment in the first plurality of sentiment risk scores is based on searching an approval list using at least one of the first plurality of sentiments as a search query; based on determining that the first URL is the cybersquatting attempt of the second URL, modifying an aggregate risk score by a first value from the plurality of risk tolerance values; based on determining that the similarity score is below a similarity threshold, modifying the aggregate risk score by a second value from the plurality of risk tolerance values; based on the plurality of sentiment risk scores, modifying the aggregate risk score by a third value from the plurality of risk tolerance values; determining that the aggregate risk score is within a first range of predetermined values; and in response to determining that the aggregate risk score is within the first predetermined range of values, associating at least one of the first webpage or the first domain name associated with the first URL with an exclusion list associated with the first advertiser.
1 . A method for providing contextual information associated with webpages, the method comprising:
receiving a plurality of risk tolerance values from a first advertiser;
accessing a first webpage through a first universal resource locator (URL), wherein the first webpage contains at least one dynamic advertising region;
determining, using a machine learning model,
(i) whether the first URL is likely to be a cybersquatting attempt of a second URL by parsing the first URL into a plurality of URL components and performing a pattern matching of each of the plurality of URL components against a set of URLS in which a first domain name associated with the first URL is compared to a second domain name associated with the second URL;
(ii) in response to determining that the first URL is likely to be the cybersquatting attempt of the second URL that the first URL is the cybersquatting attempting of the second URL based on the pattern matching, whether the first URL is associated with a valid identity certificate by inspecting the first webpage associated with the first URL;
(iii) in response to determining that the first URL is associated with the valid identity certificate, a first plurality of sentiments associated with content items within the first webpage, a plurality of keywords associated with the first webpage, and a similarity score between the first plurality of sentiments and the plurality of keywords; and
(iv) a plurality of sentiment risk scores, wherein a sentiment risk score for at least one sentiment in the first plurality of sentiment risk scores is based on searching an approval list using at least one of the first plurality of sentiments as a search query;
based on determining that the first URL is the cybersquatting attempt of the second URL, modifying an aggregate risk score by a first value from the plurality of risk tolerance values;
based on determining that the similarity score is below a similarity threshold, modifying the aggregate risk score by a second value from the plurality of risk tolerance values;
based on the plurality of sentiment risk scores, modifying the aggregate risk score by a third value from the plurality of risk tolerance values;
determining that the aggregate risk score is within a first range of predetermined values; and
in response to determining that the aggregate risk score is within the first predetermined range of values, associating at least one of the first webpage or the first domain name associated with the first URL with an exclusion list associated with the first advertiser.
2 . The method of claim 1 , wherein the method further comprises inhibiting the first advertiser from placing a bid for advertising in the at least one dynamic advertising region based on at least one of the first webpage or the first domain name being included on the exclusion list.
3 . The method of claim 1 , wherein the method further comprises inhibiting the first advertiser from placing a bid for advertising on a second webpage based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, wherein the second webpage is accessed at a third URL, wherein the domain name associated with the first URL is identical to a domain name associated with the third URL.
4 . The method of claim 1 , wherein the method further comprises, based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, adding the first plurality of sentiments associated with the first webpage and the first webpage to a training dataset for the machine learning model.
5 . The method of claim 1 , wherein the approval list comprises a database of web content that is approved by the first advertiser.
6 . A system for providing contextual information associated with webpages, the system comprising:
a hardware processor that is configured to:
receive a plurality of risk tolerance values from a first advertiser;
access a first webpage through a first universal resource locator (URL), wherein the first webpage contains at least one dynamic advertising region;
determine, using a machine learning model,
(i) whether the first URL is likely to be a cybersquatting attempt of a second URL by parsing the first URL into a plurality of URL components and performing a pattern matching of each of the plurality of URL components against a set of URLS in which a first domain name associated with the first URL is compared to a second domain name associated with the second URL;
(ii) in response to determining that the first URL is likely to be the cybersquatting attempt of the second URL that the first URL is the cybersquatting attempting of the second URL based on the pattern matching, whether the first URL is associated with a valid identity certificate by inspecting the first webpage associated with the first URL;
(iii) in response to determining that the first URL is associated with the valid identity certificate, a first plurality of sentiments associated with content items within the first webpage, a plurality of keywords associated with the first webpage, and a similarity score between the first plurality of sentiments and the plurality of keywords; and
(iv) a plurality of sentiment risk scores, wherein a sentiment risk score for at least one sentiment in the first plurality of sentiment risk scores is based on searching an approval list using at least one of the first plurality of sentiments as a search query;
based on determining that the first URL is the cybersquatting attempt of the second URL, modify an aggregate risk score by a first value from the plurality of risk tolerance values;
based on determining that the similarity score is below a similarity threshold, modify the aggregate risk score by a second value from the plurality of risk tolerance values;
based on the plurality of sentiment risk scores, modify the aggregate risk score by a third value from the plurality of risk tolerance values;
determine that the aggregate risk score is within a first range of predetermined values; and
in response to determining that the aggregate risk score is within the first predetermined range of values, associate at least one of the first webpage or the first domain name associated with the first URL with an exclusion list associated with the first advertiser.
7 . The system of claim 6 , wherein the hardware processor is further configured to inhibit the first advertiser from placing a bid for advertising in the at least one dynamic advertising region based on at least one of the first webpage or the first domain name being included on the exclusion list.
8 . The system of claim 6 , wherein the hardware processor is further configured to inhibit the first advertiser from placing a bid for advertising on a second webpage based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, wherein the second webpage is accessed at a third URL, wherein the domain name associated with the first URL is identical to a domain name associated with the third URL.
9 . The system of claim 6 , wherein the hardware processor is further configured to, based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, add the first plurality of sentiments associated with the first webpage and the first webpage to a training dataset for the machine learning model.
10 . The system of claim 6 , wherein the approval list comprises a database of web content that is approved by the first advertiser.
11 . A non-transitory computer-readable medium containing computer executable instructions that, when executed by a processor, cause the processor to perform a method for providing contextual information associated with webpages, the method comprising:
receiving a plurality of risk tolerance values from a first advertiser;
accessing a first webpage through a first universal resource locator (URL), wherein the first webpage contains at least one dynamic advertising region;
determining, using a machine learning model,
(i) whether the first URL is likely to be a cybersquatting attempt of a second URL by parsing the first URL into a plurality of URL components and performing a pattern matching of each of the plurality of URL components against a set of URLS in which a first domain name associated with the first URL is compared to a second domain name associated with the second URL;
(ii) in response to determining that the first URL is likely to be the cybersquatting attempt of the second URL that the first URL is the cybersquatting attempting of the second URL based on the pattern matching, whether the first URL is associated with a valid identity certificate by inspecting the first webpage associated with the first URL;
(iii) in response to determining that the first URL is associated with the valid identity certificate, a first plurality of sentiments associated with content items within the first webpage, a plurality of keywords associated with the first webpage, and a similarity score between the first plurality of sentiments and the plurality of keywords; and
based on determining that the first URL is the cybersquatting attempt of the second URL, modifying an aggregate risk score by a first value from the plurality of risk tolerance values;
based on determining that the similarity score is below a similarity threshold, modifying the aggregate risk score by a second value from the plurality of risk tolerance values;
based on the plurality of sentiment risk scores, modifying the aggregate risk score by a third value from the plurality of risk tolerance values;
determining that the aggregate risk score is within a first range of predetermined values; and
in response to determining that the aggregate risk score is within the first predetermined range of values, associating at least one of the first webpage or the first domain name associated with the first URL with an exclusion list associated with the first advertiser.
12 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises inhibiting the first advertiser from placing a bid for advertising in the at least one dynamic advertising region based on at least one of the first webpage or the first domain name being included on the exclusion list.
13 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises inhibiting the first advertiser from placing a bid for advertising on a second webpage based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, wherein the second webpage is accessed at a third URL, wherein the domain name associated with the first URL is identical to a domain name associated with the third URL.
14 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises, based on at least one of the first webpage and the domain name associated with the first URL being on the exclusion list, adding the first plurality of sentiments associated with the first webpage and the first webpage to a training dataset for the machine learning model.
15 . The non-transitory computer-readable medium of claim 11 , wherein the approval list comprises a database of web content that is approved by the first advertiser.