IP Library Granted Patent US 12,445,488
Granted Patent B2
US 12,445,488 · App. 18/309,240 · Granted Oct 14, 2025

Webpage phishing auto-detection

Inventor: Karthik Shourya Kaligotla (Frisco, TX)
H04L63/1483H04L63/1416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,488
App. No.
18/309,240
Granted
Oct 14, 2025
Kind
B2
Abstract

Embodiments relate to systems and method for phishing webpage detection, the system comprising a weblink retrieval unit configured to obtain a weblink of a website of interest, a website capturing unit configured to capture a first image of the website of interest, a detection unit configured to detect a similarity between the first image of the website of interest and a list of pre-configured legitimate websites corresponding to a label in a database, an IP address grabbing unit configured to grab an IP address of the website of interest if the similarity is detected, a comparison unit configured to compare the IP address of the website of interest with a database of legitimate IP addresses in at least one of a static context and a dynamic context to identify a difference and an alarm unit configured to generate a phishing website alarm based on a presence the difference.

Claims (48)

1. A system comprising:

a processor; and a non-transitory memory comprising instructions that when executed by the processor, performs:

obtaining a weblink of a website of interest;

capturing a first image of the website of interest by rendering and screenshotting the website of interest;

detecting, using a machine learning model, a similarity between the first image of the website of interest and a second image of legitimate websites corresponding to a label in a database;

grabbing a first IP address of the website of interest if the similarity is detected;

comparing the first IP address of the website of interest with a database of legitimate IP addresses in at least one of a static context and a dynamic context to identify a difference; and

generating a phishing website alarm based on a presence of the difference,

wherein the machine learning model is trained using a first dataset in an image database; wherein the first dataset is expanded to a third dataset by adding a second dataset to the first dataset, the second dataset is obtained by performing augmentation and creating variations to the first dataset; and

wherein the machine learning model is retrained using the third dataset in the image database.

2. The system of claim 1 , further comprises sending information comprising the first image and the first IP address about a known phishing website to at least one of a domain owner and a repository on an as-needed basis.

3. The system of claim 2 , further comprises storing details and information related to the known phishing websites and the legitimate websites from a public database.

4. The system of claim 1 , further comprises collecting a list of legitimate URLs corresponding to the legitimate websites.

5. The system of claim 1 , further comprises prompting a command to enter the weblink of the website of interest.

6. The system of claim 1 , wherein the processor captures the first image of the website of interest using a web crawler.

7. The system of claim 1 , wherein the processor detects the similarity between the first image of the website of interest and the legitimate websites using computer vision.

8. The system of claim 1 , wherein the system is further configured to:

retrieve a first digital certificate of the website of interest; and

compare the first digital certificate of the website of interest with a digital certificate of the legitimate websites that the first image was classified into to identify a second difference.

9. A method comprising:

obtaining a weblink of a website of interest;

retrieving a first image of the website of interest by rendering and screenshotting the website of interest;

detecting a similarity, using a machine learning model, between the first image of the website of interest and a second image of a legitimate website;

grabbing a first IP address of the website of interest if the similarity is detected;

comparing the first IP address of website of interest with a second IP address of the legitimate website to detect a difference; and generating a phishing website alarm based on a presence of the difference;

training the machine learning model using a first dataset in an image database;

expanding the first dataset to a third dataset by adding a second dataset to the first dataset, wherein the second dataset is obtained by performing augmentation and creating variations to the first dataset; and

retraining the machine learning model using the third dataset in the image database.

10. The method of claim 9 , wherein retrieving the first image of the website can be done using at least one of network scanning, and capturing a Document Object Model (DOM) of the website of interest utilizing a web scraper and creating an image with the captured Document Object model.

11. The method of claim 9 , wherein comparing the first IP address with the second IP address is done using at least one of an exact match algorithm, subnet match algorithm, geolocation match algorithm and behavior-based match algorithm.

12. The method of claim 9 , further comprises maintaining and updating the image database comprising a series of first images, second images and augmented images and associated labels of the website of interest.

13. The method of claim 12 , wherein the image database is used to update the machine learning model on similarity for images of the website of interest.

14. The method of claim 13 , further comprises periodically training the machine learning model with the image database that is maintained through automated software features.

15. The method of claim 13 , wherein the machine learning model is utilized in a workflow of an application.

16. The method of claim 15 , wherein the application resides on a browser as an add-on extension.

17. The method of claim 15 , further comprises:

retrieving a first digital certificate of the website of interest; and

comparing the first digital certificate of the website of interest with a digital certificate of the legitimate website that the first image was classified into, to identify a second difference.

18. A non-transitory computer-readable storage medium, storing executable instructions, when executed by a processor, causing the processor to implement a machine learning (ML)-based phishing protection method, the method comprising:

obtaining a weblink of a website of interest;

capturing a first image of the website of interest by rendering and screenshotting the website of interest;

detecting a similarity, using a machine learning model, between the first image of the website of interest and a second image of a legitimate website;

grabbing a first IP address of the website of interest if the similarity is detected;

comparing the first IP address of the website of interest with a second IP address of the legitimate website to detect a difference;

generating a phishing website alarm based on a presence of the difference;

training the machine learning model using a first dataset in an image database;

expanding the first dataset to a third dataset by adding a second dataset to the first dataset, wherein the second dataset is obtained by performing augmentation and creating variations to the first dataset; and

retraining the machine learning model using the third dataset in the image database.

Continuity (1)
Related Publication 20230344868A1 · Oct 26, 2023
References Cited (15)
US 8856937B1 · Wuest et al. · 2014 [cited by applicant]
US 10834128B1 · Rajagopalan et al. · 2020 [cited by applicant]
US 11381597B2 · Lancioni et al. · 2022 [cited by applicant]
US 11444978B1 · Liao et al. · 2022 [cited by applicant]
US 11483343B2 · Kohavi · 2022 [cited by applicant]
US 11570211B1 · Liu · 2023 [cited by applicant]
US 11582226B2 · Starov et al. · 2023 [cited by applicant]
US 11595438B2 · Quint et al. · 2023 [cited by applicant]
US 20150067839A1 · Wardman et al. · 2015 [cited by applicant]
US 20210312611A1 · Neumann · 2021 [cited by examiner]
US 20210344711A1 · Cleveland et al. · 2021 [cited by applicant]
US 20220174092A1 · Farjon · 2022 [cited by examiner]
US 20220385694A1 · Zverkov et al. · 2022 [cited by applicant]
US 20230081266A1 · Crume · 2023 [cited by examiner]
Maurya, Swati, Harpreet Singh Saini, and Anurag Jain. “Browser extension based hybrid anti-phishing framework using feature selection.” International Journal of Advanced Computer Science and Applications 10.11 (2019). (… [cited by examiner]