Local detection of fraudulent websites using lightweight machine learning models
Abstract
This disclosure describes a fraudulent website detection system that provides a framework for locally detecting fraudulent websites on a client device. For example, using a local lightweight machine learning model, the fraudulent website detection system can detect and respond to fraudulent websites in real time. In some examples, the fraudulent website detection system is integrated into a web browser to promptly identify fraudulent websites. Moreover, the fraudulent website detection system, operating on multiple client devices, can collaborate with an online threat detection system to quickly notify other client devices about fraudulent websites and to utilize aggregated reports to improve the lightweight machine learning model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for determining one or more fraudulent websites locally on a computing device, comprising:
capturing an image of a website that is loaded on a client device; generating a set of classification scores for the website based on website request information and the image using a threat assessment machine learning model executed locally on the client device; determining a website threat score for the website based on aggregating a subset of the set of classification scores; and based on the website threat score for the website satisfying one or more threat thresholds, performing one or more actions reporting the website as fraudulent.
2 . The computer-implemented method of claim 1 , further comprising determining to use the threat assessment machine learning model to assess the website for fraudulent behavior in response to detecting the website that is loaded on the client device.
3 . The computer-implemented method of claim 2 , further comprising determining to use the threat assessment machine learning model based on verifying one or more low-computational filter conditions.
4 . The computer-implemented method of claim 3 , wherein the one or more low-computational filter conditions include:
a first filter condition that verifies whether the website is associated with a commonly accessed website; a second filter condition that verifies whether the website is not included on a fraudulent website list; a third filter condition that verifies whether the website includes a threshold number of grammar and typographical errors; a fourth filter condition that provides verification based on a previous website that linked to the website; and a fifth filter condition that provides verification based on client device permissions requested by the website.
5 . The computer-implemented method of claim 1 , wherein capturing the image of the website that is loaded on the client device includes capturing a screen capture of the website as it appears within a browser window to a user.
6 . The computer-implemented method of claim 1 , further comprising:
detecting one or more client device permissions requested by the website; determining one or more accepted permissions associated with the one or more client device permissions requested; and generating the website request information to indicate the one or more client device permissions and the one or more accepted permissions.
7 . The computer-implemented method of claim 1 , further comprising:
providing a corpus of fraudulent-based website information to a large generative model with instructions to determine website classification types and corresponding fraudulent associations; selecting a set of website classification types; and generating the threat assessment machine learning model based on the set of website classification types to generate a classification score for each of the website classification types for candidate websites.
8 . The computer-implemented method of claim 1 , wherein the threat assessment machine learning model determines a fraudulent website verdict that the website is fraudulent before the client device receives additional user input associated with the website.
9 . The computer-implemented method of claim 1 , further comprising converting the image of the website to converted text before providing the converted text of the image to the threat assessment machine learning model.
10 . The computer-implemented method of claim 1 , wherein classification types of the threat assessment machine learning model have a binary value indicating a fraudulent association.
11 . The computer-implemented method of claim 10 , wherein determining the website threat score for the website includes:
generating a classification type subset based on identifying fraudulent associations for each of the classification types; and generating the subset of the set of classification scores based on classification scores generated from the classification type subset.
12 . The computer-implemented method of claim 1 , further comprising:
determining that the website threat score for the website satisfies a user threat threshold; and based on the user threat threshold being satisfied, notifying a user associated with the client device that the website is fraudulent.
13 . The computer-implemented method of claim 1 , further comprising:
determining that the website threat score for the website satisfies a global threat threshold; and based on the global threat threshold being satisfied, reporting the website to a fraudulent listener service.
14 . The computer-implemented method of claim 13 , wherein reporting the website includes providing a fraudulent website verdict, the image of the website, and the website request information to the fraudulent listener service.
15 . The computer-implemented method of claim 14 , further comprising receiving, from the fraudulent listener service, one or more websites to add to a fraudulent website list for blocking fraudulent websites.
16 . A system comprising:
a processing system; and a computer memory comprising instructions that, when executed by the processing system, cause the system to perform operations of:
capturing an image of a website that is loaded on a client device;
generating a set of classification scores for the website based on website request information and the image using a threat assessment machine learning model executed locally on the client device;
determining a website threat score for the website based on aggregating a subset of the set of classification scores; and
based on the website threat score for the website satisfying one or more threat thresholds, performing one or more actions reporting the website as fraudulent.
17 . The system of claim 16 , wherein the operations further include:
determining that the website threat score for the website satisfies a user threat threshold; and based on the user threat threshold being satisfied, notifying a user associated with the client device that the website is fraudulent.
18 . The system of claim 16 , wherein the operations further include:
determining that the website threat score for the website satisfies a global threat threshold; and based on the global threat threshold being satisfied, reporting the website to a fraudulent listener service.
19 . A computer-implemented method for determining one or more fraudulent websites locally on a computing device, comprising:
capturing an image of a website that is loaded on a client device; generating a set of classification scores for the website based on website request information and the image using a threat assessment classification machine learning model executed locally on the client device; determining a website threat score for the website based on aggregating a subset of classification scores; and based on determining that the website threat score for the website satisfies a user threat threshold, determining a fraudulent website verdict for the website and preventing a user associated with the client device from further accessing the website.
20 . The computer-implemented method of claim 19 , wherein the threat assessment classification machine learning model does not use a remote resource to determine the set of classification scores for the website.Join the waitlist — get patent alerts
Track US2025358315A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.