US2025358315A1PendingUtilityA1

Local detection of fraudulent websites using lightweight machine learning models

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: May 20, 2024Filed: May 20, 2024Published: Nov 20, 2025
Est. expiryMay 20, 2044(~17.8 yrs left)· nominal 20-yr term from priority
H04L 63/1483G06V 30/19173
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure describes a fraudulent website detection system that provides a framework for locally detecting fraudulent websites on a client device. For example, using a local lightweight machine learning model, the fraudulent website detection system can detect and respond to fraudulent websites in real time. In some examples, the fraudulent website detection system is integrated into a web browser to promptly identify fraudulent websites. Moreover, the fraudulent website detection system, operating on multiple client devices, can collaborate with an online threat detection system to quickly notify other client devices about fraudulent websites and to utilize aggregated reports to improve the lightweight machine learning model.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for determining one or more fraudulent websites locally on a computing device, comprising:
 capturing an image of a website that is loaded on a client device;   generating a set of classification scores for the website based on website request information and the image using a threat assessment machine learning model executed locally on the client device;   determining a website threat score for the website based on aggregating a subset of the set of classification scores; and   based on the website threat score for the website satisfying one or more threat thresholds, performing one or more actions reporting the website as fraudulent.   
     
     
         2 . The computer-implemented method of  claim 1 , further comprising determining to use the threat assessment machine learning model to assess the website for fraudulent behavior in response to detecting the website that is loaded on the client device. 
     
     
         3 . The computer-implemented method of  claim 2 , further comprising determining to use the threat assessment machine learning model based on verifying one or more low-computational filter conditions. 
     
     
         4 . The computer-implemented method of  claim 3 , wherein the one or more low-computational filter conditions include:
 a first filter condition that verifies whether the website is associated with a commonly accessed website;   a second filter condition that verifies whether the website is not included on a fraudulent website list;   a third filter condition that verifies whether the website includes a threshold number of grammar and typographical errors;   a fourth filter condition that provides verification based on a previous website that linked to the website; and   a fifth filter condition that provides verification based on client device permissions requested by the website.   
     
     
         5 . The computer-implemented method of  claim 1 , wherein capturing the image of the website that is loaded on the client device includes capturing a screen capture of the website as it appears within a browser window to a user. 
     
     
         6 . The computer-implemented method of  claim 1 , further comprising:
 detecting one or more client device permissions requested by the website;   determining one or more accepted permissions associated with the one or more client device permissions requested; and   generating the website request information to indicate the one or more client device permissions and the one or more accepted permissions.   
     
     
         7 . The computer-implemented method of  claim 1 , further comprising:
 providing a corpus of fraudulent-based website information to a large generative model with instructions to determine website classification types and corresponding fraudulent associations;   selecting a set of website classification types; and   generating the threat assessment machine learning model based on the set of website classification types to generate a classification score for each of the website classification types for candidate websites.   
     
     
         8 . The computer-implemented method of  claim 1 , wherein the threat assessment machine learning model determines a fraudulent website verdict that the website is fraudulent before the client device receives additional user input associated with the website. 
     
     
         9 . The computer-implemented method of  claim 1 , further comprising converting the image of the website to converted text before providing the converted text of the image to the threat assessment machine learning model. 
     
     
         10 . The computer-implemented method of  claim 1 , wherein classification types of the threat assessment machine learning model have a binary value indicating a fraudulent association. 
     
     
         11 . The computer-implemented method of  claim 10 , wherein determining the website threat score for the website includes:
 generating a classification type subset based on identifying fraudulent associations for each of the classification types; and   generating the subset of the set of classification scores based on classification scores generated from the classification type subset.   
     
     
         12 . The computer-implemented method of  claim 1 , further comprising:
 determining that the website threat score for the website satisfies a user threat threshold; and   based on the user threat threshold being satisfied, notifying a user associated with the client device that the website is fraudulent.   
     
     
         13 . The computer-implemented method of  claim 1 , further comprising:
 determining that the website threat score for the website satisfies a global threat threshold; and   based on the global threat threshold being satisfied, reporting the website to a fraudulent listener service.   
     
     
         14 . The computer-implemented method of  claim 13 , wherein reporting the website includes providing a fraudulent website verdict, the image of the website, and the website request information to the fraudulent listener service. 
     
     
         15 . The computer-implemented method of  claim 14 , further comprising receiving, from the fraudulent listener service, one or more websites to add to a fraudulent website list for blocking fraudulent websites. 
     
     
         16 . A system comprising:
 a processing system; and   a computer memory comprising instructions that, when executed by the processing system, cause the system to perform operations of:
 capturing an image of a website that is loaded on a client device; 
 generating a set of classification scores for the website based on website request information and the image using a threat assessment machine learning model executed locally on the client device; 
 determining a website threat score for the website based on aggregating a subset of the set of classification scores; and 
 based on the website threat score for the website satisfying one or more threat thresholds, performing one or more actions reporting the website as fraudulent. 
   
     
     
         17 . The system of  claim 16 , wherein the operations further include:
 determining that the website threat score for the website satisfies a user threat threshold; and   based on the user threat threshold being satisfied, notifying a user associated with the client device that the website is fraudulent.   
     
     
         18 . The system of  claim 16 , wherein the operations further include:
 determining that the website threat score for the website satisfies a global threat threshold; and   based on the global threat threshold being satisfied, reporting the website to a fraudulent listener service.   
     
     
         19 . A computer-implemented method for determining one or more fraudulent websites locally on a computing device, comprising:
 capturing an image of a website that is loaded on a client device;   generating a set of classification scores for the website based on website request information and the image using a threat assessment classification machine learning model executed locally on the client device;   determining a website threat score for the website based on aggregating a subset of classification scores; and   based on determining that the website threat score for the website satisfies a user threat threshold, determining a fraudulent website verdict for the website and preventing a user associated with the client device from further accessing the website.   
     
     
         20 . The computer-implemented method of  claim 19 , wherein the threat assessment classification machine learning model does not use a remote resource to determine the set of classification scores for the website.

Join the waitlist — get patent alerts

Track US2025358315A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.