Methods and systems for detecting malicious webpages
Abstract
Methods and systems are disclosed for training a malicious webpages detector for detecting malicious webpages, based on a training set comprising a plurality of samples representing malicious and non-malicious webpages. Text content can be extracted from the source code of each sample, and/or non-text content can be extracted from each sample, in order to train respectively at least a first deep learning neural network and a second deep learning neural network of the malicious webpages detector. A malicious webpages detector can detect whether or not a webpage is malicious, by extracting text content from the source code of the webpage, and/or non-text content from the webpage, thereafter providing prospects that the webpage is malicious based on the extracted data.
Claims
exact text as granted — not AI-modified1 . A method of detecting a malicious webpage using a malicious webpages detector, wherein the malicious webpages detector comprises at least a first deep learning neural network and a second deep learning neural network operable on at least a processing unit, the method comprising, for a webpage:
extracting text content from at least a source code of said webpage, providing first prospect of whether said text content constitutes malicious content, using the first deep learning neural network, determining non-text content of said webpage, wherein the non-text content comprises at least one file including a binary content usable to display at least one of an image and an animated content of the webpage, extracting at least part of the binary content of the file, feeding the binary content extracted from the file to the second deep learning neural network, providing second prospect of whether said non-text content constitutes malicious content, using the second deep learning neural network, and detecting whether the webpage is malicious based on at least one of the first prospect and the second prospect.
2 . The method according to claim 1 , wherein:
the first prospect comprises at least one of a probability that the text content constitutes malicious content, and a probability that the text content does not constitute malicious content, the second prospect comprises at least one of a probability that the non-text content constitutes malicious content and a probability that the non-text content does not constitute malicious content, and wherein a webpage is detected as malicious based on a comparison of at least one of the first prospect and the second prospect with a criterion.
3 . The method of claim 1 , comprising, following the detection of a malicious webpage, performing a security action to avoid a connection of a user to said webpage or to limit the connection of the user to said webpage.
4 . The method of claim 1 , wherein extracting the text content of the source code of a webpage comprises extracting the whole raw text content of the source code of the webpage, or at least part of it.
5 . The method of claim 1 , wherein the malicious webpages detector is operable for at least one of:
a plurality of different browsers used to access the webpage, and a plurality of different operating systems on which a browser is used to access the webpage, and a plurality of different programming languages of webpages.
6 . The method of claim 1 , wherein the text content comprises non-obfuscated content and obfuscated content, or only obfuscated content, the method comprising:
deobfuscating said obfuscated content, feeding the non-obfuscated content and the deobfuscated content, or only the deobfuscated content, to the first deep learning neural network, and providing first prospects of whether said text content constitutes malicious content, using the first deep learning neural network.
7 . The method of claim 1 , wherein at least one of (i) and (ii) is met:
(i) the text content comprises text content without semantic meaning; (ii) the binary content comprises raw binary content without semantic meaning.
8 . A system operative to detect a malicious webpage, comprising at least a first deep learning neural network and a second deep learning neural network operable on a processing unit, the system being configured, for a webpage, to:
extract text content from at least a source code of said webpage, provide first prospect of whether said text content constitutes malicious content, using the first deep learning neural network, determine non-text content of said webpage, wherein the non-text content comprises at least one file including a binary content usable to display at least one of an image and an animated content of the webpage, extract at least part of the binary content of the file, feed the binary content extracted from the file to the second deep learning neural network, provide second prospects of whether said non-text content constitutes malicious content, using the second deep learning neural network, and detect whether the webpage is malicious based on at least one of the first prospect and the second prospect.
9 . The system according to claim 8 , wherein:
the first prospect comprises at least one of a probability that the text content constitutes malicious content and a probability that the text content does not constitute malicious content, the second prospect comprises at least one of a probability that the non-text content constitutes malicious content and a probability that the non-text content does not constitute malicious content, and wherein a webpage is detected as malicious based on a comparison of at least one of the first prospect and the second prospect with a criterion.
10 . The system of claim 8 , configured to, following the detection of a malicious webpage, perform a security action to avoid a connection of a user to said webpage or to limit the connection of the user to said webpage.
11 . The system of claim 8 , wherein extracting the text content of the source code of a webpage comprises extracting the whole raw text content of the source code of the webpage, or at least part of it.
12 . The system of claim 8 , said system being operable for at least one of:
a plurality of browsers used to access the webpage, and a plurality of operating systems of the user accessing the webpage, and a plurality of programming languages of the webpage.
13 . The system of claim 8 , wherein said system is located in at least one of a plug-in of a web browser and a proxy.
14 . The system of claim 8 , wherein the text content comprises non-obfuscated content and obfuscated content, or only obfuscated content, the system being configured to:
deobfuscate said obfuscated content, feed the non-obfuscated content and the deobfuscated content, or only the deobfuscated content, to the first deep learning neural network, and provide first prospect of whether said text content constitutes malicious content, using the first deep learning neural network.
15 . The system of claim 8 , wherein at least one of (i) and (ii) is met:
(i) the text content comprises text content without semantic meaning; (ii) the binary content comprises raw binary content without semantic meaning.
16 . A system operative to detect a malicious webpage, comprising at least a deep learning neural network operable on a processing unit, the system being configured, for a webpage, to:
determine non-text content of said webpage, wherein the non-text content comprises at least one file including a binary content usable to display at least one of an image and an animated content of the webpage, extract at least part of the binary content of the file, feed the binary content extracted from the file to the deep learning neural network, provide prospect of whether said non-text content constitutes malicious content, using the deep learning neural network. detect whether the webpage is malicious based at least on the prospect.
17 . The system of claim 16 , wherein the binary content comprises raw binary content without semantic meaning.
18 . The system of claim 16 , said system being operable for at least one of:
a plurality of browsers used to access the webpage, and a plurality of operating systems of the user accessing the webpage, and a plurality of programming languages of the webpage.Join the waitlist — get patent alerts
Track US2021006577A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.