A Survey of Machine Learning Methods and Challenges for Windows Malware Classification

06/15/2020
by   Edward Raff, et al.
0

Malware classification is a difficult problem, to which machine learning methods have been applied for decades. Yet progress has often been slow, in part due to a number of unique difficulties with the task that occur through all stages of the developing a machine learning system: data collection, labeling, feature creation and selection, model selection, and evaluation. In this survey we will review a number of the current methods and challenges related to malware classification, including data collection, feature extraction, and model construction, and evaluation. Our discussion will include thoughts on the constraints that must be considered for machine learning based solutions in this domain, and yet to be tackled problems for which machine learning could also provide a solution. This survey aims to be useful both to cybersecurity practitioners who wish to learn more about how machine learning can be applied to the malware problem, and to give data scientists the necessary background into the challenges in this uniquely complicated space.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/28/2022

Multi-feature Dataset for Windows PE Malware Classification

This paper describes a multi-feature dataset for training machine learni...
research
08/03/2018

Machine Learning Aided Static Malware Analysis: A Survey and Tutorial

Malware analysis and detection techniques have been evolving during the ...
research
04/04/2019

Malware Detection using Machine Learning and Deep Learning

Research shows that over the last decade, malware has been growing expon...
research
06/09/2023

A Survey on Cross-Architectural IoT Malware Threat Hunting

In recent years, the increase in non-Windows malware threats had turned ...
research
07/11/2017

A Survey on Resilient Machine Learning

Machine learning based system are increasingly being used for sensitive ...
research
11/06/2017

Computer activity learning from system call time series

Using a previously introduced similarity function for the stream of syst...

Please sign up or login with your details

Forgot password? Click here to reset