Data Augmentation for Opcode Sequence Based Malware Detection

06/22/2021
by   Niall McLaughlin, et al.
0

Data augmentation has been successfully used in many areas of deep-learning to significantly improve model performance. Typically data augmentation simulates realistic variations in data in order to increase the apparent diversity of the training-set. However, for opcode-based malware analysis, where deep learning methods are already achieving state of the art performance, it is not immediately clear how to apply data augmentation. In this paper we study different methods of data augmentation starting with basic methods using fixed transformations and moving to methods that adapt to the data. We propose a novel data augmentation method based on using an opcode embedding layer within the network and its corresponding opcode embedding matrix to perform adaptive data augmentation during training. To the best of our knowledge this is the first paper to carry out a systematic study of different augmentation methods applied to opcode sequence based malware classification.

READ FULL TEXT
research
10/11/2018

Efficient Augmentation via Data Subsampling

Data augmentation is commonly used to encode invariances in learning met...
research
11/15/2022

Local Magnification for Data and Feature Augmentation

In recent years, many data augmentation techniques have been proposed to...
research
08/20/2017

Improving Deep Learning using Generic Data Augmentation

Deep artificial neural networks require a large corpus of training data ...
research
12/08/2017

Music Transcription by Deep Learning with Data and "Artificial Semantic" Augmentation

In this progress paper the previous results of the single note recogniti...
research
05/22/2023

Revisiting Data Augmentation in Model Compression: An Empirical and Comprehensive Study

The excellent performance of deep neural networks is usually accompanied...
research
04/17/2023

Transformer with Selective Shuffled Position Embedding using ROI-Exchange Strategy for Early Detection of Knee Osteoarthritis

Knee OsteoArthritis (KOA) is a prevalent musculoskeletal disorder that c...
research
11/28/2022

Exoplanet Detection by Machine Learning with Data Augmentation

It has recently been demonstrated that deep learning has significant pot...

Please sign up or login with your details

Forgot password? Click here to reset