Search
Search Results
-
A METHOD FOR DETECTING COMPUTER ATTACKS BASED ON H-DDPM NETWORK TRAFFIC DATA AUGMENTATION MODEL
А. V. Balyberdin2026-07-07Abstract ▼This paper examines the problem of improving the detection of computer attacks (CA) by an intrusion detection system (IDS) under conditions of significant network traffic data imbalance. Based on an analysis of methods for reducing data imbalance, it is concluded that classical methods of balancing and generative augmentation do not preserve the statistical structure of multidimensional tabular data, including their fractal properties and self-similarity, which reduces the quality of classifier training. This paper proposes a method for detecting computer attacks (CA) based on the H-DDPM data augmentation model, a modification of the DDPM diffusion probabilistic model, in which the variance of the added Gaussian noise in the forward process depends on the Hurst exponent H for each CA class. The method includes data preprocessing, the formation of time series using sliding windows, H estimation using DFA and R/S methods, and the generation of synthetic data for training the LSTM classifier. The method is evaluated using the general performance metrics Accuracy, Recall, F1, ROC-AUC, and G-means, as well as Precision, Recall, and
F1-score for each class. Experiments were conducted on the CSE-CSE-CIC-IDS2018 and UNSW-NB15 datasets. A comparison was made with other methods, such as SMOTE, GAN, and DDPM. The experimental results show that H-DDPM improves the efficiency of CA detection, outperforming similar methods in terms of imbalance-sensitive metrics. Furthermore, experimental validations demonstrate that directly using the Hurst H exponent for CA classes in the H-DDPM model improves the recall and balanced quality of CA detection. It is noted that H-DDPM has an impact on CA classification, manifested by an increase in false positives and a decrease in the ROC-AUC metric, which requires additional tuning of the classifier model hyperparameters and filtering of synthetic data -
MODIFIED WORD SENSE DISAMBIGUATION METHOD BASED ON DISTRIBUTED REPRESENTATION METHODS
Y. A. Kravchenko, Mansour Ali Mahmoud, Mohammad Juman Hussain2021-08-11Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. This paper describes a modified Word Sense Disambiguation (WSD) method which extends
two well-known variations of the Lesk WSD approach. Given a word and its context, Lesk
bases its calculations on the overlap between the context of a word and each definition of its senses
(gloss) in order to select the proper meaning. The main contribution of the proposed method is
the adoption of the concept of “similarity” between definition and context instead of "overlap", in
addition to expanding the definition with examples provided by WordNet for each sense of the
target word. The proposed method is also characterized by the use of text similarity measurement
functions defined in a distributed semantic space. The proposed method has been tested on five
different benchmark datasets for words sense disambiguation tasks and compared with several
basic methods, including simple Lesk, extended Lesk, WordNet 1st sense, Babelfy and UKB. The
results show that proposed method outperforms most basic methods with the exception of Babelfy
and the WN 1st sense methods. -
ON THE SIMILARITY FUNCTION OF GRAPHIC REPRESENTATIONS OF EXECUTIVE FILES IN THE OBFUSCING TRANSFORMATION EVALUATION MODEL
P.D. Borisov , Y.V. Kosolapov264-2732025-07-24Abstract ▼Obfuscation of program code is used to complicate its analysis in a model when the analyst has full access to the program. Obfuscation is usually divided into cryptographically secure and heuristically resistant. In the first case, the complexity of the analysis is comparable to the difficulty of solving some known mathematical problem. In the second case, the resistance is usually justified by the lack of effective techniques for analyzing the obfuscation method known at the time of its creation. Cryptographically secure obfuscation has not yet found practical application, while heuristically resistant is widely used. Previously, the authors proposed a model for assessing the efficiency and resistance of heuristic obfuscating transformations based on the use of a similarity function. In this paper, such a similarity function is constructed using machine learning methods based on a comparison of the graphical representation of program executable files. In particular, the comparison is performed using a convolutional network with four convolutional layers, an RMSprop optimizer, an NLLLoss loss function, and two outputs of a fully connected layer. The proposed function is used in the implementation of a model for evaluating the efficiency and resistance of obfuscating transformations. In addition to the similarity function, the implementation of the model also includes: a basic set of obfuscating transformations provided by the Hikari obfuscator; a set of obfuscating transformation sequences based on the basic set; a test set of programs for training models based on the CoreUtils, PolyBench and HashCat program sets; approximation of the most "understandable" version of the program using the smallest version of the program (searched among the versions obtained using various optimization options of the GCC, Clang and AOCC compilers); a program deobfuscation scheme based on the optimizing compiler from LLVM. The results of an experimental study with the implemented model showed that it is impractical to use the constructed similarity function in the framework of the evaluation model due to its low accuracy, but it is possible to use it when constructing more complex functions.








