Search
Search Results
-
AGGLOMERATIVE CLUSTERIZATION ALGORITHMS FOR THE PROBLEMS OF ANALYSIS OF LINGUISTIC EXPERT INFORMATION
F.S. Bulyga, V.M. Kureichik2022-01-31Abstract ▼This article discusses and presents the main problems and principles of the data clustering
process, in particular, the principles and tasks of clustering text arrays of linguistic expert information.
In the course of this work, the main difficulties arising in the design of such systems were
identified, for example: the need for preprocessing data, reducing the size of the initial sample,
etc. To effectively perform the presented tasks, the implemented solution must have an integrated
approach that takes into account the efficiency indicators of methods aimed at solving individual
subtasks, as well as the ability to provide high efficiency indicators for the implementation of each
stage of the clustering process. In the presented work, various groups of hierarchical clustering
algorithms are considered, in particular, a subgroup of agglomerative clustering algorithms was
considered in relation to the problems of clustering linguistic expert information. In the described
work, a formal statement of the text clustering problem is given, and the main group of implemented
solutions based on the principles of agglomerative clustering is determined: ROCK, CURE,
CHAMELEON. A detailed review of each of the presented algorithms is carried out, and the main
advantages and disadvantages of each of them are formulated. The advantage of this work can be
considered the totality of the presented data on the algorithms, as well as the results of a comparative analysis, which make it possible to further assess the feasibility and potential probability of
using these solutions from the presented group of agglomerative clustering algorithms. The novelty
of this work lies in the formation of an overview analysis of existing approaches in the field of
hierarchical clustering for solving the problems of cluster analysis of linguistic expert information,
as well as the formation of the results of the comparative analysis of the considered algorithms. -
METHODS AND ALGORITHMS FOR TEXT DATA CLUSTERING (REVIEW)
V.V. Bova, Y.A. Kravchenko, S.I. Rodzin2022-11-01Abstract ▼The article deals with one of the important tasks of artificial intelligence – machine processing
of natural language. The solution of this problem based on cluster analysis makes it possible
to identify, formalize and integrate large amounts of linguistic expert information under conditions
of information uncertainty and weak structure of the original text resources obtained from
various subject areas. Cluster analysis is a powerful tool for exploratory analysis of text data,
which allows for an objective classification of any objects that are characterized by a number of
features and have hidden patterns. A review and analysis of modern modified algorithms for agglomerative
clustering CURE, ROCK, CHAMELEON, non-hierarchical clustering PAM, CLARA
and the affine transformation algorithm used at various stages of text data clustering, the effectiveness
of which is verified by experimental studies, is carried out. The paper substantiates the
requirements for choosing the most efficient clustering method for solving the problem of increasing the efficiency of intellectual processing of linguistic expert information. Also, the paper considers
methods for visualizing clustering results for interpreting the cluster structure and dependencies
on a set of text data elements and graphical means of their presentation in the form of
dendograms, scatterplots, VOS similarity diagrams, and intensity maps. To compare the quality of
the algorithms, internal and external performance metrics were used: "V-measure", "Adjusted
Rand index", "Silhouette". Based on the experiments, it was found that it is necessary to use a
hybrid approach, in which, for the initial selection of the number of clusters and the distribution of
their centers, use a hierarchical approach based on sequential combining and averaging the characteristics
of the closest data of a limited sample, when it is not possible to put forward a hypothesis
about the initial number of clusters. Next, connect iterative clustering algorithms that provide
high stability with respect to noise features and the presence of outliers. Hybridization increases
the efficiency of clustering algorithms. The research results showed that in order to increase the
computational efficiency and overcome the sensitivity when initializing the parameters of clustering
algorithms, it is necessary to use metaheuristic approaches to optimize the parameters of the
learning model and search for a global optimal solution. -
METHODOLOGY FOR CONSTRUCTING AND EVALUATING AN ONTOLOGICAL PROFILE FOR CONTENT PERSONALIZATION SYSTEMS: STAGES AND EVALUATION CRITERIA
Z.H. Mohammad248-2622025-12-30Abstract ▼This article presents the development and testing of a methodology for building an ontological profile designed for content personalization systems. It details the modular architecture of a web-based personalization system, illustrating the text processing and analysis methods and algorithms employed at each stage, and provides a step-by-step procedure for ontology creation. The methodology encompasses primary data processing, including the extraction of keywords and phrases, followed by their hierarchical clustering to reveal the semantic structure of the domain. Subsequent stages involve defining thresholds to filter out insignificant connections, and extracting and formalizing relationships between concepts using natural language processing techniques such as word-sense disambiguation and semantic similarity-based relationship extraction. An integrated pipeline was developed to implement this process, combining improved algorithms proposed by the author in previous studies, namely, an algorithm for extracting key phrases from individual text based on semantic similarity and a modified algorithm for word sense disambiguation. This pipeline also optimally integrated all necessary natural language processing tools, ensuring the efficient operation of these methods in the process of automatically constructing an ontology from text. The study places particular emphasis on a comprehensive evaluation of the resulting ontology using a specialized set of criteria designed to objectively assess the profile's quality, completeness, and consistency. A important component of the work is a computational experiment that clearly demonstrates the impact of each data processing stage on the final quality and efficacy of the ontology. The results show that the proposed method enables the construction of a practical, scalable, and relevant ontology, suitable for industrial deployment and integration into personalization systems to enhance their accuracy and adaptability
-
VERIFICATION OF DYNAMIC BIOMETRIC PARAMETERS OF A PERSONALITY BASED ON A PROBABLE NEURAL NETWORK
Y.A. Bryuhomitsky2021-01-19Abstract ▼Biometric identity verification is used primarily for access to computer and mobile systems, as
well as for remote (voice) verification. In fact, the most widespread systems are biometric verification
systems based on a fixed passphrase, which are quite simple to implement, but very vulnerable to
attacks of reproduction of a compromised short text. To eliminate this drawback, it is proposed to
carry out identity verification using a text that is arbitrary in terms of volume, content and language
(text-independent biometric verification). This paper proposes a generalized approach to solve the
problem of identity verification by dynamic biometric parameters of different modality (keyboard
writing, handwriting, voice). The presentation of dynamic biometrics signals is carried out by converting
them into a sequences of information units, each of which contains the same number of counts
of biometric signal of corresponding modality. The solution to this problem is carried out by monitoring
the degree of concentration of closely located information units (clusters) at certain points of the
multidimensional feature space. Such control is implemented on a probabilistic neural network thatstatistically evaluates the probability density of the distribution of information units in the corresponding
clusters with the subsequent determination of the total probability density for the entire
class of objects. The advantages of the proposed approach are: generalization of substantially different
methods of text-independent identity verification by dynamic biometric parameters of different
modality; the ability to make a verification decision for a fixed time of receipt of biometric data, determined
by the size of the model used; the ability to set the verification accuracy by changing the
dimension of the layer of probabilistic network samples. The disadvantage of the proposed approach
is the need for software implementation of a large-scale neural network. However, this drawback is
quickly leveled with an increase in the productivity of computer technology. -
SOLUTION OF THE PROBLEM OF INTELLECTUAL DATA ANALYSIS BASED ON BIOINSPIRED ALGORITHM
E.V. Kuliev, D.Y. Zaporozhets, Y.A. Kravchenko, М.М. Semenova2022-01-31Abstract ▼The article discusses a bioinspired algorithm for solving the problems of intellectual analysis.
The integration of bioinspired algorithms for solving data mining problems is a promising
area of research. As a bioinspired algorithm, an algorithm based on the adaptive behavior of an
ant colony is considered. The ant colony algorithm allows for a high-quality search for promising
solutions to obtain optimal and quasi-optimal solutions. The algorithm has the ability to search for
suitable logical conditions. The ant colony algorithm is based on the example of the behavior of
living ants in nature. Ants are able to find the shortest solution by adapting to changes in the environment.
The authors proposed a modified ant colony algorithm for solving the problem of data
mining. The clustering problem was chosen as the task of data mining. Clustering is a combining
of similar objects into groups, is one of the fundamental tasks in the field of data analysis and
Data Mining. The list of application areas where it is applied is wide: image segmentation, marketing,
anti-fraud, forecasting, text analysis and many others. The solution to this problem is of particular relevance in the context of the constantly growing volume of generated, transmitted and
processed data. Classical clustering methods are optimized by combining with the proposed
bioinspired optimization algorithm - the ant algorithm. The proposed method is a model in which
ants are represented as agents that randomly move in the solution space with some restrictions
(for example, obstacles in their path). To determine the effectiveness of the developed modified ant
algorithm (ALA) with the clustering algorithm, the authors carried out a series of computational
experiments. For comparison, we took the genetic algorithm, the monkey algorithm and the wolf
algorithm. The simulation results prove that the clustering-based ant algorithm gives better results
than other proposed algorithms. -
TEXT VECTORIZATION USING DATA MINING METHODS
Ali Mahmoud Mansour , Juman Hussain Mohammad, Y. A. Kravchenko2021-07-18Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. Traditional text vectorization methods such as TF-IDF and bag-of-words are effective and
characterized by intuitive interpretability, but suffer from the «curse of dimensionality», and they
are unable to capture the meanings of words. On the other hand, modern distributed methods effectively
capture the hidden semantics, but they are computationally intensive, time-consuming,
and uninterpretable. This article proposes a new text vectorization method called Bag of weighted
Concepts BoWC that presents a document according to the concepts’ information it contains. The
proposed method creates concepts by clustering word vectors (i.e. word embedding) then uses the
frequencies of these concept clusters to represent document vectors. To enrich the resulted document
representation, a new modified weighting function is proposed for weighting concepts based
on statistics extracted from word embedding information. The generated vectors are characterized
by interpretability, low dimensionality, high accuracy, and low computational costs when used in
data mining tasks. The proposed method has been tested on five different benchmark datasets in
two data mining tasks; document clustering and classification, and compared with several baselines,
including Bag-of-words, TF-IDF, Averaged GloVe, Bag-of-Concepts, and VLAC. The results
indicate that BoWC outperforms most baselines and gives 7 % better accuracy on average -
DEVELOPMENT OF ALGORITHMS OF INTELLIGENT SERVICE FOR INFORMATION SEARCH AND MONITORING
M. S. Anferova, A. M. Belevtsev2021-08-11Abstract ▼This paper describes the problem of strategic analysis and the choice of directions for the development
of an innovative enterprise in the conditions of transition to the 6th technological order and
industry 4.0. In these conditions, search and analytical processing of information cannot be fully performed
without the use of automated information and analytical systems, including those based on artificial
intelligence. During the analysis, the main priority functions that the developed services should
provide were identified. The main difficulties in the development of these services are identified, such as:
pre-processing of data and automated checking of the relevance of databases. To effectively solve thetasks set, the intelligent monitoring and information retrieval service should use an integrated approach,
taking into account the effectiveness of applying methods for individual subtasks, and ensure high efficiency
of implementing all stages of the intelligent monitoring procedure. In this regard, this paper describes
not only the development of a general intelligent search algorithm, but also individual block
algorithms necessary to ensure the priority functions of the service being developed. The paper presents
the following algorithms: an information search algorithm necessary to solve the problem of full-text
search of documents within the database of information resources of the information and analytical
complex; an algorithm for the procedure for entering new documents; an algorithm for pre-processing
data that includes stemming and removing punctuation marks for subsequent text analysis; an algorithm
for evaluating the ranking and relevance of information, including vectorization of documents; an algorithm
for clustering information search results based on the Kohonen neural network; the algorithm for
checking the relevance of information is to check whether the local copy of the document corresponds to
the current version on the source's web resource. The Python programming language for the implementation
of the presented algorithm is proposed and justified. The system provides automated continuous
monitoring with a high frequency of sending a request without the participation of an operator, which
will increase the quality and efficiency of information search in conditions of a large volume of unstructured
information. -
MODULE FOR ADJUSTING PARAMETERS OF ALGORITHMS FOR AUTOMATIC DETECTION AND TRACKING OF OBJECTS FOR OPTOELECTRONIC SYSTEMS
V. А. Tupikov, V. А. Pavlova, А.I. Lizin, P.А. Gessen71-812022-04-20Abstract ▼In order to create an innovative module for automatic correction of algorithms for automatic
detection and tracking of objects with real-time training, a study of world experience in the field
of general-purpose automatic tracking with the ability to recognize the tracking object for use in
embedded computing devices of optoelectronic systems of promising robotic complexes was carried
out. Based on the conducted research, methods and approaches have been selected and tested
that allow with the greatest accuracy, while maintaining high computational efficiency, to provide
on-the-fly training of classifiers (online learning) without a priori knowledge of the type of tracking object and to ensure subsequent correction during tracking and detection of the original object
in case of its short-term loss. Such methods include a histogram of directional gradients – a descriptor
of key features based on the analysis of the distribution of brightness gradients of an object
image. Its use allows you to reduce the amount of information used without losing key data
about the object and increase the speed of image processing. The article substantiates the choice
of one of the classification algorithms in real time, which allows solving the problem of binary
classification - the method of support vectors. Due to the high speed of data processing and the
need for a small amount of initial training data to build a separating hyperplane, on the basis of
which the classification of objects takes place, this method is chosen as the most suitable for solving
the task. To implement online training, a modification of the support vector machine was chosen,
implementing stochastic gradient descent at each step of the algorithm – Pegasos. Another
auxiliary method is the clustering method of key points – this ensures an accelerated selection of
objects for classification and training. The authors of the study carried out the development and
semi-natural modeling of the proposed module, evaluated the effectiveness of its work in the tasks
of correcting and detecting the object of interest in real time with preliminary online training in
the process of tracking the object. The developed algorithm has shown high efficiency in solving
the problem. In conclusion, proposals are presented to further improve the accuracy and probability
of detecting an object of interest by the developed algorithm, as well as to improve its performance
by optimizing calculations. -
ANALYSIS OF REQUIREMENTS AND DEVELOPMENT OF ALGORITHMS FOR INTELLIGENT MONITORING SERVICES
М.S. Anferova, А.М. Belevtsev2022-08-09Abstract ▼The paper considers the problems of strategic analysis and the choice of directions for the development
of innovative enterprises in the conditions of transition to the 6th technological order and industry
4.0. The main levels of analysis are determined. The objectives of the strategic analysis are outlined
based on the scale of the research being conducted. The analysis tasks are highlighted, the solution of
which will allow achieving the set goals. The complexity of solving global monitoring tasks, which are
caused by a large volume of heterogeneous and unstructured information, is shown. In these conditions,
thematic search and analytical processing of information cannot be performed without the use of automated information and analytical systems and the creation of search services based on artificial intelligence.
A general monitoring procedure is proposed. The main stages of monitoring technological trends
are defined, the tasks to be solved within a specific stage and the planned result are shown. Based on the
general monitoring procedure, the main priority functions that the developed services should have are
determined. As well as the problems of their development and structuring of the received information in
the form of information objects and clustering of documents. In contrast to the well-known global monitoring
systems, in which the search is based on indicators: an increase in the use of keywords, an increase
in the number of new authors, quoting works from related fields. Algorithms are proposed that
provide the definition of reference topics, assessment of ranking and relevance of information. The description
of the algorithms is given on the example of creating a summary information table, with the
help of which the interrelationships of documents of scientific and technological development in each
direction of monitoring and the search for specific documents in the database are formed. The construction
of search services based on the presented algorithms will ensure the allocation of reference topics
of documents, provide more reliable results of clustering of unstructured information and the formation
of scientific and technological trends in information and analytical complexes. To implement the algorithm,
it is proposed to use the Python programming language. The implementation of these algorithms
will improve the quality and efficiency of information retrieval in conditions of a large volume of unstructured
information. -
METHOD OF AUTOMATIC OPTIMIZATION OF THE FUZZY RULE BASE OF AN INTELLIGENT CONTROLLER BASED ON SUBTRACTIVE CLUSTERING
А.S. Ignatyeva , V.V. Shadrina , D.S. Ignatyev , А.V. Maksimov181-1972025-07-24Abstract ▼The aim of the work is to develop a method for optimizing the fuzzy rule base of an intelligent controller for controlling a technical object using subtractive clustering. The article provides an overview and a brief analysis of the state of affairs in the field of optimizing the operation of intelligent control systems. To achieve the goal of the study, a hybrid model has been developed in which the technical object is controlled using a classical PI controller and a fuzzy PI controller with a generated structure of a Cygeno-type fuzzy inference system and a developed model of an adaptive neuro-fuzzy inference system. This configuration of the model allows you to form a fuzzy rule base that does not depend on the expert's knowledge in the subject area. The article proposes a new method for optimizing the fuzzy controller rule base based on clustering methods, in particular subtractive clustering, which allows you to reduce the number of fuzzy logical inference rules and increase the performance of the technical object control system. First, a hybrid model synthesized on the basis of the values of the fuzzy and classical controllers before applying subtractive clustering was simulated. The application of subtractive clustering according to the method developed in the study for the values of the classical and fuzzy controllers allowed us to achieve their quantitative reduction by 1.7 and 5.25 times, respectively. Then, the hybrid model synthesized on the basis of the values of the fuzzy and classical controllers after applying subtractive clustering was simulated. The results obtained in the process of simulation showed high efficiency of the proposed method for optimizing the fuzzy controller rule base. Due to the application of subtractive clustering in the hybrid model for the intelligent controller, it was possible to significantly reduce the number of membership functions required to describe the input linguistic variables (from five to four) and reduce the number of fuzzy logical inference rules (from twenty-five to sixteen). The analysis of the resulting graphs of transient processes obtained for the hybrid models before and after applying subtractive clustering showed that the main indicators of the quality of the control process remain unchanged with a significant reduction in the calculations performed.
-
STOCHASTIC DYNAMIC MODEL OF UNDERWATER WIRELESS SENSOR NETWORK BASED ON LOUVAIN CLUSTERING ALGORITHM
А.М. Maevsky , V.А. Ryzhov , Т. А. Fedorova , I. V. Kozhemyakin , N.М. Burov62-812025-07-24Abstract ▼Underwater wireless sensor networks (UWSNs) play an important role in monitoring ocean processes, underwater navigation, environmental control and security. However, underwater environment features such as high signal attenuation, limited energy resources and changing network topology create significant challenges in organizing efficient data transmission. To optimize network operation and extend its service life, a clustering method is used to group nodes, reduce the load on communication channels and improve energy efficiency. However, in the event of network node failure, static clustering becomes ineffective, which requires the implementation of dynamic reclustering. The procedure of redistributing node roles and rebuilding the network topology allows maintaining communication stability and minimizing data losses, taking into account the energy balance of the entire network as a whole. This paper examines modern approaches to clustering and reclustering in UWSNs taking into account the energy balance, node failure probability and interference in the transmission medium. The development of adaptive UWSN control methods is an urgent task aimed at increasing the reliability, energy efficiency and durability of underwater communication networks. The article presents a stochastic cross-level model for dynamic three-dimensional PBSNs of arbitrary topology. The model uses a new clustering/reclustering technique based on the Louvain algorithm, a routing protocol built on the Dijkstra method, and a time-domain management (TDMA) method. The proposed PBSN operating model is the basis for the developed simulation complex, which allows assessing the efficiency and reliability of the network, taking into account the loss of connectivity and vulnerabilities for PBSNs of various scales and purposes. As part of the research, a parametric analysis of systematic calculations of the PBSN functional characteristics was performed. The results of the analysis showed that the proposed simulation model provides an increase in the autonomous network operation time and a decrease in the number of lost messages compared to the models of other authors
-
MULTIMODAL DATA FEATURE EXTRACTION METHOD FOR NETWORK ATTACK CLASSIFICATION
A.V. Balyberdin6-162025-07-24Abstract ▼An intrusion detection system (IDS) is an important component of corporate data network (CDN) protection. IDS analyzes network traffic and detects network attacks. Depending on the detection methods, IDS can be classified into the following types of systems: signature-based analysis systems, anomaly detection systems (ADS), and hybrid systems combining the aforementioned approaches. Recently, anomaly detection systems (IDS) have been actively developing. For anomaly detection systems, network attacks are anomalous behavior of network traffic consisting of a set of features or event attributes. Modern IDS are based on machine and deep learning methods, and therefore the detection of network attacks and anomalies is formulated as a classification and clustering problem. To solve these problems, methods for optimizing the feature space of network traffic are required. The aim of the work is to develop a feature extraction method based on a multimodal approach to representing network traffic data for classifying network attacks. The paper considers the analysis of relevant studies on feature extraction methods from various fields. The objective of the study is to improve classification efficiency using a multimodal representation of network traffic features. The result of the work is a method for extracting data features based on two modalities: a spectral representation of network traffic features and an image feature matrix. The novelty of the presented method lies in the application of the windowed Fourier transform method for network traffic events, followed by the calculation of spectral features for discrete signals, as well as the transformation of data features into an image matrix and its expansion to optimize the feature space using a convolutional neural network (CNN). Evaluation of the multimodal method showed that this method increased the classification accuracy for unbalanced classes of network attacks








