Search
Search Results
-
COMPARATIVE ANALYSIS OF METHODS OF VECTORIZATION OF HIGH DIMENSIONAL TEXT DATA
F.S. Bulyga, V. М. Kureichik2023-06-07Abstract ▼The presented publication is devoted to an overview of the problem of presenting textual information
for the subsequent implementation of cluster analysis in the framework of processing
and managing high-dimensional information. Modern requirements for analytical, search and
recommendation information systems demonstrate the weak formation of a holistic solution that
can provide a sufficient level of speed and quality of the results obtained within the framework of
the current information technology market. The search for a solution to the presented problem
entails the need to conduct an objective analysis of existing solutions for representing textual information
in vector space, in order to form a holistic view of the advantages and disadvantages of
the analyzed approaches, as well as the formation of criteria that allow one to implement their
own approach, devoid of identified weaknesses. The presented work is analytical, and allows you
to get an idea of the current state and elaboration of the identified problem within a limited subject
area. Clustering of text data is the automatic formation of subsets, the elements of which are instances
of documents of some researched, unstructured sample of a fixed dimension. This process
can be classified as unsupervised learning, which implies the absence of an expert who personally
assigns class indices to the original sample of documents. However, the implementation of cluster
analysis of text data without any pre-processing is impossible. To do this, it is necessary to ensure
standardization and reduction of input data to a single format and form. Within the framework of
this stage of the implementation of cluster analysis, the presented publication discusses methods
for preprocessing text data. The novelty of the presented publication lies in the formation of the
theoretical basis of the main methods of text data vectorization, by systematizing and objectifying
the proposed assumptions, by conducting a series of experimental studies. The main difference of
this work from the already published scientific works is the systematization and analysis of modern
solutions, as well as the hypotheses about the relevance and effectiveness of our own hybridized
approach designed for text data vectorization. -
AGGLOMERATIVE CLUSTERIZATION ALGORITHMS FOR THE PROBLEMS OF ANALYSIS OF LINGUISTIC EXPERT INFORMATION
F.S. Bulyga, V.M. Kureichik2022-01-31Abstract ▼This article discusses and presents the main problems and principles of the data clustering
process, in particular, the principles and tasks of clustering text arrays of linguistic expert information.
In the course of this work, the main difficulties arising in the design of such systems were
identified, for example: the need for preprocessing data, reducing the size of the initial sample,
etc. To effectively perform the presented tasks, the implemented solution must have an integrated
approach that takes into account the efficiency indicators of methods aimed at solving individual
subtasks, as well as the ability to provide high efficiency indicators for the implementation of each
stage of the clustering process. In the presented work, various groups of hierarchical clustering
algorithms are considered, in particular, a subgroup of agglomerative clustering algorithms was
considered in relation to the problems of clustering linguistic expert information. In the described
work, a formal statement of the text clustering problem is given, and the main group of implemented
solutions based on the principles of agglomerative clustering is determined: ROCK, CURE,
CHAMELEON. A detailed review of each of the presented algorithms is carried out, and the main
advantages and disadvantages of each of them are formulated. The advantage of this work can be
considered the totality of the presented data on the algorithms, as well as the results of a comparative analysis, which make it possible to further assess the feasibility and potential probability of
using these solutions from the presented group of agglomerative clustering algorithms. The novelty
of this work lies in the formation of an overview analysis of existing approaches in the field of
hierarchical clustering for solving the problems of cluster analysis of linguistic expert information,
as well as the formation of the results of the comparative analysis of the considered algorithms.








