Search
Search Results
-
AGGLOMERATIVE CLUSTERIZATION ALGORITHMS FOR THE PROBLEMS OF ANALYSIS OF LINGUISTIC EXPERT INFORMATION
F.S. Bulyga, V.M. Kureichik2022-01-31Abstract ▼This article discusses and presents the main problems and principles of the data clustering
process, in particular, the principles and tasks of clustering text arrays of linguistic expert information.
In the course of this work, the main difficulties arising in the design of such systems were
identified, for example: the need for preprocessing data, reducing the size of the initial sample,
etc. To effectively perform the presented tasks, the implemented solution must have an integrated
approach that takes into account the efficiency indicators of methods aimed at solving individual
subtasks, as well as the ability to provide high efficiency indicators for the implementation of each
stage of the clustering process. In the presented work, various groups of hierarchical clustering
algorithms are considered, in particular, a subgroup of agglomerative clustering algorithms was
considered in relation to the problems of clustering linguistic expert information. In the described
work, a formal statement of the text clustering problem is given, and the main group of implemented
solutions based on the principles of agglomerative clustering is determined: ROCK, CURE,
CHAMELEON. A detailed review of each of the presented algorithms is carried out, and the main
advantages and disadvantages of each of them are formulated. The advantage of this work can be
considered the totality of the presented data on the algorithms, as well as the results of a comparative analysis, which make it possible to further assess the feasibility and potential probability of
using these solutions from the presented group of agglomerative clustering algorithms. The novelty
of this work lies in the formation of an overview analysis of existing approaches in the field of
hierarchical clustering for solving the problems of cluster analysis of linguistic expert information,
as well as the formation of the results of the comparative analysis of the considered algorithms. -
METHODS AND ALGORITHMS FOR TEXT DATA CLUSTERING (REVIEW)
V.V. Bova, Y.A. Kravchenko, S.I. Rodzin2022-11-01Abstract ▼The article deals with one of the important tasks of artificial intelligence – machine processing
of natural language. The solution of this problem based on cluster analysis makes it possible
to identify, formalize and integrate large amounts of linguistic expert information under conditions
of information uncertainty and weak structure of the original text resources obtained from
various subject areas. Cluster analysis is a powerful tool for exploratory analysis of text data,
which allows for an objective classification of any objects that are characterized by a number of
features and have hidden patterns. A review and analysis of modern modified algorithms for agglomerative
clustering CURE, ROCK, CHAMELEON, non-hierarchical clustering PAM, CLARA
and the affine transformation algorithm used at various stages of text data clustering, the effectiveness
of which is verified by experimental studies, is carried out. The paper substantiates the
requirements for choosing the most efficient clustering method for solving the problem of increasing the efficiency of intellectual processing of linguistic expert information. Also, the paper considers
methods for visualizing clustering results for interpreting the cluster structure and dependencies
on a set of text data elements and graphical means of their presentation in the form of
dendograms, scatterplots, VOS similarity diagrams, and intensity maps. To compare the quality of
the algorithms, internal and external performance metrics were used: "V-measure", "Adjusted
Rand index", "Silhouette". Based on the experiments, it was found that it is necessary to use a
hybrid approach, in which, for the initial selection of the number of clusters and the distribution of
their centers, use a hierarchical approach based on sequential combining and averaging the characteristics
of the closest data of a limited sample, when it is not possible to put forward a hypothesis
about the initial number of clusters. Next, connect iterative clustering algorithms that provide
high stability with respect to noise features and the presence of outliers. Hybridization increases
the efficiency of clustering algorithms. The research results showed that in order to increase the
computational efficiency and overcome the sensitivity when initializing the parameters of clustering
algorithms, it is necessary to use metaheuristic approaches to optimize the parameters of the
learning model and search for a global optimal solution. -
METHODOLOGY FOR CONSTRUCTING AND EVALUATING AN ONTOLOGICAL PROFILE FOR CONTENT PERSONALIZATION SYSTEMS: STAGES AND EVALUATION CRITERIA
Z.H. Mohammad248-2622025-12-30Abstract ▼This article presents the development and testing of a methodology for building an ontological profile designed for content personalization systems. It details the modular architecture of a web-based personalization system, illustrating the text processing and analysis methods and algorithms employed at each stage, and provides a step-by-step procedure for ontology creation. The methodology encompasses primary data processing, including the extraction of keywords and phrases, followed by their hierarchical clustering to reveal the semantic structure of the domain. Subsequent stages involve defining thresholds to filter out insignificant connections, and extracting and formalizing relationships between concepts using natural language processing techniques such as word-sense disambiguation and semantic similarity-based relationship extraction. An integrated pipeline was developed to implement this process, combining improved algorithms proposed by the author in previous studies, namely, an algorithm for extracting key phrases from individual text based on semantic similarity and a modified algorithm for word sense disambiguation. This pipeline also optimally integrated all necessary natural language processing tools, ensuring the efficient operation of these methods in the process of automatically constructing an ontology from text. The study places particular emphasis on a comprehensive evaluation of the resulting ontology using a specialized set of criteria designed to objectively assess the profile's quality, completeness, and consistency. A important component of the work is a computational experiment that clearly demonstrates the impact of each data processing stage on the final quality and efficacy of the ontology. The results show that the proposed method enables the construction of a practical, scalable, and relevant ontology, suitable for industrial deployment and integration into personalization systems to enhance their accuracy and adaptability








