Skip to main content Skip to main navigation menu Skip to site footer
##common.pageHeaderLogo.altText##
Izvestiya SFedU
Engineering sciences
  • Current
  • Previous issues
    • Archive
    • Issues 1995 – 2019
  • Editorial Board
  • About journal
    • Officially
    • The main tasks
    • Main sections
    • Specialties of the Higher Attestation Commission of the Russian Federation
    • Editor-in-Chief
ISSN 1999-9429 print
ISSN 2311-3103 online
  • Login
  1. Home /
  2. Search

Search

Advanced filters
Published After
Published Before

Search Results

##search.searchResults.foundPlural##
  • AGGLOMERATIVE CLUSTERIZATION ALGORITHMS FOR THE PROBLEMS OF ANALYSIS OF LINGUISTIC EXPERT INFORMATION

    F.S. Bulyga, V.M. Kureichik
    2022-01-31
    Abstract ▼

    This article discusses and presents the main problems and principles of the data clustering
    process, in particular, the principles and tasks of clustering text arrays of linguistic expert information.
    In the course of this work, the main difficulties arising in the design of such systems were
    identified, for example: the need for preprocessing data, reducing the size of the initial sample,
    etc. To effectively perform the presented tasks, the implemented solution must have an integrated
    approach that takes into account the efficiency indicators of methods aimed at solving individual
    subtasks, as well as the ability to provide high efficiency indicators for the implementation of each
    stage of the clustering process. In the presented work, various groups of hierarchical clustering
    algorithms are considered, in particular, a subgroup of agglomerative clustering algorithms was
    considered in relation to the problems of clustering linguistic expert information. In the described
    work, a formal statement of the text clustering problem is given, and the main group of implemented
    solutions based on the principles of agglomerative clustering is determined: ROCK, CURE,
    CHAMELEON. A detailed review of each of the presented algorithms is carried out, and the main
    advantages and disadvantages of each of them are formulated. The advantage of this work can be
    considered the totality of the presented data on the algorithms, as well as the results of a comparative analysis, which make it possible to further assess the feasibility and potential probability of
    using these solutions from the presented group of agglomerative clustering algorithms. The novelty
    of this work lies in the formation of an overview analysis of existing approaches in the field of
    hierarchical clustering for solving the problems of cluster analysis of linguistic expert information,
    as well as the formation of the results of the comparative analysis of the considered algorithms.

  • METHODS AND ALGORITHMS FOR TEXT DATA CLUSTERING (REVIEW)

    V.V. Bova, Y.A. Kravchenko, S.I. Rodzin
    2022-11-01
    Abstract ▼

    The article deals with one of the important tasks of artificial intelligence – machine processing
    of natural language. The solution of this problem based on cluster analysis makes it possible
    to identify, formalize and integrate large amounts of linguistic expert information under conditions
    of information uncertainty and weak structure of the original text resources obtained from
    various subject areas. Cluster analysis is a powerful tool for exploratory analysis of text data,
    which allows for an objective classification of any objects that are characterized by a number of
    features and have hidden patterns. A review and analysis of modern modified algorithms for agglomerative
    clustering CURE, ROCK, CHAMELEON, non-hierarchical clustering PAM, CLARA
    and the affine transformation algorithm used at various stages of text data clustering, the effectiveness
    of which is verified by experimental studies, is carried out. The paper substantiates the
    requirements for choosing the most efficient clustering method for solving the problem of increasing the efficiency of intellectual processing of linguistic expert information. Also, the paper considers
    methods for visualizing clustering results for interpreting the cluster structure and dependencies
    on a set of text data elements and graphical means of their presentation in the form of
    dendograms, scatterplots, VOS similarity diagrams, and intensity maps. To compare the quality of
    the algorithms, internal and external performance metrics were used: "V-measure", "Adjusted
    Rand index", "Silhouette". Based on the experiments, it was found that it is necessary to use a
    hybrid approach, in which, for the initial selection of the number of clusters and the distribution of
    their centers, use a hierarchical approach based on sequential combining and averaging the characteristics
    of the closest data of a limited sample, when it is not possible to put forward a hypothesis
    about the initial number of clusters. Next, connect iterative clustering algorithms that provide
    high stability with respect to noise features and the presence of outliers. Hybridization increases
    the efficiency of clustering algorithms. The research results showed that in order to increase the
    computational efficiency and overcome the sensitivity when initializing the parameters of clustering
    algorithms, it is necessary to use metaheuristic approaches to optimize the parameters of the
    learning model and search for a global optimal solution.

  • METHODOLOGY FOR CONSTRUCTING AND EVALUATING AN ONTOLOGICAL PROFILE FOR CONTENT PERSONALIZATION SYSTEMS: STAGES AND EVALUATION CRITERIA

    Z.H. Mohammad
    248-262
    2025-12-30
    Abstract ▼

    This article presents the development and testing of a methodology for building an ontological profile designed for content personalization systems. It details the modular architecture of a web-based personalization system, illustrating the text processing and analysis methods and algorithms employed at each stage, and provides a step-by-step procedure for ontology creation. The methodology encompasses primary data processing, including the extraction of keywords and phrases, followed by their hierarchical clustering to reveal the semantic structure of the domain. Subsequent stages involve defining thresholds to filter out insignificant connections, and extracting and formalizing relationships between concepts using natural language processing techniques such as word-sense disambiguation and semantic similarity-based relationship extraction. An integrated pipeline was developed to implement this process, combining improved algorithms proposed by the author in previous studies, namely, an algorithm for extracting key phrases from individual text based on semantic similarity and a modified algorithm for word sense disambiguation. This pipeline also optimally integrated all necessary natural language processing tools, ensuring the efficient operation of these methods in the process of automatically constructing an ontology from text. The study places particular emphasis on a comprehensive evaluation of the resulting ontology using a specialized set of criteria designed to objectively assess the profile's quality, completeness, and consistency. A important component of the work is a computational experiment that clearly demonstrates the impact of each data processing stage on the final quality and efficacy of the ontology. The results show that the proposed method enables the construction of a practical, scalable, and relevant ontology, suitable for industrial deployment and integration into personalization systems to enhance their accuracy and adaptability

1 - 3 of 3 items

links

For authors
  • Submit article
  • Author Guidelines
  • Editorial Policy
  • Reviewing
  • Ethics of scientific publications
  • Open access policy
  • Supporting documents
Language
  • English
  • русский

journal

* not an advertisement

index

Индексация журнала
* not an advertisement
Information
  • For Readers
  • For Authors
  • For Librarians
Address: 347900, Taganrog, Chekhov St., 22, A-211 Phone: +7 (8634) 37-19-80 E-mail: iborodyanskiy@sfedu.ru
Publication is free
More information about the publishing system, Platform and Workflow by OJS/PKP.
logo Developed by RDCenter