Skip to main content Skip to main navigation menu Skip to site footer
##common.pageHeaderLogo.altText##
Izvestiya SFedU
Engineering sciences
  • Current
  • Previous issues
    • Archive
    • Issues 1995 – 2019
  • Editorial Board
  • About journal
    • Officially
    • The main tasks
    • Main sections
    • Specialties of the Higher Attestation Commission of the Russian Federation
    • Editor-in-Chief
ISSN 1999-9429 print
ISSN 2311-3103 online
  • Login
  1. Home /
  2. Search

Search

Advanced filters
Published After
Published Before

Search Results

##search.searchResults.foundPlural##
  • MODIFIED WORD SENSE DISAMBIGUATION METHOD BASED ON DISTRIBUTED REPRESENTATION METHODS

    Y. A. Kravchenko, Mansour Ali Mahmoud, Mohammad Juman Hussain
    2021-08-11
    Abstract ▼

    In the text mining tasks, textual representation should be not only efficient but also interpretable,
    as this enables an understanding of the operational logic underlying the data mining
    models. This paper describes a modified Word Sense Disambiguation (WSD) method which extends
    two well-known variations of the Lesk WSD approach. Given a word and its context, Lesk
    bases its calculations on the overlap between the context of a word and each definition of its senses
    (gloss) in order to select the proper meaning. The main contribution of the proposed method is
    the adoption of the concept of “similarity” between definition and context instead of "overlap", in
    addition to expanding the definition with examples provided by WordNet for each sense of the
    target word. The proposed method is also characterized by the use of text similarity measurement
    functions defined in a distributed semantic space. The proposed method has been tested on five
    different benchmark datasets for words sense disambiguation tasks and compared with several
    basic methods, including simple Lesk, extended Lesk, WordNet 1st sense, Babelfy and UKB. The
    results show that proposed method outperforms most basic methods with the exception of Babelfy
    and the WN 1st sense methods.

  • A TRANSFORMER-BASED ALGORITHM FOR CLASSIFYING LONG TEXTS

    Ali Mahmoud Mansour
    2024-08-12
    Abstract ▼

    The article is devoted to the urgent problem of representing and classifying long text documents using
    transformers. Transformers-based text representation methods cannot effectively process long sequences
    due to their self-attention process, which scales quadratically with the sequence length. This limitation
    leads to high computational complexity and the inability to apply such models for processing long
    documents. To eliminate this drawback, the article developed an algorithm based on the SBERT transformer,
    which allows building a vector representation of long text documents. The key idea of the algorithm
    is the application of two different procedures for creating a vector representation: the first is based
    on text segmentation and averaging of segment vectors, and the second is based on concatenation of segment
    vectors. This combination of procedures allows preserving important information from long documents.
    To verify the effectiveness of the algorithm, a computational experiment was conducted on a group
    of classifiers built on the basis of the proposed algorithm and a group of well-known text vectorization
    methods, such as TF-IDF, LSA, and BoWC. The results of the computational experiment showed that
    transformer-based classifiers generally achieve better classification accuracy results compared to classical
    methods. However, this advantage is achieved at the cost of higher computational complexity and,
    accordingly, longer training and application times for such models. On the other hand, classical text
    vectorization methods, such as TF-IDF, LSA, and BoWC, demonstrated higher speed, making them more
    preferable in cases where pre-encoding is not allowed and real-time operation is required. The proposed
    algorithm has proven its high efficiency and led and led to an increase in the classification accuracy of the
    BBC dataset by 0.5% according to the F1 criterion.

  • TEXT VECTORIZATION USING DATA MINING METHODS

    Ali Mahmoud Mansour , Juman Hussain Mohammad, Y. A. Kravchenko
    2021-07-18
    Abstract ▼

    In the text mining tasks, textual representation should be not only efficient but also interpretable,
    as this enables an understanding of the operational logic underlying the data mining
    models. Traditional text vectorization methods such as TF-IDF and bag-of-words are effective and
    characterized by intuitive interpretability, but suffer from the «curse of dimensionality», and they
    are unable to capture the meanings of words. On the other hand, modern distributed methods effectively
    capture the hidden semantics, but they are computationally intensive, time-consuming,
    and uninterpretable. This article proposes a new text vectorization method called Bag of weighted
    Concepts BoWC that presents a document according to the concepts’ information it contains. The
    proposed method creates concepts by clustering word vectors (i.e. word embedding) then uses the
    frequencies of these concept clusters to represent document vectors. To enrich the resulted document
    representation, a new modified weighting function is proposed for weighting concepts based
    on statistics extracted from word embedding information. The generated vectors are characterized
    by interpretability, low dimensionality, high accuracy, and low computational costs when used in
    data mining tasks. The proposed method has been tested on five different benchmark datasets in
    two data mining tasks; document clustering and classification, and compared with several baselines,
    including Bag-of-words, TF-IDF, Averaged GloVe, Bag-of-Concepts, and VLAC. The results
    indicate that BoWC outperforms most baselines and gives 7 % better accuracy on average

1 - 3 of 3 items

links

For authors
  • Submit article
  • Author Guidelines
  • Editorial Policy
  • Reviewing
  • Ethics of scientific publications
  • Open access policy
  • Supporting documents
Language
  • English
  • русский

journal

* not an advertisement

index

Индексация журнала
* not an advertisement
Information
  • For Readers
  • For Authors
  • For Librarians
Address: 347900, Taganrog, Chekhov St., 22, A-211 Phone: +7 (8634) 37-19-80 E-mail: iborodyanskiy@sfedu.ru
Publication is free
More information about the publishing system, Platform and Workflow by OJS/PKP.
logo Developed by RDCenter