Search
Search Results
-
MODIFIED WORD SENSE DISAMBIGUATION METHOD BASED ON DISTRIBUTED REPRESENTATION METHODS
Y. A. Kravchenko, Mansour Ali Mahmoud, Mohammad Juman Hussain2021-08-11Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. This paper describes a modified Word Sense Disambiguation (WSD) method which extends
two well-known variations of the Lesk WSD approach. Given a word and its context, Lesk
bases its calculations on the overlap between the context of a word and each definition of its senses
(gloss) in order to select the proper meaning. The main contribution of the proposed method is
the adoption of the concept of “similarity” between definition and context instead of "overlap", in
addition to expanding the definition with examples provided by WordNet for each sense of the
target word. The proposed method is also characterized by the use of text similarity measurement
functions defined in a distributed semantic space. The proposed method has been tested on five
different benchmark datasets for words sense disambiguation tasks and compared with several
basic methods, including simple Lesk, extended Lesk, WordNet 1st sense, Babelfy and UKB. The
results show that proposed method outperforms most basic methods with the exception of Babelfy
and the WN 1st sense methods. -
A TRANSFORMER-BASED ALGORITHM FOR CLASSIFYING LONG TEXTS
Ali Mahmoud Mansour2024-08-12Abstract ▼The article is devoted to the urgent problem of representing and classifying long text documents using
transformers. Transformers-based text representation methods cannot effectively process long sequences
due to their self-attention process, which scales quadratically with the sequence length. This limitation
leads to high computational complexity and the inability to apply such models for processing long
documents. To eliminate this drawback, the article developed an algorithm based on the SBERT transformer,
which allows building a vector representation of long text documents. The key idea of the algorithm
is the application of two different procedures for creating a vector representation: the first is based
on text segmentation and averaging of segment vectors, and the second is based on concatenation of segment
vectors. This combination of procedures allows preserving important information from long documents.
To verify the effectiveness of the algorithm, a computational experiment was conducted on a group
of classifiers built on the basis of the proposed algorithm and a group of well-known text vectorization
methods, such as TF-IDF, LSA, and BoWC. The results of the computational experiment showed that
transformer-based classifiers generally achieve better classification accuracy results compared to classical
methods. However, this advantage is achieved at the cost of higher computational complexity and,
accordingly, longer training and application times for such models. On the other hand, classical text
vectorization methods, such as TF-IDF, LSA, and BoWC, demonstrated higher speed, making them more
preferable in cases where pre-encoding is not allowed and real-time operation is required. The proposed
algorithm has proven its high efficiency and led and led to an increase in the classification accuracy of the
BBC dataset by 0.5% according to the F1 criterion. -
TEXT VECTORIZATION USING DATA MINING METHODS
Ali Mahmoud Mansour , Juman Hussain Mohammad, Y. A. Kravchenko2021-07-18Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. Traditional text vectorization methods such as TF-IDF and bag-of-words are effective and
characterized by intuitive interpretability, but suffer from the «curse of dimensionality», and they
are unable to capture the meanings of words. On the other hand, modern distributed methods effectively
capture the hidden semantics, but they are computationally intensive, time-consuming,
and uninterpretable. This article proposes a new text vectorization method called Bag of weighted
Concepts BoWC that presents a document according to the concepts’ information it contains. The
proposed method creates concepts by clustering word vectors (i.e. word embedding) then uses the
frequencies of these concept clusters to represent document vectors. To enrich the resulted document
representation, a new modified weighting function is proposed for weighting concepts based
on statistics extracted from word embedding information. The generated vectors are characterized
by interpretability, low dimensionality, high accuracy, and low computational costs when used in
data mining tasks. The proposed method has been tested on five different benchmark datasets in
two data mining tasks; document clustering and classification, and compared with several baselines,
including Bag-of-words, TF-IDF, Averaged GloVe, Bag-of-Concepts, and VLAC. The results
indicate that BoWC outperforms most baselines and gives 7 % better accuracy on average








