Search
Search Results
-
DEVELOPMENT OF A CHATBOT FOR CLASSIFICATION AND ANALYSIS OF NATURAL LANGUAGE TEXTS USING LOCAL LARGE LANGUAGE MODELS
Juman Hussain Mohammad , Juman Hussain Mohammad , Y.А. Kravchenko159-1712025-07-24Abstract ▼This paper explores local large language models (LLMs) and their application in text classification tasks, while also comparing their performance with traditional methods. The paper provides a comprehensive review of several key local LLMs, with particular focus on their architectural advantages, characteristics, and application domains. Specifically, we examine models with varying numbers of parameters, their ability to adapt to specialized domains, and their computational requirements when deployed on local hardware. Special emphasis is placed on the trade-offs between performance and resource efficiency. As a practical contribution, we developed a chatbot that utilizes local LLMs (such as DeepSeek, Gemma, and Llama2 via Ollama) to classify incoming texts into predefined categories, demonstrating the operation of these models without cloud computing. The system features a modular architecture that allows for easy integration of new models and comparison of their effectiveness. The computational experiment involves evaluating the accuracy and inference speed of local LLMs compared to simpler methods such as Sentence-BERT, TF-IDF and BoWC, highlighting scenarios in which local models outperform or underperform traditional approaches. Testing was conducted using the benchmark BBC dataset. The results show that language models (including 7-billion parameter models) demonstrate strong and logically consistent classification performance in natural language text processing. However, their results are not perfect for benchmark datasets. Notably, we identified cases where all tested models, including traditional methods, misclassified documents, suggesting potential issues with data labeling. These findings indicate the need to reconsider benchmark labels in standard datasets, particularly for domains with subjective categories where expert evaluations may vary significantly. On the other hand, while local LLMs lag behind cloud-based solutions in speed, their advantages in data privacy and offline operation make them suitable for specialized tasks. This is particularly valuable in medical and financial institutions where protection of sensitive information is critical, and where local models can be fine-tuned for specific business processes without the constraints of cloud APIs.
-
METHODOLOGY FOR CONSTRUCTING AND EVALUATING AN ONTOLOGICAL PROFILE FOR CONTENT PERSONALIZATION SYSTEMS: STAGES AND EVALUATION CRITERIA
Z.H. Mohammad248-2622025-12-30Abstract ▼This article presents the development and testing of a methodology for building an ontological profile designed for content personalization systems. It details the modular architecture of a web-based personalization system, illustrating the text processing and analysis methods and algorithms employed at each stage, and provides a step-by-step procedure for ontology creation. The methodology encompasses primary data processing, including the extraction of keywords and phrases, followed by their hierarchical clustering to reveal the semantic structure of the domain. Subsequent stages involve defining thresholds to filter out insignificant connections, and extracting and formalizing relationships between concepts using natural language processing techniques such as word-sense disambiguation and semantic similarity-based relationship extraction. An integrated pipeline was developed to implement this process, combining improved algorithms proposed by the author in previous studies, namely, an algorithm for extracting key phrases from individual text based on semantic similarity and a modified algorithm for word sense disambiguation. This pipeline also optimally integrated all necessary natural language processing tools, ensuring the efficient operation of these methods in the process of automatically constructing an ontology from text. The study places particular emphasis on a comprehensive evaluation of the resulting ontology using a specialized set of criteria designed to objectively assess the profile's quality, completeness, and consistency. A important component of the work is a computational experiment that clearly demonstrates the impact of each data processing stage on the final quality and efficacy of the ontology. The results show that the proposed method enables the construction of a practical, scalable, and relevant ontology, suitable for industrial deployment and integration into personalization systems to enhance their accuracy and adaptability
-
MODIFIED WORD SENSE DISAMBIGUATION METHOD BASED ON DISTRIBUTED REPRESENTATION METHODS
Y. A. Kravchenko, Mansour Ali Mahmoud, Mohammad Juman Hussain2021-08-11Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. This paper describes a modified Word Sense Disambiguation (WSD) method which extends
two well-known variations of the Lesk WSD approach. Given a word and its context, Lesk
bases its calculations on the overlap between the context of a word and each definition of its senses
(gloss) in order to select the proper meaning. The main contribution of the proposed method is
the adoption of the concept of “similarity” between definition and context instead of "overlap", in
addition to expanding the definition with examples provided by WordNet for each sense of the
target word. The proposed method is also characterized by the use of text similarity measurement
functions defined in a distributed semantic space. The proposed method has been tested on five
different benchmark datasets for words sense disambiguation tasks and compared with several
basic methods, including simple Lesk, extended Lesk, WordNet 1st sense, Babelfy and UKB. The
results show that proposed method outperforms most basic methods with the exception of Babelfy
and the WN 1st sense methods. -
TO ESTIMATION OF ATTRACTION AREA OF EQUILIBRIUM IN NONLINEAR CONTROL SYSTEMS
Almashaal Mohammad Jalal6-142025-07-31Abstract ▼Designing nonlinear control systems is still difficult, so many researchers are trying to find
some useful ways and methods to solve this problem. As a result of such research, some methods
have been seen trying to design a good enough control system for nonlinear plants. But a disadvantage
of these methods is the complexity, so it created a need to compare some methods to determine
which one is the easiest method to design a control system for nonlinear plants. It was
found a way to compare two methods, which is comparing the regions of initial conditions of the systems which are designed using these methods. Two analytical nonlinear control systems design
methods are compared on the example of the design control systems mobile robots. The algebraic
polynomial-matrix method uses a quasilinear model, and the feedback linearization method uses
particular feedback. Both considered methods give a bounded domain of equilibrium attraction,
therefore the obtained control systems can be operated only with bounded initial conditions. The
numerical example of designing the control systems for one object by these methods and the estimates
of the attraction areas of the system’s equilibriums in these systems are given in the paper.
As a result of this paper, it was found that using the algebraic polynomial-matrix method will get a
bigger section of initial conditions of the plant’s variable than the same section which is given by
the feedback linearization method. -
KEYPHRASE EXTRACTION BASED ON LARGE LANGUAGE MODELS
Mohammad Juman Hussain2024-11-10Abstract ▼The article addresses the current problem of extracting key phrases from natural language texts,
which is a critical task in the field of natural language processing and text mining. It examines in detail
the main approaches to extracting key phrases (keywords), including both traditional methods and modern
approaches based on artificial intelligence. The paper discusses a set of widely used methods in this field,
such as TF-IDF, RAKE, YAKE, and linguistic parser-based methods. These methods are based on statistical
principles and/or graph structures, but they often face problems related to their insufficient ability to
take into account the context of the text. The GPT-3 large language model demonstrates superior contextual
understanding compared to traditional methods for key phrase extraction. This advanced capability
allows GPT-3 to more accurately identify and extract relevant key phrases from text. The comparative
analysis using the Inspec benchmark dataset reveals GPT-3's significantly higher performance in terms of
Mean Average Precision (MAP@K). However, it should be noted that despite high accuracy and extraction
quality, the use of large language models may be limited in real-time applications due to their longer
response time compared to classical statistical methods. Thus, the article emphasizes the need for further
research in this area to optimize key phrase extraction algorithms, taking into account real-time requirements
and text context -
TO ESTIMATION OF ATTRACTION AREA OF EQUILIBRIUM OF NONLINEAR CONTROL SYSTEMS
Mohammad Jalal Almashaal2023-02-27Abstract ▼Designing nonlinear control systems is still difficult so many researchers are trying to find
some useful ways and methods to solve this problem. As a result of such research, some methods
have been seen trying to design a good enough control system for nonlinear plants. But a disadvantage
of these methods is the complexity, so it created a need to compare some methods to determine
which one is the easiest method to design a control system for nonlinear plants. It was
found a way to compare two methods, which is comparing the regions of initial conditions of the
systems which are designed using these methods. Two analytical nonlinear control systems design
methods are compared on the example of the design control systems mobile robots. The algebraic
polynomial-matrix method uses a quasilinear model, and the feedback linearization method uses
particular feedback. Both considered methods give a bounded domain of equilibrium attraction,
therefore the obtained control systems can be operated only with bounded initial conditions.
The numerical example of designing the control systems for one object by these methods and the
estimates of the attraction areas of the system’s equilibriums of these systems are given in the
paper. As a result of this paper, it was found that using the algebraic polynomial-matrix method
will get a bigger cross section of initial conditions of the plant’s variable than the same cross section
which is given by the feedback linearization method. -
TEXT VECTORIZATION USING DATA MINING METHODS
Ali Mahmoud Mansour , Juman Hussain Mohammad, Y. A. Kravchenko2021-07-18Abstract ▼In the text mining tasks, textual representation should be not only efficient but also interpretable,
as this enables an understanding of the operational logic underlying the data mining
models. Traditional text vectorization methods such as TF-IDF and bag-of-words are effective and
characterized by intuitive interpretability, but suffer from the «curse of dimensionality», and they
are unable to capture the meanings of words. On the other hand, modern distributed methods effectively
capture the hidden semantics, but they are computationally intensive, time-consuming,
and uninterpretable. This article proposes a new text vectorization method called Bag of weighted
Concepts BoWC that presents a document according to the concepts’ information it contains. The
proposed method creates concepts by clustering word vectors (i.e. word embedding) then uses the
frequencies of these concept clusters to represent document vectors. To enrich the resulted document
representation, a new modified weighting function is proposed for weighting concepts based
on statistics extracted from word embedding information. The generated vectors are characterized
by interpretability, low dimensionality, high accuracy, and low computational costs when used in
data mining tasks. The proposed method has been tested on five different benchmark datasets in
two data mining tasks; document clustering and classification, and compared with several baselines,
including Bag-of-words, TF-IDF, Averaged GloVe, Bag-of-Concepts, and VLAC. The results
indicate that BoWC outperforms most baselines and gives 7 % better accuracy on average








