Search
Search Results
-
METHOD FOR SEARCHING SEQUENTIAL PATTERNS OF USER'S BEHAVIOR ON THE INTERNET
V.V. Kureychik, V. V. Bova, Y.A. Kravchenko2020-11-22Abstract ▼One of the important tasks of data mining is to isolate patterns and detect related events in
sequential data based on the analysis of sequential patterns. The article examines the possibility of
using sequential patterns to analyze the events of search and cognitive activity of users when interacting
with Internet resources of an open information and educational environment. Searching
for sequential patterns is a complex computational task whose goal is to retrieve all frequent sequences
representing potential relationships within elements from a transactional database of
sequences of search activity events with a given minimum support. To solve it, the article proposes
a method for searching for patterns in sequences of events to detect hidden patterns that indicate
possible levels of vulnerability when performing information search tasks in the Internet space.
A mathematical model of user behavior in a search session based on the theory of sequential patterns
is described. To improve the computational efficiency of the method, a modified algorithm
for generating sequential patterns has been developed, at the first stage of which AprioriAll is
performed, which forms frequent candidate sequences of all possible lengths, and at the second
stage, a genetic algorithm for optimizing the input parameters of the feature space of the generated
set to search for maximum patterns. A series of computational experiments were carried out on
test data from the MSNBC corpus, the SPMF open source data mining library. The comparative
analysis was carried out with the VMSP and GSP algorithms. The research results confirmed the
efficiency of the search for maximum sequential patterns by the proposed algorithm in terms of the
execution time and the number of extracted patterns. The results of the experimental studies of the
method showed that to increase the stability and accuracy of the work, the sample size obtained as
a result of the GA operation will reduce the required number of scans of the pattern database,
providing acceptable computational costs comparable to the VMSP algorithm and the GSP algorithm
that exceeds the search time for sequential patterns. an average of more than 150 %. -
BASIC APPROACHES TO EXTRACTING TEXTUAL INFORMATION (OVERVIEW)
V.V. Kureichik, P.S. Gerasimenko2024-10-08Abstract ▼This article is devoted to the review of known and modern approaches, methods and algorithms of
full-text search. A brief history of the solution of the problem of search in unstructured text data, its development
and relevance are described. The main task of search in text data is formulated. The definition of
the database index is given. The target function of the search information system is defined in general
terms and possible compromise variations of its parameters when solving various applied problems are
described. A generalized architecture of a modern search information system is given with the division of
the search problem into two phases: the primary extraction of relevant records and their subsequent ranking
to form the final search results. The article provides basic descriptions of the main algorithms and
methods of full-text search, such as: search by terms (logical search), search using trees and their varieties
(B-trees, UB-trees, tries), search based on n-grams (including search based on frequency representation),
use of the vector space model (VSM), search based on an inverted (reverse) index, search using the apparatus of fuzzy logic and bioinspired methods. The main advantages and disadvantages of these methods
are given, their applicability in various conditions is described, and possible methods for optimizing
the search for text data to improve the accuracy, speed of search and efficiency of resource use are considered.
Possible promising directions in the field of solving the problem of primary information extraction
are presented. Some methods for determining the similarity of text records for solving the ranking
problem based on the apparatus of fuzzy logic are given. The article touches upon the issues of increasing
the relevance of primary extraction using artificial intelligence methods, neural networks, fuzzy logic and
bioinspired methods, in particular methods for expanding the search query and/or expanding the processed
text records. The influence of the boundary conditions of the search system construction on increasing
its efficiency is described. In conclusion, the article summarizes the review and discusses the prospects
for further development of various full-text search methods. -
IMPLICIT THREATS IDENTIFICATION BASED ON ANALYSIS OF USER ACTIVITY ON THE INTERNET SPACE
V. V. Bova , D. Y. Zaporozhets, Y.A. Kravchenko , E. V. Kuliev , V. V. Kureichik , N. A. Lyz2020-10-11Abstract ▼The article is devoted to the problem of identifying implicit information threats of a user's
search activity in the internet space based on an analysis of his activity in the course of this interaction.
The use of knowledge stored in the Internet space for the implementation of criminal intentions
poses a threat to the whole society. Identifying malicious intent in the users’ actions of the
global information network is not always a trivial task. The proven technologies for analyzing the
context of user interests fail in the case of cautious and competent actions of attackers who do not
explicitly demonstrate the goal they are pursuing. The paper analyzes the threats associated with
certain scenarios for the implementation of search procedures that manifest themselves in search
activities. Criteria of inefficient and effective search scenarios estimation are described. Among
the signs indicating the possibility of a threat, the following main ones are highlighted: avoiding
solving the problem in aimless navigation or attractive resources, superficial search, lack of
meaningful immersion in solving the search problem, and chaotic actions during the search.
To determine the presence of adverse signs, a system of indicators is built. The features of an effective
scenario for organizing a search in the Internet space are formulated, options for the presence
of implicit threats for a similar situation are described.An approach for identification the
described threats is presented taking into account the specified criteria for evaluating various
scenarios of user behavior in the global information space. A machine learning algorithm has
been developed to identify problem scenarios by comparing with key behavioral patterns. The
software implementation of the subsystem for identifying information threats has been created,
experimental studies have been conducted to confirm the effectiveness of the subsystem. Experimental
studies were carried out on the basis of processing open data from social networks, as well
as using analysis of user search activity in the university corporate information environment.








