Search
Search Results
-
CLASSIFICATION OF PROCESSING NODES IN BIG DATA SYSTEMS ACCORDING TO THE ZERO TRUST APPROACH
М.А. Poltavtseva , D. V. Ivanov55-622025-07-24Abstract ▼Data cybersecurity is one of the most important factors for the successful implementation of the national project ‘Data Economy and Digital Transformation of the State’. The challenges of building secure big data systems lie in their heterogeneous nature, large number of heterogeneous tools, high connectivity and high trust between distributed components. Reducing the internal trust and reducing the attack surface according to the zero-trust approach is necessary to increase the security of such systems with the least impact on their performance. The aim of the paper is to create a method for dynamic classification of nodes and data processing components in heterogeneous big data systems based on the application of different approaches to trust reduction with respect to the objects realising the information processing process. The paper considers the zero trust approach as applied to the class of systems under study, as well as the task of extended implementation of the principle of minimum privilege to reduce the attack surface. The authors present a classification of nodes - handlers based on their operations with data, unified according to the previously developed conceptual data model. A comparison of nodes and security methods applied to them based on the need for access to semantics and data components to perform operations is proposed. Based on this classification, a method of dynamic node type determination during system operation is developed for situations of changing component composition of a big data processing system, typical for multi-component distributed highly loaded systems. The results of the work are a part of the complex consistency approach to the construction of secure big data processing systems.
-
ALGORITHM OF ENSURING THE SECURITY OF CONFIDENTIAL DATA OF THE MEDICAL INFORMATION SYSTEM FOR STORAGE AND PROCESSING OF EXAMINATION RESULTS
L.K. Babenko, A.S. Shumilin, D.M. Alekseev2021-01-19Abstract ▼The objectives of the study are to develop and assess the effectiveness of the structure of a
cloud platform for storing, processing and organizing medical data, determining a method of protection,
in particular, ensuring confidentiality when transferring and storing examination results.
To achieve this goal, the tasks of analyzing existing models of information processes and structures
in the subject area are being solved, the features of the means for accumulating and processing medical data stored in electronic information systems for patient registration, the architecture
of a cloud platform for distributed data storage and an algorithm for ensuring the safety of
medical data stored in the cloud are being developed. the platform in electronic form in the form
of initial physiological signals (EEG, ECG, EMG, EOG, etc.) recorded during patient examinations;
an integrated cloud platform for distributed storage, analysis and systematization of medical
data and a security system using the developed protection method are being created; the effectiveness
of the proposed algorithm for protecting confidential medical information is analyzed in the
context of integration into the developed cloud platform. The proposed method for protecting a
medical information system involves the use of an original DICOM file and subsequently a converted
PNG image, which is subjected to a pixel encryption algorithm. An algorithm based on
chaos theory is used to encrypt the image. The capabilities of chaos systems can significantly increase
productivity. Hierarchical division of data streams into levels and standardization of data
transfer protocols, as well as their storage formats, allow to form a universal, flexible and reliable
medical information system. The proposed architecture has the ability to integrate into existing
medical systems. In the course of the work, it was found that the considered protection method is
an effective way to ensure the confidentiality of medical system data. -
ALGORITHM OF PROTECTING CONFIDENTIAL DATA IN THE CLOUD MEDICAL INFORMATION SYSTEM
L.K. Babenko, A.S. Shumilin, D.M. Alekseev2021-12-24Abstract ▼The aim of the work is the development and implementation of the architecture of a cloud
storage system, systematization and processing of survey results (for example, EEG) and an algorithm
for ensuring the protection of confidential data based on a completely homomorphic cryptosystem.
The object of the research is the technologies of storage, transmission, processing and
protection of confidential information in distributed medical information systems. The architecture
of a cloud platform for distributed storage, processing, systematization and protection of confidential
data (results of medical examinations) has been developed, which makes it possible to interact
with various medical information systems and diagnostic hardware in order to generate big data.
An algorithm has been developed to ensure the safety of medical data stored in a cloud platform in electronic form, recorded during patient examinations in order to calculate the average value for
each of the brain activity rhythms (based on the results of a series of examinations over a long
period of time) using a fully homomorphic encryption algorithm. Based on the test results (analysis
of the execution time of such operations as: encryption, decryption, addition, multiplication,
signal-to-noise ratio of ciphertext to plaintext), the optimal algorithm. According to the results of
the work, it is shown that the fully homomorphic encryption scheme CKKS is the most effective,
especially in the context of the criticality of the requirements for a high level of security of confidential
data, which determines the choice of this scheme for the implementation of the algorithm
proposed in this work. -
DEVELOPMENT OF ALGORITHMS OF INTELLIGENT SERVICE FOR INFORMATION SEARCH AND MONITORING
M. S. Anferova, A. M. Belevtsev2021-08-11Abstract ▼This paper describes the problem of strategic analysis and the choice of directions for the development
of an innovative enterprise in the conditions of transition to the 6th technological order and
industry 4.0. In these conditions, search and analytical processing of information cannot be fully performed
without the use of automated information and analytical systems, including those based on artificial
intelligence. During the analysis, the main priority functions that the developed services should
provide were identified. The main difficulties in the development of these services are identified, such as:
pre-processing of data and automated checking of the relevance of databases. To effectively solve thetasks set, the intelligent monitoring and information retrieval service should use an integrated approach,
taking into account the effectiveness of applying methods for individual subtasks, and ensure high efficiency
of implementing all stages of the intelligent monitoring procedure. In this regard, this paper describes
not only the development of a general intelligent search algorithm, but also individual block
algorithms necessary to ensure the priority functions of the service being developed. The paper presents
the following algorithms: an information search algorithm necessary to solve the problem of full-text
search of documents within the database of information resources of the information and analytical
complex; an algorithm for the procedure for entering new documents; an algorithm for pre-processing
data that includes stemming and removing punctuation marks for subsequent text analysis; an algorithm
for evaluating the ranking and relevance of information, including vectorization of documents; an algorithm
for clustering information search results based on the Kohonen neural network; the algorithm for
checking the relevance of information is to check whether the local copy of the document corresponds to
the current version on the source's web resource. The Python programming language for the implementation
of the presented algorithm is proposed and justified. The system provides automated continuous
monitoring with a high frequency of sending a request without the participation of an operator, which
will increase the quality and efficiency of information search in conditions of a large volume of unstructured
information. -
ANALYSIS OF REQUIREMENTS AND DEVELOPMENT OF ALGORITHMS FOR INTELLIGENT MONITORING SERVICES
М.S. Anferova, А.М. Belevtsev2022-08-09Abstract ▼The paper considers the problems of strategic analysis and the choice of directions for the development
of innovative enterprises in the conditions of transition to the 6th technological order and industry
4.0. The main levels of analysis are determined. The objectives of the strategic analysis are outlined
based on the scale of the research being conducted. The analysis tasks are highlighted, the solution of
which will allow achieving the set goals. The complexity of solving global monitoring tasks, which are
caused by a large volume of heterogeneous and unstructured information, is shown. In these conditions,
thematic search and analytical processing of information cannot be performed without the use of automated information and analytical systems and the creation of search services based on artificial intelligence.
A general monitoring procedure is proposed. The main stages of monitoring technological trends
are defined, the tasks to be solved within a specific stage and the planned result are shown. Based on the
general monitoring procedure, the main priority functions that the developed services should have are
determined. As well as the problems of their development and structuring of the received information in
the form of information objects and clustering of documents. In contrast to the well-known global monitoring
systems, in which the search is based on indicators: an increase in the use of keywords, an increase
in the number of new authors, quoting works from related fields. Algorithms are proposed that
provide the definition of reference topics, assessment of ranking and relevance of information. The description
of the algorithms is given on the example of creating a summary information table, with the
help of which the interrelationships of documents of scientific and technological development in each
direction of monitoring and the search for specific documents in the database are formed. The construction
of search services based on the presented algorithms will ensure the allocation of reference topics
of documents, provide more reliable results of clustering of unstructured information and the formation
of scientific and technological trends in information and analytical complexes. To implement the algorithm,
it is proposed to use the Python programming language. The implementation of these algorithms
will improve the quality and efficiency of information retrieval in conditions of a large volume of unstructured
information. -
ESTIMATING THE EFFECTIVENESS OF THE METHOD FOR SEARCHING THE ASSOCIATIVE RULES FOR THE TASKS OF PROCESSING BIG DATA
V. V. Bova, E.V. Kuliev, S.N. Scheglov2020-07-20Abstract ▼The modern databases have significant volume and consist of large masses of information.
One of the popular methods of knowledge identification in terms of tasks of analysis and processing
of large data volumes is composed of the algorithms for searching the associative rules.
The paper solves the problem of building the bases of associative rules for the analysis of the unstructured
large data volumes on the basis of searching different regularities considering the importance
of their characteristics. The authors propose the method for synthesizing the bases and
building the transaction database to calculate the threshold values of support and application of
criteria of estimating implicit associations. This allows us to extract repeated and implicit associative
rules. To improve the computational effectiveness of extracting the associative rules, the paper
applies the genetic algorithm for optimization of input parameters of the characteristic searching
space. The developed method shortens the time of rules extraction, reduces the number of generated
common rules, and avoid the resource-consuming procedure of pre-processing the synthesized
rule base. The authors developed the program and algorithmic module to carry out the experimental
research of the proposed method for synthesizing the associative rules on the basis of filtering
the input parameters of the search model for solving the tasks of processing the unstructured
data. The experiments conducted on the test transaction bases allow us to clarify the theoretical
estimations of time complexity of the proposed method that used the genetic algorithm to calculate
the weighed support of the set of rules considering the assessment of a priori informative content
of the characteristics included in the dataset. The time complexity of the developed method is estimated
as О(I2). The comparative analysis is performed using the test data of the Retail Data
with the algorithms Apriori and Frequent Pattern-Growth. The results have proven the effectiveness
of the search method on big sets of transactions. The method allows us to reduce the cardinal
of an irredundant set of extracted associative rules in more than 40% in comparison with the popular
algorithms. The experiments have shown that the method can be effective for the tasks of
knowledge discovery in terms of processing large volumes of data. -
ANALYTICAL REVIEW OF THE DECISION TREE ALGORITHM IN DATA INTELLIGENCE TECHNOLOGY
E.V. Kuliev, V.A. Semenov, A.V. Kotelva, S.V. Ignateva2022-05-26Abstract ▼The decision algorithm is the preferred filtering algorithm in data mining technology, and
its results are usually chosen in the form of "if-then" rules. Algorithm C4.5 is one of the decision
algorithms that takes advantage of the ease of understanding and increasing importance, and also
takes advantage of the advanced information rate gain of its advanced ID3 algorithm. After the
theoretical analysis of the information, the algorithm C4.5 is selected to analyze the results of
performance appraisal, and enterprise performance appraisal decisions by collecting data, preprocessing
data, calculating information gain and determining selection parameters. The system isdeveloped in B/S architecture, an R&D project management platform that can perform evaluation
analysis with decision analysis results evaluation tools and web coverage. The system includes
information storage, task management, reporting, receipt and presentation control, information
visualization and other functions of the management information system functions. They can realize
project management functions, such as creating and managing a project, flow tasks, filling and
managing information about functions, creating a performance evaluation system, creating reports
of various sizes, building management. decision decision algorithm as the core technology,
the system acquires scientific significant project management information with high data accuracy,
and realizes visualization, which can help the enterprise to have a good management system in
large areas. Task management, reporting, audit control, information visualization and other functions
of the system's management reporting management functions are included. -
USE OF PARALLEL COMPUTING FOR SECURITY METHOD IMPLEMENTATION BASED ON THE SHAMIR SCHEME IN A MEDICAL INFORMATION SYSTEM
L. K. Babenko , A.S. Shumilin2023-10-23Abstract ▼Medical information systems currently are becoming the most popular tools for processing,
storing, organizing, and transmitting patient medical data. Medical examinations can be presented
in the form of files in various formats and vary greatly in terms of size (from a few bytes to hundreds
of gigabytes). For example, some binary files are small and lightweight because they contain
only doctors' conclusions in the form of a text description. However, records of night video
monitoring of a patient or DICOM files of human organs CT scans containing several hundred
slices, can reach hundreds of gigabytes in size. Accordingly, large files require significant computing
resources when transferred from server to server. In addition, when using the security method,
which is an algorithm of a secret sharing (medical output file) according to the Shamir sharing
scheme, operations to split the secret into parts and merge the parts together may take longer in
serial operation than in parallel way. Therefore, it seems possible to speed up the processing of
big data without reducing the level of security. The main purpose of the work is to confirm the
hypothesis of reducing time to perform the operation of splitting and merging parts of a secret
based on parallel computing tools withing implementing the security method according to the
Shamir secret sharing scheme in a medical information system. The object of the study is a security
method developed by the author for implementation in the information security subsystems of a
medical information system. As part of the study, author analyzed the most effective tools for parallelizing
processes (like MPI and OpenMP). MPI has been used as a tool as much more suitable
for the current purpose. Moreover, several waves of experiments have been run (analysis of time
depending on the number of parallel streams and the number of characters contained in the
DICOM file) and allowed us to prove a concept of parallelizing the secret exchange algorithm
based on the Shamir scheme, achieving almost linear acceleration using the MPI library








