Search
Search Results
-
METHOD FOR DETECTING FEATURE POINTS OF AN IMAGE USING A SIGN REPRESENTATIONS
A. N. Karkishchenko, V. B. Mnukhin2020-11-22Abstract ▼The aim of the study is to develop a method for detecting feature points of a digital image
that is stable with respect to a certain class of brightness transformations. The need for such a
method is due to the needs of detecting feature points of images in video surveillance systems and
face recognition, often working in a changing light environment. A feature of the proposed method
that distinguishes it from a number of well-known approaches to the problem of distinguishing
characteristic points is the use of the so-called sign representation of images. In contrast to the
usual defining of a digital image by a discrete brightness function, with a sign representation, the
image is set in the form of an oriented graph corresponding to the binary relation of the increase
in brightness on a set of pixels. Thus, the sign representation determines not a single image, but a
set of images, the brightness functions of which are connected by strictly monotonic brightness
transformations. It is this property of the sign representation that determines its effectiveness for
solving the problems caused by the goal set above. A feature of the method under consideration is
a special approach to the interpretation of the characteristic points of the image. This concept in
image processing theory is not strictly defined; we can say that the characteristic point is characterized
by increased "complexity" of the image structure in its vicinity. Since the sign representation
of the image can be represented in the form of a directed graph, in this paper, to evaluate the
complexity measure of the local neighborhood of its vertices, it is proposed to use the ranking
method known in the spectral theory of graphs based on the Perron-Frobenius theorem. Its essence
lies in the fact that the value of the component of the so-called Perron eigenvector of the
adjacency matrix of this graph acts as a measure of the complexity of the vertex. To conduct experimental
studies of the proposed approach, a set of programs was developed, the results of
which confirm the efficiency of the method and demonstrate that with its help it is possible to obtain
results close to the expected ones on model examples. The paper also offers a number of recommendations
on the use of this method. -
THE APPLICATION OF COMPLEX DESCRIPTORS IN SOLVING A SLAM TASK
V. P. Noskov, А. N. Kuryanov2022-04-21Abstract ▼The actual problem of determining all six coordinates (three linear and three angular) of the
current position of a mobile robot (unmanned aerial vehicle) from video rangefinder images of the
external environment (volumetric colored point clouds) formed by an onboard integrated vision
system built on the basis of a 3D rangefinder sensor (lidar) and a color video camera while moving
(flying) in an unknown environment is considered. An algorithm of video navigation based on
the use of complexed (video-rangefinder) descriptors is proposed, for the description of which
visual and geometric parameters are used. The rules for the formation of a complex descriptor are
formulated, which ensure the allocation of special (central) points of the descriptor using the
Sobel operator and the calculation of brightness and geometric parameters in its local area. The
addition of the brightness parameters of the descriptor provided by the video camera with the geometric
parameters provided by the rangefinder sensor removes the problem of invariance of the
descriptor to the scale and thereby significantly reduces the complexity of calculations when selecting
it. The rules for finding complexed descriptors corresponding to each other in a sequence
of complexed images are described, based on calculating the difference in brightness and geometric
parameters of the compared descriptors. The estimation of the error in solving the navigation
problem using the integrated descriptors was performed depending on the error of the sensors of
the vision system and the geometric dimensions of the descriptor. By constructing histograms of
the solution of the navigation problem for each coordinate of the control object for all pairs of
descriptors corresponding to each other, a statistically stable high reliability of the solution of the
complete navigation problem has been achieved. At the same time, the error in solving the navigation
task turned out to be an order of magnitude smaller than the error in the formation of complex
images by the technical vision system. The use of complex descriptors made it possible, with a
relatively small amount of calculations, to solve the complete navigation problem with acceptable
accuracy, which provides a solution to the SLAM problem on the onboard computations at the
pace of movement of the control object. The effectiveness of the proposed algorithmic and developed
software and hardware is confirmed by field experiments conducted in real conditions of
various environments. -
IMAGE MATCHING SYSTEM WITH USING INTUITIONISTIC FUZZY SETS
К.I. Morev286-2982026-04-29Abstract ▼This paper presents a fully learnable system for solving the problem of matching two images. All the main elements of the system are trainable, i.e. their final form corresponds to the target dataset on which the training was carried out. The fact that the system is trainable, the methods used in training and the architecture of the system allow using the system to solve a large number of various computer vision problems. The system consists of a convolutional neural network that serves both to extract key points and their descriptors, as well as a trainable matcher of the extracted key points based on their description and mutual arrangement in the observed scene. The used convolutional neural network processes full-size images and calculates both the location of interest points with pixel accuracy and the descriptors associated with them in a single forward pass. Matching key points is a separate step and is performed after the forward pass of the neural network. In the process of training the model for calculating the positions of key points and their descriptors, a method for forming a training sample is used, called homographic adaptation - an approach that helps to increase the repeatability and accuracy of detecting key points. The process of training the feature point detection model consists of obtaining new weights in the process of additional training of the base detector, which represents the initialization weights of the model. The final feature point detection model, trained on the universal MS-COCO image set using homographic adaptation, repeatedly outperforms the original base detector in terms of the number, reliability and repeatability of feature points, and also outperforms any other traditional corner detector based on classical approaches
-
ONBOARD ACTIVE-PULSE UNDERWATER VISION SYSTEM THROUGH THE AIR-WATER BOUNDARY
Y. К. Gruzevich, Y.N. Gordienko, P. S. Alkov, D.V. Volkov, М.S. Khodakovskaya2025-04-27Abstract ▼The objective of this work is to create a system for detecting underwater objects intended for installation
on surface platforms (aircraft or remotely piloted aircraft). Systems of this type can be used to solve
a wide range of problems in various areas of the national economy: searching for rare fish and marine
mammals, determining their migration routes, diagnostics and laying underwater pipelines and fiber optic
cable networks, monitoring seawater pollution, searching for sunken ships and archaeological treasures,
and carrying out rescue operations. To solve this problem, the process of laser radiation propagation to
an object across the air-water interface was described, a rough sea surface was modeled, and a number of
mathematical assumptions and approximations were proposed. In the practical part, a structural diagram
of a laser optical-television active-pulse underwater vision system was developed, including receiving and
transmitting channels, as well as a control device consisting of an image processing unit and a control controller. The receiving channel includes an electron-optical converter of the III+ generation, highly
sensitive in the spectral range of sea water transparency. The main element of the transmitting channel is
a highly efficient pulse laser emitting in the spectral range of sea water transparency. The assembled device
has undergone field tests, as a result of which it became clear that detection and recognition of underwater
targets from an aircraft through the air-water interface using the generated image is possible,
the maximum detection range and recognition of underwater targets of the active-pulse underwater vision
system from an aircraft through the air-water interface is mainly determined by: attenuation of optical
radiation in sea water and the power of the illuminating laser pulse radiation, At the same time, a distinctive
feature of the active-pulse underwater vision system is that the increase in the detection and recognition
range is almost directly proportional to a certain level of laser radiation power, and a further increase
in power leads to an insignificant increase in range -
THE PARALLEL-PIPELINED IMPLEMENTATION OF THE FRACTAL IMAGE COMPRESSION AND DECOMPRESSION FOR RECONFIGURABLE COMPUTING SYSTEMS
M.D. Chekina2021-02-25Abstract ▼Fractal algorithms find an increasing number of areas of application - from computer
graphics to modeling complex physical processes, but their software implementation requires
significant computing power. Fractal image compression is characterized by a high degree of data
compression with good quality of the reconstructed image. The aim of this work is to improve the
performance of reconfigurable computing systems (RCS) when implementing fractal compression
and decompression of images. The paper describes the developed methods of fractal compression
and subsequent decompression of images, implemented in a parallel-pipeline method for RCS. Themain idea of parallel implementation of fractal image compression is reduced to parallel execution
of pairwise comparison of domain and rank blocks. For best performance, the maximum
number of pairs must be compared simultaneously. In the practical implementation of fractal image
compression on the DCS, such critical resources as the number of input channels and the
number of FPGA logical cells are taken into account. For the problem of fractal image compression,
data channels are a critical resource; therefore, the parallel organization of computations is
replaced by parallel-pipeline, after the performance reduction of the parallel computational structure
is performed. Each operand goes into the computational structure sequentially (bit by bit) to
save computational resources and reduce equipment downtime. To store the coefficients of the
iterated functions system encoding the image, a data structure has been introduced that specifies
the relation between the numbers of rank and domain blocks and the corresponding parameters.
For the convenience of subsequent decompression, the elements of the array encoding the compressed
image are ordered by the numbers of the rank blocks, which avoids double indirect addressing
in the computational structure. Applying this approach for parallel-pipeline programs
allows scaling computing structure to plurality programmable logic arrays (FPGAs). A practical
implementation performed on a reconfigurable computer Tertius-2 containing eight FPGAs provides
an acceleration of 15000 times compared to a universal multi-core processor and 18–25
times compared to existing solutions for FPGAs. The implementation of image decompression on a
reconfigurable computer shows an acceleration of 380 times in comparison with the similar implementation
for a multi-core general-purpose processor. -
CLASSIFICATION OF RADAR IMAGES OF MULTI-ROTOR UNMANNED AERIAL VEHICLES USING THE YOLO11 ALGORITHM
V.А. Derkachev171-1802025-07-24Abstract ▼This article discusses a classifier of radar images of unmanned aerial vehicles based on a neural network built on the YOLO algorithm version 11. Solving the problem of detecting and classifying unmanned aerial vehicles has become one of the priority tasks at present. The increase in the number of modifications of unmanned aerial vehicles greatly complicates the use of statistical classification methods, which requires the use of new approaches to solving the classification problem. The development of neural network methods, simultaneously with an increase in the performance of computers for training, on the one hand, and embedded solutions, on the other, allows for the classification of aircraft using radar images in real time. The use of the YOLO11 algorithm allows, in addition to determining the class of the target, to estimate the range to the observed object. The use of radar images is justified due to the fact that visual observation is not always possible due to difficult weather conditions and darkness. To train the neural network, it is proposed to use a set of radar images obtained using the author's model of data generation with an arbitrary configuration of unmanned aerial vehicles. The neural network of the Detection YOLO11s class (9.4 million parameters) was trained on a sample of radar images of two classes, a total of 8192. As a result of training, an accuracy of 0.99 was obtained for classification in 2 classes of objects (on test model data). Tests were conducted using natural data taken using the TI IWR1642 millimeter-range radar system, as a result of which error-free classification of objects on a small sample was achieved
-
MULTI-AGENT SYSTEM USING ARTIFICIAL INTELLIGENCE TO PROCESS IMAGES FROM THE DRONE'S TECHNICAL VISION CAMERAS
А. L. Verevkin , I.E. Josephs , V.V. Misyura , L.S. Verevkina198-2122025-07-24Abstract ▼Multi-agent technology with drones, modern sensors, precise GPS and artificial intelligence, have led to a breakthrough in the field of cyber-physical systems. This article presents a multi-agent system using artificial intelligence to process images from technical vision cameras installed on a drone. A block diagram of a multi-agent system on a drone was developed based on an effective and simple platform taken from the ARRISE 410 octocopter – an agricultural sprayer drone with: intelligent control system; omnidirectional digital microwave radar; 6-axis high-precision accelerometer; electronic level for measuring tilt; real-time optical camera 1 with a first-person view; control panel equipped with the latest Light Bridge 2 signal transmission system; remote control has a design protected from dust and water. The kit must be supplemented with: hyperspectral HS - camera for scanning, its power module and the ability to interface with the ARRISE 410 drone systems, an information compression module. Model for studying the throughput on the DJI Agras T20 hexacopter DJI Agras T20, MikrotikRB411 5G network card, Raspberry Pi 3 microcomputer, 1 Mpix RGB camera, built-in on-board computer Raspberry Pi OV5647 v1.3 and hyperspectral HS - camera 2 Resonon Pika L shoots hyperspectral data with 281 spectral bands with spectral wavelengths from 400 to 1000 nm and a spatial resolution of 900 hyperspectral pixels per image line. The article solves the problem of experimentally and computationally determining the required compression of information obtained from hyperspectral and optical range cameras with transmission through a telecom operator and the Internet for image processing by an artificial Internet
-
ALGORITHM FOR PRE-PROCESSING VIDEO IMAGES TO INCREASE THE ACCURACY OF SMALL OBJECT DETECTION
V.V. Kovalev, N.E. Sergeev2021-12-24Abstract ▼Recognition of certain patterns in video images captured by a camera is carried out using
training methods based on convolutional neural networks. The larger the number of images with
multiple features and the more diverse the training sample of video images, the better the convolutional
neural networks extract features from the sequence of video images that were not included in
the training sample. This is a consequence of increasing the accuracy of detecting visual images on
video images containing features of target images. However, there are limitations in improving the
detection performance when the size of the image to be detected is much smaller than the background
area, or when the image is described with little information. To solve problems of this kind, the authors
of the article have developed an algorithm for the spatio-temporal integration of information
about the movement of dynamic images. The algorithm processes a fixed number of video images at
certain points in time and extracts new independent signs of motion of dynamic images based on
space-time processing of video images. Further, it combines new local motion features with the original
video image features. This allows you to add a motion feature of dynamic images while preserving
the original image features that describe static images. Areas of the video image that characterize
the motion feature are displayed in a «color» cluster. The use of pre-processing is aimed at improving the accuracy of pattern detection, provided there are dynamic visual images on a static background.
If the camera is in scan mode, a static background can be provided with a video stabilizer.
Experimentally, estimates of integral criteria for the accuracy of detection neural network algorithms
have been obtained, showing an increase in the accuracy of detecting visual images using
the algorithm for spatial-temporal integration of motion information. -
COMPARATIVE ANALYSIS OF TWO FILTERING METHODS TO ELIMINATE NOISE IN AN IMAGE OF DIFFERENT DEGREES OF NOISE
K.O. Sever, I.I. Turulin, D.A. Guzhva2021-08-11Abstract ▼In modern photography and video technology, any image in the process of its creation is
distorted by various types of noise. There are various types of noise, but in practice, impulsive and
Gaussian noise models are the most common. Attenuation of the effect of noise is achieved by filtering.
At the moment, there is no universal filter that suppresses noise data at various intens ities
of distortion. Therefore, an important aspect is to determine the field of application of each
type of filter when suppressing noise in the image and creating a filter, consisting of a combination
of different filtering methods for optimal image cleaning. The article presents a comparative
analysis of median filtering and Wiener filtering to eliminate impulse and Gaussian noise in
the image with different degrees of noise. For modeling, we used one image, separately distorted
by impulse and separately by Gaussian noise with pixel distortion probabilities from 1% to
99% inclusive. Filtration was performed with windows equal to 3x3 and 5x5. As a result, we
obtained numerical estimates of the image filtering quality based on the peak signal-to-noise
ratio (PSNR). On the basis of the data obtained, the application of the investigated filters, their
modifications, advantages and disadvantages were analyzed, as well as recommendations for
their use were given. As a result of a comparative analysis of the studied types of filtering for
noisy images, it was found that the median filter with a 3x3 window copes better with image
cleaning from low-intensity impulse noise and with a 5x5 window - with image cleaning with an
average noise intensity. Also, the median filter does a better job of filtering out Waussian noise
at its medium and high rms deviations. The Wiener filter with 3x3 and 5x5 windows better fi lters
Gaussian noise at small values of its root-mean-square deviation. Also, the Wiener filter
copes better with impulse noise with relatively high noise power. -
ERROR ESTIMATION FOR MULTIPLE COMPARISON OF NOISY IMAGES
A.N. Karkishchenko, V. B. Mnukhin2021-07-18Abstract ▼The aim of this work is to study the effect of noise on the image on the quality of comparison of a
finite set of images of the same shape and size. This task inevitably arises when analyzing scenes, detecting
individual objects, detecting symmetry, etc. The noise factor must be taken into account, since the
difference between digital objects can be caused not only by the mismatch of the compared images of
real objects, but also by distortions due to noise, which is almost always takes place. This differenceturns out to be proportional to the level of the noise component. The main result of this article is an
analytical estimate for the probability of a given level of error, which may arise in the multiple comparison
of a finite set of commensurate digital images. This estimate is based on a low-level comparison,
which is a pixel-by-pixel calculation of image differences using the Euclidean metric. In this case, a
standard assumption is made about the independent normal noise of image intensities with zero mathematical
expectation and a priori established standard deviation in each pixel. The evidence presented in
the article allows us to assert that the obtained estimate should be regarded as sufficiently "cautious"
and it can be expected that in reality the scatter of the measure caused by noise in the image will be
significantly less than the theoretically found boundary. The estimates obtained in this work are also
useful for detecting various types of symmetry in images, which, as a rule, lead to the need to calculate
the difference of an arbitrary number of commensurate digital areas. In addition, they can be used as
theoretically grounded threshold values in tasks requiring a decision on the coincidence or difference of
images. Such threshold values inevitably appear at various stages of processing noisy images, and the
question of their specific values, as a rule, remains open; at best, heuristic considerations are proposed
for their selection. -
PROSPECTS OF MALE-CLASS UAVS USING FOR THE HUGE TERRITORIES AERIAL SURVEY
А. М. Fedulin, D.M. Driagin2021-04-04Abstract ▼The aim of the study is to estimate the MALE-class (Medium Altitude Long Endurance) UAV
(Unmanned Air Vehicles) using possibility to solve the problem of regular aerial survey of huge
areas relative to other means used for this, such as: small-sized UAVs, satellite remote sensing
and manned aircrafts. Considered is the issue of practical construction of onboard computer vision
system based on a UAV “Orion” wit a ta eoff weig t of more t an a ton, w ic pro ides
aerial photography in the visible and near infrared range and airborne laser scanning of the underlying
surface with automatic processing of the received data on board in near real-time mode
detecting the changes occurred since the previous survey. It has been determined the key components
of the computer vision system both the hardware and software platform required highperformance
computing and big-data storage. It has been presented a promising architecture,
given estimates for its search performance, weight and power consumption, determined the typical
flight altitude, which provides the input data spatial resolution, which is necessary for objectoriented
change detection algorithms, based on a convolutional neural networks machine learning.
It has been proposed organizational and technical solutions to speed up the data processing
cycle, taking into account the requirements of the legislation regarding the declassification of
aerial survey data. The results obtained confirm that after the issuance of the Orion UAV by the
Federal Air Transport Agency of the aircraft type certificate, which gives the right to perform
commercial flights in the shared airspace of the Russian Federation, it will be possible to implement
an aerial survey complex of high productivity and degree of autonomy using cut of the edge
CV & ML technologies. It seems the tactical, technical and economic capabilities of which proposed
will be orders of magnitude superior to the currently existing solutions especially for hardto-
reach regions. -
IMPROVING REAL PERFORMANCE OF RCS WHEN SOLVING DIGITAL IMAGE PROCESSING TASKS USING FAST FOURIER TRANSFORM
A.V. Chkan2021-02-25Abstract ▼The article discusses the issues of digital processing of images of large dimensions in real
time using reconfigurable computer systems (RCS) on the base of programmable logic arrays
(FPGAs). RCS belongs to the class of high-performance multiprocessor computing systems that
have a programmable architecture that allows configuring the structure of a computer system and
optimally adjusting it to the algorithms of the solved task. At the same time, optimization of the
computational structure of the task reduced to the development and implementation of parallel
algorithms corresponding to the specifics of the RCS architecture used. All this allows to effectively
using RCS to solve a wide class of digital signal processing tasks. Offered are methods of increasing
specific and real performance of RCS when solving digital image processing problems
using fast Fourier transform (FFT). Using the example of a procedure for filtering images in the
frequency domain, the main computational steps and methods for optimizing them based on the
properties of the FFT algorithm are discussed. The use of optimization allows to significantly reducing
both the amount of computation and the amount of hardware resources of the FPGA andincrease the performance of RCS for image processing tasks. The FPGA resources freed because
of the optimization of the computational structure can be uses to further parallelize calculations
and accelerate the processing of incoming data. The advantages of presenting data in fixed-point
format when performing calculations on RCS are showed. The use of a fixed point allows not only
to increase the specific and real performance of a computer system compared to a floating point
due to the properties of the format, but also to use arbitrary data bit capacity, which is relevant for
most digital signal processing tasks. The solution to the problem of overflow of the bit grid when
using the fixed-point format using data bit scaling is discussed. -
RESTORATION OF DEFECTS AND BLIND ZONE ON IMAGES OF UNDERLYING SURFACE FOR ONBOARD RADAR SYSTEMS OF MAPPING BASED ON DOPPLER BEAM SHARPENING
R.R. Ibadov, S.R. Ibadov, V.P. Fedosov2021-02-13Abstract ▼The problem of forming a radar image (RI) of the earth's surface in real time remains one of
the most urgent in solving radio imaging problems, despite the appearance of a large number of
publications in this area, reflecting a whole range of new methods and algorithms for processing
trajectory signals in order to improve the quality of images. The main goal in the formation of
radar images is to achieve the maximum resolution and image quality under real constraints associated
with the drift of the parameters of the received trajectory signal (synthesis time), measurement
inaccuracy and variability of flight characteristics (speed, acceleration, flight trajectory),
exposure to a wide range of noise and interference, both external and internal, against the background
of a low-power received signal from remote radio reflectors (energy resources). The article
investigates an algorithm for constructing and restoring images of the underlying surface and
develops its software implementation. The effectiveness of the new approach is shown using several examples for various areas of the underlying surface with a blind spot. The subject of the research
is methods and algorithms for constructing a terrain map and reconstructing lost image
areas. The research object is a set of test images. The result of the research is the development of a
method for image restoration in order to restore the lost area. The novelty of the work is an algorithm
that improves the quality of image restoration based on a neural network. The results obtained
make it possible to restore the areas. Evaluation of the efficiency of the image restoration
method was carried out using a statistical criterion - the root mean square error of the processing
result from the true image. As a result of solving the tasks, we can draw conclusions: A method
was developed for constructing and restoring images of the underlying surface based on the
search for similar blocks with their subsequent combining by a neural network. Analysis of the
results of the study showed that the proposed method improves the quality of image reconstruction. -
DEVELOPMENT AND ANALYSE VISUAL NAVIGATION SYSTEM FOR AIR AND GROUND-BASED ROBOTS
V. P. Noskov, Y. S. Barichev, О.P. Goydin, А.N. Kuryanov2025-04-27Abstract ▼The work is devoted to solving urgent problems of joint autonomous visual navigation for air and
ground-based robots in urbanized environments. These environments are highly demanded for special
operations, including dense urban areas and buildings, where the use of traditional remote control devices
is limited due to the presence of shielded areas. The proposed solution addresses group navigation tasks
based on data from onboard vision systems during operational reconnaissance of the working area by an
unmanned aerial vehicle (UAV). The results of this reconnaissance enable autonomous movement and
flight, both for individual heterogeneous robotic systems and for groups.The navigation algorithms are
based on methods for extracting a horizontal reference surface and horizontal sections of the external
environment from a volumetric point cloud generated by an onboard lidar. These methods allow for the
precise and rapid determination of all six coordinates of the control object. Cases where the navigation
task cannot be fully solved due to specific environmental characteristics are also considered. To address these challenges, methods are proposed to enhance lidar rangefinding data by integrating video camera
data. An accuracy assessment of the video navigation solutions is provided, obtained through mathematical
modeling of the external environment and the generation of video data. To ensure safe autonomous
flight and movement of robotic systems in urban environments, methods for reducing video navigation
errors are proposed. These methods utilize a specially designed bank of reference images with known
coordinates of their formation. The effectiveness of the applied methods and the proposed video navigation
algorithms is confirmed by experimental studies of the corresponding software and hardware in real
urbanized environments -
CLASSIFIER OF IMAGES OF AGRICULTURAL CROPS SEEDS USING A CONVOLUTION NEURAL NETWORK
V. A. Derkachev, V. V. Bakhchevnikov, A. N. Bakumenko2020-11-22Abstract ▼This article discusses the creation of a convolutional neural network architecture that classifies
images of crops (in particular wheat) for subsequent use in an optical seed separator (photo
separator). Interest in the design of neural networks for classifying images has recently increased
significantly, which is associated both with the development of the theory of deep neural networks
and the increased computing power of desktop computers, as well as the transfer of computing to
graphic processors. The aim of the article is to develop the architecture of a neural network that
allows the separation of the input flow of wheat seeds into two classes: “good” seeds and “bad”
(with defects in shape and color) seeds. The architecture of the resulting neural network is convolutional,
because, unlike a fully connected one, this class of neural networks is within certain limits
immune to changes in the scale and angle of rotation of objects in the input data. In the work,
for the formation of training, validation and test samples, seed images obtained using a household
camera were used, which negatively affected the results of training and testing the neural network
regarding the possible result of application in a real photo separator. The architecture of the developed
neural network is preliminarily optimized for use on FPGAs, however, in the considered
case, the transition from the values of weighting factors from the data type from a floating point to
an integer type has not been made, which can lead to a decrease in the accuracy of the neural
network, while significantly reducing the amount of resources FPGA. Application of the proposed
architecture allows one to obtain a fairly accurate estimate of classified wheat seeds from verification
and test data sets. -
IMAGE MATCHING USING DIFFERENT KEYPOINTS TYPES
K. I. Morev , A.V. Bozhenyuk2020-10-11Abstract ▼The work is devoted to experiments with various methods of selecting special points on images,
followed by their description with a binary descriptor and comparison by a full search method.
This paper actively uses the method of describing the neighborhood of singular points, based
on the construction of a binary string that characterizes changes in the brightness of pixels in the
described neighborhood. The resulting string is obtained by comparing the brightness of pixels
according to a specific template. Today, the use of special points when working with images allows
you to develop applied methods in various areas of computer vision with increased requirements
for working time and resistance to sudden changes in scenes. The paper presents the results
of experiments with special points of various classes, the classification is given in section 1. During
the experiments, methods implemented in the OpenCV library were used. The paper provides
brief descriptions of the methods used in experiments. Section 1 of the paper offers a classification
of modern types of singular points of images and provides a brief description of popular methods
for detecting the described types of singular points. In section 2, the authors give a General description
of methods for working with special image points. Section 3 describes the experiments
that are being carried out with the comparison of special points of different types described by a
single descriptor, and reveals their results. The experiments performed allow us to identify the
strengths and weaknesses of bundles of different types of singular points when comparing them. -
IMAGE RECOGNITION OF AGRICULTURAL CROPS, PLANTS AND FORESTS
I. B. Abbasov, Ratnadeep R. Deshmukh2020-10-11Abstract ▼The paper provides an overview of some studies on the recognition of images of crops,
plants and forests. These image recognition systems use various methods of pre-processing, computer
vision, and deep learning. Recently recognition systems based on mobile devices are increasing,
which increases their availability and wide distribution. The articles on recognition,
classification of fruits and fruits in orchards, the creation of a data bank of these agricultural
products (apples, pears, kiwi) to assess ripening and yield are considered. The works devoted to
the automation of harvesting grain crops are described on the example of the work of a combine
harvester using machine vision. Crop production plays an important role in providing feed for
animal husbandry; articles on the recognition of agricultural plants based on leaf images are
analyzed. Also, by the condition of the leaves of potato bushes, you can determine their disease,
assess the condition of the soil. The work on the development of mobile systems for monitoring and
recognition of the process of growing mushrooms based on the "green house" technology for
farms is presented. Using remote diagnostics, you can analyze and monitor the state of the surface
of land and seas. For remote environmental monitoring of the landscape of the earth's surface,
work is described on the recognition, classification of forests, water resources using hyperspectral
analysis of satellite images. -
REALTIME NEURAL NETWORK ALGORITHM FOR FULL-FRAME MARINE SURFACE OBJECTS RECOGNITION
V.A. Tupikov, V.A. Pavlova, V.A. Bondarenko, N.G. Holod2020-07-10Abstract ▼The article explores modern neural network architectures for the automatic detection and recognition of marine surface objects and obstacles of given classes throughout the full image area, applicable for execution in real or near real time on an optoelectronic vision system to au-tomate and improve the safety of civil marine navigation. A formal statement of the problem of automatic detection of objects on images is given. The state-of-the-art algorithms for detecting objects in images based on use of artificial convolutional neural networks were reviewed, their comparison was made and a reasonable choice was made in favor of the most efficient neuralnetwork architecture in terms of computational complexity to recognition accuracy. The subject area is studied, as well as publicly available databases of surface objects suitable for use in the training of algorithms using artificial neural networks. The article concluded that there is insuffi-cient labeled data for training neural network algorithms, as a result of which the authors inde-pendently collected research images and video sequences, prepared and labeled the collected data containing surface marine objects and other obstacles that represent a navigation hazard for ships. Based on the selected neural network architecture, a new neural network algorithm for automatic full-frame detection and recognition of surface objects was developed, and an artificial neural network was trained using the prepared database of images of typical objects. The resulting algorithm was tested by the authors on a validation data set, the quality of its work was estimated using various metrics, and the algorithm’s performance was measured. Conclusions are made about the necessity to expand the collected database of images of typical marine objects, further steps are proposed to improve the accuracy of the developed software and algorithmic complex and its implementation to be used in a marine optoelectronic machine vision system for automa-tion and improving the safety of civil navigation.
-
DISTRIBUTED SYSTEM FOR BARCODE RECOGNITION USING NEURAL NETWORKS
А.Y. Yurchenko , М.Y. Polenov70-792025-10-01Abstract ▼This work presents a distributed software-hardware system for automated barcode recognition on moving objects in industrial environments. The primary objective of the research is to develop a reliable and adaptive solution capable of consistently reading barcodes regardless of the orientation, speed, or height of objects moving along a conveyor belt. The main focus is not on achieving maximum processing speed, but rather on providing a wide field of view and ensuring reliable recognition of moving objects. Unlike traditional scanners that require precise positioning and expensive hardware, the proposed approach leverages a single network camera and a server equipped with neural processing modules, providing a cost-effective and versatile alternative suitable for a wide range of industrial applications. A key component of the system architecture is a neural image restoration module based on the MPRNet model, which effectively reduces motion blur and optical distortions in video frames. After preprocessing, frames are passed to an object detection module built upon the YOLO architecture, which has been adapted specifically for barcode recognition. Detected barcode data is stored in a database using an ORM interface, enabling seamless integration with existing enterprise systems. To prevent frame loss and maintain high throughput, the system incorporates asynchronous processing mechanisms using multithreading and buffered queues. The relevance of this research stems from the widespread use of barcodes as the primary method of product marking in industrial settings and the increasing demand for automation in product tracking and inventory control. Despite the availability of various vision-based and scanning solutions, most existing systems are not designed to handle unstable or low-quality video streams. The proposed system demonstrates robustness to visual distortions and motion-related artifacts, making it suitable for deployment in real production environments. Its affordability and adaptability also open up possibilities for implementation in logistics, warehousing, and supply chain management.
-
NOISE GENERATION METHOD BASED ON A SET OF NOISY IMAGES WITHOUT CLEAN EXAMPLES
А.S. Kovalenko , Y. М. Demyanenko243-2542025-11-10Abstract ▼In this work, a novel method is proposed for noise generation from noisy images that does not require aligned pairs of clean and noisy data. Unlike traditional approaches demanding matched image sets or a priori noise models, the developed technique models complex noise characteristics intrinsic to specific CMOS sensors solely from observed noisy data. Noise synthesis is achieved via a U‑Net‑like generative adversarial architecture based on StyleGANv2, featuring a modified discriminator conditioned on camera parameters and input image metadata. Special emphasis is placed on preserving the spatial–color structure and textural details of each image, enforced through a dedicated loss function that ensures fidelity to the original color rendering and fine-grained patterns. Training of the noise generator is performed without any paired clean and noisy images, which proves particularly valuable when handling real-world datasets acquired from multiple camera models under varied lighting conditions. The experimental section presents a detailed comparative analysis of the synthesized images using PSNR and SSIM metrics, along with an evaluation of the noise distribution based on intensity statistics and spectral characteristics. It is demonstrated that the generated dataset functions effectively as a standalone training corpus for denoising neural networks and, when combined with a real dataset (e.g., SIDD), yields further enhancements in denoising performance. Results indicate that combined training on the union of generated and real examples produces an average PSNR improvement of 1.5 dB compared to existing methods reliant on aligned data. Independence from the specific optical characteristics of any given sensor significantly broadens the method’s applicability. These findings confirm the utility of the proposed approach for realistic noise synthesis and removal in scenarios lacking clean reference images, and they open avenues for future research into adaptive noise-model generation
-
MODERN APPROACHES TO NATURAL FIRE MONITORING AND FORECASTING: REVIEW AND CONCEPT OF AUTONOMOUS UAV-BASED SYSTEM
N.D. Boldyrev , V. V. Gilka , А.S. Kuznetsova , D.А. Morozov58-802025-12-30Abstract ▼Natural fires cause serious damage to ecosystems, the economy, and public safety every year, and timely detection of fires and prediction of their development increases the speed of response to threats and allows for optimal allocation of resources during emergency response. Existing monitoring methods are limited by the speed of detecting fire outbreaks and the speed of their further spread, which reduces the effectiveness of rescue services. To solve this problem, heterogeneous data sources can be used, including unmanned aerial vehicles (UAVs), distributed sensor networks, mobile field observation systems, ground-based thermal imaging stations, etc., which can contribute to a more accurate analysis of the current situation and improve the reliability of predictive models of fire spread. The aim of the study was to develop a concept for an automated approach to monitoring and predicting wildfires based on unmanned aerial vehicles. We believe that this approach will improve the speed of detecting fire outbreaks and the accuracy of predicting their spread. The tasks include analyzing existing monitoring methods, developing a concept for a system that integrates multispectral imaging, optimized data transmission, automatic segmentation, and forecasting based on machine learning, as well as ensuring interaction between the operator and alert specialists. The work used methods of collecting, analyzing, and transmitting data from UAVs, processing multispectral images, machine learning and neural networks for fire detection, image segmentation algorithms and simulation modeling for fire spread prediction, data visualization to support decision-making by operators and administrators, logging and analysis of results for model training, software engineering, and human-computer interaction technologies. The system will reduce the time required to detect and predict fires, enable operators to launch multiple drones simultaneously, and automate the processing of data received from them. Process automation will reduce emergency response times and staffing levels, improve resource allocation, increase forecast accuracy, and improve the timeliness of emergency service notifications. This will help reduce damage from wildfires and improve the safety of people and ecosystems. Despite the progress made in addressing this challenge, the comprehensive system described in this article does not yet exist in its entirety in Russia, the CIS countries, or in Western and Asian countries. Although individual components, such as UAVs for monitoring and artificial intelligence (AI) for data analysis, are already in active use, there is currently no integrated solution that combines all elements (drone control, near real-time fire spread prediction, data transmission, and interaction with emergency services). does not currently exist. This concept represents a new approach that could become a breakthrough technology for combating natural disasters.
-
MODERN APPROACHES TO FACE RECOGNITION IN LOW-LIGHT CONDITIONS: A REVIEW AND THE CONCEPT OF A HYBRID END-TO-END ARCHITECTURE
D. А. Morozov , V.V. Gilka , А. S. Kuznetsova113-1332026-07-07Abstract ▼The article addresses the problem of reliable face recognition in critical areas such as video surveillance and biometric authentication under low-light conditions. Existing approaches typically separate the tasks of image enhancement and face identification, which leads to error accumulation and loss of informative features. The aim of this work is to overcome this limitation by developing and theoretically substantiating a hybrid end-to-end architecture in which image enhancement and face recognition are solved jointly. The study provides a systematic review of modern methods, including classical algorithms (such as histogram equalization and noise suppression) and advanced deep neural networks (including EnlightenGAN, Zero-DCE, ArcFace, and RetinaFace). The main contribution is the integration of generative and identification modules into a single computational graph. The key result of the study is the demonstration that joint optimization of all processing stages within a unified model, unlike fragmented solutions, fundamentally changes the approach to the problem. Theoretical analysis and comparative evaluation of existing concepts show that the proposed architecture ensures a more efficient gradient flow during training, leading to the formation of higher-quality and noise-robust identity features. It is shown that this approach prevents error accumulation between stages and minimizes information loss. The novelty of the work lies in the holistic, end-to-end view of the face recognition problem under low-light conditions. The practical significance is confirmed by the applicability of the architecture in real systems, where its implementation can potentially improve reliability and processing speed by combining heterogeneous tasks into a single optimizable framework.
The article addresses the problem of reliable face recognition in critical areas such as video surveillance and biometric authentication under low-light conditions. Existing approaches typically separate the tasks of image enhancement and face identification, which leads to error accumulation and loss of informative features. The aim of this work is to overcome this limitation by developing and theoretically substantiating a hybrid end-to-end architecture in which image enhancement and face recognition are solved jointly. The study provides a systematic review of modern methods, including classical algorithms (such as histogram equalization and noise suppression) and advanced deep neural networks (including EnlightenGAN, Zero-DCE, ArcFace, and RetinaFace). The main contribution is the integration of generative and identification modules into a single computational graph. The key result of the study is the demonstration that joint optimization of all processing stages within a unified model, unlike fragmented solutions, fundamentally changes the approach to the problem. Theoretical analysis and comparative evaluation of existing concepts show that the proposed architecture ensures a more efficient gradient flow during training, leading to the formation of higher-quality and noise-robust identity features. It is shown that this approach prevents error accumulation between stages and minimizes information loss. The novelty of the work lies in the holistic, end-to-end view of the face recognition problem under low-light conditions. The practical significance is confirmed by the applicability of the architecture in real systems, where its implementation can potentially improve reliability and processing speed by combining heterogeneous tasks into a single optimizable framework.
-
IMPROVING SEGMENTATION IN MULTIPHASE CT IMAGES USING TRAINABLE PHASE SUPERIMAGING
S. V. Ermolenko , I. L. Kashirina77-902026-09-10Abstract ▼Joint Automated analysis of multiphase CT scans, acquired at different time points after contrast agent administration, is highly important for accurate pathology detection. However, it faces a fundamental problem of spatial misalignment between phases due to patient breathing and movement, which leads to a significant reduction in the accuracy of automatic segmentation. Existing pre-registration (phase alignment) methods require manual parameter tuning and are not integrated into trainable pipelines, hindering the full automation of the segmentation process. The aim of this study was to develop a differentiable method for aligning multiphase CT images based on trainable linear affine transformations, fully embedded into the segmentation model training pipeline. Unlike traditional approaches, the proposed method implements a differentiable affine registration module (including translation, rotation, and scaling), whose parameters are optimized via gradient descent without manual tuning and are directly integrated into the computational graph of the nnU-Net model. The study compared the proposed method with a baseline multiphase segmentation approach (based on simple phase concatenation without registration) and an affine registration method implemented using the ITK.Elastix library on the open abdominal CT dataset WAW-TACE. Compared to segmentation without registration, a substantial improvement in quality metrics was achieved: an increase in the Dice coefficient by 43.19% and in the ROC-AUC metric by 13.55%. Compared to Elastix, the improvements were 13.43% in Dice, 6.86% in ROC-AUC, and 19.6% in the accuracy of pathology count detection per CT scan. The practical significance of the research lies in the development of a ready-to-use PyTorch module for integration into CT image analysis pipelines. It enables fully automated registration without the need for manual hyperparameter tuning and offers high computational efficiency
-
ON THE INFLUENCE OF NOISE ON THE RECOGNITION OF THREEFOLD ROTATIONAL SYMMETRY IN HEXAGONAL IMAGES
A.N. Karkishchenko, V.B. Mnukhin2021-01-19Abstract ▼The article presents an algebraic approach to the representation and processing of digital
images defined on hexagonal lattices. The described approach is based on the representation of images
as functions on finite fields of “Eisenstein's integers”. As it turns out, the elements of such fields
naturally correspond to the pixels of hexagonal images of certain sizes. The exponential and logarithmic
transformations in the Eisenstein fields are described. A method for detecting the centers of
threefold rotational symmetry in grayscale images is presented and the corresponding normalized
measure of symmetry is introduced. The main purpose of the work is to study the effect of noise on the
image on the quality of the symmetry assessment using the introduced measure. The noise factor must
be taken into account, since a decrease in the measure can be caused not only by the incomplete
symmetry of the real object, but also by distortions due to noise, which is almost always the case.
Obviously, this difference will be proportional to the level of the noise component. Analytical estimates
of the effect of noise on the criterion for detecting symmetry are obtained in this work. If images
are subject to random noise, then the measure of symmetry of local image areas will be a random
variable, the distribution law of which is determined by the distribution laws of noise components. At
the same time, the standard for image processing assumption is made in the work about the model of
normal and independent noise level of the brightness function. The peculiarity of the introduced
threefold rotational symmetry measure does not allow directly applying standard methods to obtain
probabilistic estimates. For this purpose, an assessment of the cumulative probability distribution
function was carried out, on the basis of which an expression was obtained for the probabilities of
deviation of the symmetry measure from the true value by a given value. By virtue of the a priori
assumptions made, the obtained estimate should be considered as rather "cautious" and it can be
expected that in reality the spread of the measure caused by noise in the image will be significantly
less than the theoretically established boundaries.








