Search
Search Results
-
GAME-THEORETIC AND REINFORCEMENT LEARNING-BASED ALGORITHM FOR INTENTIONAL JAMMING MITIGATION
К. S. Grigoryan , Е. S. Basan66-772026-09-10Abstract ▼Intentional jamming represents a serious threat to the security and availability of wireless communication systems. Modern wireless networks, including cognitive radio, sensor networks, and Internet of Things infrastructures, are particularly vulnerable because adversaries can dynamically adapt their jamming strategies. The objective of this study is to develop an adaptive anti-jamming algorithm capable of maintaining communication reliability under dynamic interference conditions. To achieve this goal, the interaction between the legitimate transmitter and the jammer is modeled as a Markov Stackelberg game, where the legitimate node acts as a leader and the jammer acts as a follower. Reinforcement learning is used to determine the optimal strategy of the leader in a stochastic environment, while robustness against channel uncertainty is ensured through a SOCP (Second-order cone programming) formulation that guarantees the required quality-of-service constraints. The learning process is implemented using the SAC (Soft Actor-Critic) algorithm, which enables stable policy optimization in continuous action spaces and stochastic environments.
The research tasks include the formalization of the anti-jamming interaction as a Markov decision process, the integration of reinforcement learning with a robust SOCP optimization layer, and the evaluation of the proposed approach through simulation. A Monte Carlo simulation of the proposed algorithm, as well as several algorithms based on FHSS (Frequency-Hopping Spread Spectrum), was conducted. The proposed anti-jamming algorithm reduces the probability of communication outage by 0.14. The results indicate that the proposed approach improves the resilience of wireless communication systems and reduces the probability of successful denial-of-service attacks at the physical layer -
REINFORCEMENT LEARNING METHODS IN ADAPTIVE CONTROL OF NONLINEAR DYNAMIC OBJECTS
М.Y. Medvedev , V.K., А.R. Gaiduk , I.М. Medvedev , Е.Y. Kosenko2026-04-29Abstract ▼The relevance of the problem of adaptive control of nonlinear dynamic objects is ensured by the ever-increasing demands on the quality and operating conditions of technical systems. Automation and robotics lead to an increase in the complexity of the problems being solved and the need to adapt to structural and parametric uncertainty and external disturbances. In recent years, the use of reinforcement learning methods for the synthesis of adaptive control systems has gained popularity. These methods demonstrate effectiveness in controlling uncertain dynamic objects. However, there are two fundamental problems with the application of machine learning methods to control systems for dynamic objects. First, the need to ensure asymptotic stability of the desired trajectory of a closed-loop system limits the application of deep learning methods. Second, during the learning process, it is necessary that the intelligent controller does not generate controls that lead to state variables exceeding specified limits. The purpose of this article is to review and analyze recent advances in the application of reinforcement learning methods to the synthesis of adaptive control systems for nonlinear dynamic objects. Particular attention is given to Actor-Critic methods, which are structurally similar to adaptive control systems with self-adjusting parameters. Based on the analysis, the structure of an adaptive system with two-component control, including nominal and adaptive controllers, is proposed. An adaptive control algorithm is proposed, distinguished by the use of a modified Actor-Critic algorithm, distinguished by a new form of the delta error and the absence of a singularity in the neighborhood of zero. The proposed algorithm allows for a reduction in the number of adjustable parameters. The article presents conditions for Lyapunov stability of the zero-equilibrium position of a closed-loop system and an example of the synthesis and modeling of the proposed adaptive control algorithm
-
RESEARCH OF AN INTELLIGENT ADAPTIVE CONTROL ALGORITHM BASED ON THE REINFORCEMENT LEARNING METHOD
А. N. Karapeev, Е.Y. Kosenko, М. Y. Medvedev, V. K. Pshikhopov2025-04-27Abstract ▼An algorithm for adaptive control of a DC motor based on the use of machine learning technology
with reinforcement is proposed and investigated. An overview and brief analysis of the state of affairs in
the field of intelligent motor control systems is given. A mathematical model of the DC motor is presented,
and a structural scheme for training an intellectual agent is presented. An intelligent adaptive motor speed
control system is proposed. The DC motor is represented as a black box with the limited input and output.
The control system is based on a zero-order Q-learning algorithm. It is assumed that the output of the
intelligent agent is a control applied to the motor input. The intelligent system uses a tabular approximation
of the value of each of the control action. In this article, we study the effect of the discreteness of the
representation of state, the set of control effects used, the applied rewards, and the parameters of the
learning algorithm on the control error. The sensitivity of the control system to the parameters of the motor
and an unmeasured moment is investigated. Based on the results of the study, a modified algorithm is
proposed, which assumes the measurement or evaluation of the current of the motor stator. The control
algorithm provides robustness to parameters and external disturbance. Additionally, the approximation of
the control value function using polynomials and using a neural network are investigated -
MACHINE LEARNING MODEL OF SWARM EVASION FROM THE INFLUENCE OF ANTAGONISTIC ENVIRONMENT
V. К. Abrosimov, G.А. Dolgov, Е. S. Mikhailova6-192025-04-27Abstract ▼One of the priority areas of group control theory for the near future is swarm control of groups of
small unmanned aerial vehicles - micro-, mini- and nano-classes, performing a collective task under enemy
influence. Here, two antagonistic strategies collide - minimization of losses from the point of view of
the attacking swarm and maximization of such losses from the point of view of the defense system. Research objective: development of an approach to solving a practical problem - penetration of a swarm of
unmanned aerial vehicles into an object protected by a defense system. The objectives of the study were to
analyze the characteristics of the factors influencing the processes of detection, tracking, recognition of
swarm intentions by the defense system and the development of a machine learning model for creating
spatio-temporal formations that minimize the number of swarm elements affected by the defense system.
The main parameters of the defense system are the detection range and duration of swarm recognition, the
time to make a decision on the actions of the swarm, the size of the zone of destruction of defense means.
The method of machine learning on convolutional neural networks with reinforcement was chosen as the
research method. The counteraction effect against the defense system is created due to the swarm's dynamics;
it can actively maneuver, creating spatio-temporal maneuvers during the mission. To simulate the
"Swarm vs. Defense System" situation, a swarm agent (a neural network with a transformer architecture
that initiates swarm formations) and a defense system agent are introduced that recognizes the swarm and
attacks it, creating a zone of destruction in the conventional center of mass of the swarm. The swarm is
guided by a stochastic rule, asking the defense system (environment) to react to its maneuver. The environment
responds by attacking the swarm, creating a damaging factor at the point where the swarm or the
main part of the swarm is expected to be. The reward of the swarm strategy is the number of undestroyed
objects under the conditions of constraints; for the defense system, this "reward" acts as a "punishment".
An interesting phenomenon was established in the process of machine learning: each swarm element,
remaining within a given space and implementing the biological principles of swarm control without a
Leader, independently evades the area of destruction, which together creates a random spatio-temporal
formation for defense means with minimal losses of swarm elements. Thus, using the method of machine
learning with reinforcement, a model was created that allows varying the behavior of the swarm and synthesizing
spatio-temporal formations that complicate detection, tracking, recognition of intentions and
decision-making on the impact of the defense system on a swarm of attacking small unmanned aerial vehicles,
as well as significantly reducing their losses.








