REINFORCEMENT LEARNING METHODS IN ADAPTIVE CONTROL OF NONLINEAR DYNAMIC OBJECTS
Abstract
The relevance of the problem of adaptive control of nonlinear dynamic objects is ensured by the ever-increasing demands on the quality and operating conditions of technical systems. Automation and robotics lead to an increase in the complexity of the problems being solved and the need to adapt to structural and parametric uncertainty and external disturbances. In recent years, the use of reinforcement learning methods for the synthesis of adaptive control systems has gained popularity. These methods demonstrate effectiveness in controlling uncertain dynamic objects. However, there are two fundamental problems with the application of machine learning methods to control systems for dynamic objects. First, the need to ensure asymptotic stability of the desired trajectory of a closed-loop system limits the application of deep learning methods. Second, during the learning process, it is necessary that the intelligent controller does not generate controls that lead to state variables exceeding specified limits. The purpose of this article is to review and analyze recent advances in the application of reinforcement learning methods to the synthesis of adaptive control systems for nonlinear dynamic objects. Particular attention is given to Actor-Critic methods, which are structurally similar to adaptive control systems with self-adjusting parameters. Based on the analysis, the structure of an adaptive system with two-component control, including nominal and adaptive controllers, is proposed. An adaptive control algorithm is proposed, distinguished by the use of a modified Actor-Critic algorithm, distinguished by a new form of the delta error and the absence of a singularity in the neighborhood of zero. The proposed algorithm allows for a reduction in the number of adjustable parameters. The article presents conditions for Lyapunov stability of the zero-equilibrium position of a closed-loop system and an example of the synthesis and modeling of the proposed adaptive control algorithm
References
1. Zemlyakov S., Rutkovsky V. Synthesizing self-adjusting control systems with a standard model, Automa-tion and Remote Control, 1966, Vol. 27(3), pp. 407-414.
2. Parks P. Liapunov Redesign of Model Reference Adaptive Control Systems, IEEE Transactions on Automatic Control, 1966, Vol. 11(3), pp. 362-367.
3. Stafeychuk B.G., Shakirova A.Ya. Issledovanie adaptivnoy sistemy avtomaticheskogo regulirovaniya s primeneniem neyrosetevykh tekhnologiy na imitatsionnoy modeli reaktora [Study of an adaptive auto-matic control system using neural network technologies on a simulation model of a reactor], Vestnik PNIPU. Khimicheskaya tekhnologiya i biotekhnologiya [Bulletin of PNIPU. Chemical technology and biotechnology], 2019, No. 2.
4. Eremenko Yu.I., Glushchenko A.I., Fomin A.V. O primenenii neyrosetevogo nastroyshchika parametrov PI-regulyatora na teplovykh ob"ektakh gorno-metallurgicheskoy otrasli v rezhime otrabotki vozmush-cheniy [On the use of a neural network tuner for PI controller parameters at thermal facilities in the min-ing and metallurgical industry in the disturbance processing mode], GIAB [GIAB], 2017, No. 12.
5. Santalov A.A. Neyrosetevaya nastroyka adaptivnogo PID-regulyatora moshchnosti gidroagregata [Neu-ral network tuning of an adaptive PID controller for the power of a hydroelectric unit], Vestnik UlGTU [Bulletin of UlSTU], 2021, No. 3 (95).
6. Satton R.S., Barto E.Dzh. Obuchenie s podkrepleniem: Vvedenie [Reinforcement Learning: An Intro-duction. 2nd ed.: transl. from Engl. A.A. Slinkina. Moscow: DMK Press, 2020, 552 p.
7. Gaiduk A.R., Medvedev M.Y. and Pshikhopov V.K. Design Method of Quasilinear Nonaffine Nonlinear Control Systems of General Structure, IEEE Transactions on Automation Science and Engineering, 2025, Vol. 22, pp. 20208-20220.
8. Gayduk A.R., Medvedev M.Yu., Pshikhopov V.Kh., Gistsov V.G. Sintez neaffinnykh nelineynykh sistem upravleniya na osnove kvazilineynykh modeley [Design of nonaffi ne nonlinear control systems based on quasilinear models], Mekhatronika, avtomatizatsiya, upravlenie [Mekhatronika, Avtomatizatsiya, Up-ravlenie], 2025, 26(5), pp. 223-232. Available at: https://doi.org/10.17587/mau.26.223-232.
9. Karapeev A.N., Kosenko E.Yu., Medvedev M.Yu., Pshikhopov V.Kh. Issledovanie intellektual'nogo adap-tivnogo algoritma upravleniya na baze metoda obucheniya s podkrepleniem [Study of an intelligent adap-tive control algorithm based on the reinforcement learning method], Izvestiya YuFU. Tekhnicheskie nauki [Izvestiya SFedU. Engineering Sciences], 2025, No. 2, pp. 162-175.
10. Medvedev M.Yu., Pshikhopov V.Kh., & Evdokimov I.D. Algoritm robastnogo upravleniya odnomernym dinamicheskim ob"ektom na osnove tablichnogo Q-metoda obucheniya s podkrepleniem [Robust control algorithm for single input single output dynamic object based on table-based
Q-method of reinforcement learning], Informatika i avtomatizatsiya [Informatics and Automation], 2025, 24 (3), pp. 717-744. Available at: https://doi.org/10.15622/ia.24.3.1.
11. Xu B., Yang C. and Shi Z. Reinforcement Learning Output Feedback NN Control Using Deterministic Learning Technique, IEEE Transactions on Neural Networks and Learning Systems, 2014, Vol. 25 (3), pp. 635-641. DOI: 10.1109/TNNLS.2013.2292704.
12. Mu C., Ni Z., Sun C., and He H. Data-driven tracking control with adaptive dynamic programming for a class of continuous-time nonlinear systems, IEEE Transactions on Cybernetics, 2016, Vol. 47 (6), pp. 1460-1470. DOI: 10.1109/TCYB.2016.2548941.
13. Wang A., Liao X., Dong T. Event-driven optimal control for uncertain nonlinear systems with external disturbance via adaptive dynamic programming, Neurocomputing, 2018, Vol. 281, pp. 188-195. DOI: 10.1016/j.neucom.2017.12.010.
14. Biao Luo, Yin Yang, Derong Liu, and Huai-Ning Wu. Event-Triggered Optimal Control with Perfor-mance Guarantees Using Adaptive Dynamic Programming, IEEE Transactions on Neural Networks and Learning Systems, 2020, Vol. 31 (1), pp. 76-88. DOI: 10.1109/TNNLS.2019.2899594.
15. Yang X., Xu M. and Wei Q. Dynamic Event-Sampled Control of Interconnected Nonlinear Systems Us-ing Reinforcement Learning, IEEE Transactions on Neural Networks and Learning Systems, 2024, Vol. 35(1), pp. 923-937. DOI: 10.1109/TNNLS.2022.3178017.
16. Zhang H., Zhao X., Wang H., Zong G. and Xu N. Hierarchical Sliding-Mode Surface-Based Adaptive Actor–Critic Optimal Control for Switched Nonlinear Systems With Unknown Perturbation, IEEE Transactions on Neural Networks and Learning Systems, 2024, Vol. 35 (2), pp. 1559-1571. DOI: 10.1109/TNNLS.2022.3183991.
17. Dong C., Chen L. and Dai S.-L. Performance-Guaranteed Adaptive Optimized Control of Intelligent Surface Vehicle Using Reinforcement Learning, IEEE Transactions on Intelligent Vehicles, 2024,
Vol. 9(2), pp. 3581-3592. DOI: 10.1109/TIV.2023.3338486.
18. Phuong Nam Dao, Minh Hiep Phung. Nonlinear robust integral based actor–critic reinforcement learn-ing control for a perturbed three-wheeled mobile robot with mecanum wheels, Computers and Electrical Engineering, 2025, Vol. 121. 109870. DOI:10.1016/j.compeleceng.2024.109870.
19. Ding Wang, Xin P., Ren J., and Qiao J. Adaptive Critic Tracking Design for Data-Based Nonaffine Predictive Control, IEEE Transactions on Automation Science and Engineering, 2024, Vol. 21(4),
pp. 5534-5545. DOI: 10.1109/TASE.2023.3313159.
20. Berkenkamp F., Turchetta M., Schoellig A., Krause A. Safe model-based reinforcement learning with stability guarantees, Advances in Neural Information Processing Systems, 2017, Vol. 30, pp. 908-918. DOI: 10.48550/arXiv.1705.08551.
21. Han M., Zhang L., Wang J., Pan W. Actor-critic reinforcement learning for control with stability guaran-tee, IEEE Robotics and Automation Letters, 2020, Vol. 5 (4), pp. 6217-6224. DOI: 10.1109/LRA.2020.3011351.
22. Zhou Z., Liu A., and Wang D. Improved adaptive-critic-based dynamic event-triggered control for non-affine systems, Mechatronics Tech., 2024, Vol. 1 (1). DOI: 10.55092/mt20230002.
23. Cheng R., Orosz G., Murray R.M., Burdick J.W. End-to end safe reinforcement learning through barrier functions for safety critical continuous control tasks, The Thirty-Third AAAI Conference on Artificial In-telligence, 2019, pp. 3387-3395.
24. Choi J., Castaneda F., Tomlin C.J., Sreenath K. Reinforcement learning for safety-critical control under model uncertainty, using control Lyapunov functions and control barrier functions, Conference Robotics: Science and Systems, 2020.
25. Xue W., Lian B., Fan J., Kolaric P., Cha T., Lewis F. Inverse Reinforcement Q-Learning Through Ex-pert Imitation for Discrete-Time Systems, IEEE Transactions on Neural Networks and Learning Sys-tems, 2021, pp. 1-14. 10.1109/TNNLS.2021.3106635.
26. Yu Zhang, Niu B., Zhao X., Duan P., Wang H., and Gao B. Global Predefined-Time Adaptive Neural Network Control for Disturbed Pure-Feedback Nonlinear Systems with Zero Tracking Error, IEEE Transactions on Neural Networks and Learning Systems, 2023, Vol. 34 (9), pp. 6328-6338.
27. Wang D., Zhou Z., Li M., Ren J., and Qia J. Event-based robust performance guarantee for nonaffine plants via system identification, International Journal Robust Nonlinear Control, 2023, Vol. 33, pp. 53650-5387. DOI: 10.1002/rnc.6646.
28. Zanon M. Gros S. Safe Reinforcement Learning Using Robust MPC, IEEE Transactions on Automatic Control, 2020, pp. 1-1. 10.1109/TAC.2020.3024161.
29. Rizvi S. A.A., Lin Z. Output Feedback Q-Learning Control for the Discrete-Time Linear Quadratic Regu-lator Problem, IEEE Transactions on Neural Networks and Learning Systems, 2018, pp. 1-14. 10.1109/TNNLS.2018.2870075.
30. Kim J.W., Oh T.H., Son S.H., Jeong D.H., Lee J.M. Convergence analysis of the deep neural networks based globalized dual heuristic programming, Automatica, 2020, Vol. 122. DOI: 10.1016/j.automatica.2020.109222.
31. Borovik V.S., Shidlovskiy S.V. Obuchenie s podkrepleniem v sistemakh upravleniya ob"ektami s transportnym zapazdyvaniem [Reinforcement learning in plant control systems with transport lag], Avtometriya [Autometry], 2021, Vol. 57 (3), pp. 48-57. DOI: 10.15372/AUT20210306.
32. Galyaev A.A., Medvedev A.I., Nasonov I.A. Neyrosetevoy algoritm perekhvata mashinoy Dubinsa tseley, dvizhushchikhsya po izvestnym traektoriyam [Neural network algorithm for intercepting targets moving along known trajectories by a Dubins’ car], Avtomatika i telemekhanika [Automation and Telemechan-ics], 2023, No. 3, pp. 3-21. DOI: 10.31857/S0005231023030017.
33. Khapkin D.L., Feofilov S.V. Sintez ustoychivykh neyrosetevykh regulyatorov dlya ob"ektov s ogra-nichitelyami v usloviyakh nepolnoy informatsii [Synthesis of stable neural network controllers for ob-jects with constraints under conditions of incomplete information], Mekhatronika, avtomatizatsiya, up-ravlenie [Mekhatronika, Avtomatizatsiya, Upravlenie], 2024, Vol. 25 (7), pp. 345-353. Available at: https://doi.org/10.17587/mau.25.345-353.
34. Pshikhopov V., Krukhmalev V., Medvedev M., and Neydorf R. Estimation of Energy Potential for Control of Feeder of Novel Cruiser/Feeder MAAT System, SAE Technical Paper, 2012, Vol. 2012-01-2099. doi: 10.4271/2012-01-2099.
35. Pshikhopov V.Kh., Medvedev M.Yu., Gayduk A.R., Neydorf R.A., Belyaev V.E., Fedorenko R.V., Kost-yukov V.A., Krukhmalev V.A. Sistema pozitsionno-traektornogo upravleniya robotizirovannoy vozdu-khoplavatel'noy platformoy: algoritmy upravleniya [Position-trajectory control system of a robotic aero-nautical platform: control algorithms], Mekhatronika, avtomatizatsiya i upravlenie [Mekhatronika, Avtomatizatsiya, Upravlenie], 2013, No. 7, pp. 13-20.
36. Gayduk A.R., Pshikhopov V.Kh., Medvedev M.Yu., Gistsov V.G. Nepreryvnoe upravlenie nelineynymi neaffinnymi ob"ektami [Continuous control of nonlinear non-affine objects], Izvestiya YuFU. Tekhnich-eskie nauki [Izvestiya SFedU. Engineering Sciences], 2024, No. 1 (237), pp. 122-133.
37. Medvedev M.Yu., Pshikhopov V.Kh., Medvedev I.M. Algoritmy adaptivnogo upravleniya kaskadnymi dinamicheskimi ob"ektami na baze metoda obucheniya s podkrepleniem [Algorithms for adaptive control of cascade dynamic objects based on the reinforcement learning method], Mekhatronika, avtomatizatsiya i upravlenie [Mekhatronika, Avtomatizatsiya, Upravlenie], 2026, No. 3








