Designing a stabilizing controller for nonlinear systems is a challenging task, especially for high-dimensional problems with unknown dynamics. Traditional reinforcement learning algorithms applied to stabilization tasks tend to drive the system close to the equilibrium point. However, these approaches often fall short of achieving true stabilization and result in persistent oscillations around the equilibrium point. In this work, we propose a reinforcement learning algorithm that stabilizes the system by learning a local linear representation ofthe dynamics. The main component of the algorithm is integrating the learned gain matrix directly into the neural policy. We demonstrate the effectiveness of our algorithm on several challenging high-dimensional dynamical systems. In these simulations, our algorithm outperforms popular reinforcement learning algorithms, such as soft actor-critic (SAC) and proximal policy optimization (PPO), and successfully stabilizes the system. To support the numerical results, we provide a theoretical analysis of the feasibility of the learned algorithm for both deterministic and stochastic reinforcement learning settings, along with a convergence analysis of the proposed learning algorithm. Furthermore, we verify that the learned control policies indeed provide asymptotic stability for the nonlinear systems.

本研究解决了高维未知动态非线性系统控制的稳定性问题，传统强化学习算法在此任务中的表现不足。我们提出了一种新的强化学习算法，通过学习系统动力学的局部线性表示来实现稳定控制，并将学习得到的增益矩阵直接整合进神经策略中。实验结果表明，该算法在多种高维动态系统中表现优于主流强化学习算法，成功实现了系统的稳定性。

具有稳定性保证的随机强化学习在未知非线性系统控制中的应用