[关键词]
[摘要]
【目的】在无刷双馈电机(BDFM)的空间矢量调制直接转矩控制(SVM-DTC)系统中,内环通常采用转矩与磁链比例积分(PI)控制器。然而固定的PI参数难以适应系统非线性特性以及频繁变化的运行工况,难以在全工作点维持最优控制性能。为此,本文提出一种基于深度强化学习(DRL)的SVM-DTC策略,摆脱对人工参数整定的依赖,提升系统动态性能与鲁棒性。【方法】采用双延迟深度确定性策略梯度(TD3)算法,构建并训练Actor-Critic神经网络代替内环转矩和磁链PI调节器。状态空间由转矩、磁链、转矩偏差、磁链偏差、转矩偏差积分、磁链偏差积分、转速参考值以及转速反馈值构成;动作空间为控制绕组参考电压矢量的d、q轴分量。设计一种综合转矩脉动抑制和磁链偏差惩罚的奖励函数,引导TD3智能体在连续动作空间中学习最优电压决策。【结果】在转速阶跃、负载突变等典型运行工况下开展仿真,对本文所提基于DRL的SVM-DTC策略与传统SVM-DTC策略进行性能分析。仿真结果表明,所提策略与传统SVM-DTC策略的转速控制效果基本一致。【结论】经充分训练的Actor网络能够依据系统实时状态直接生成控制绕组参考电压,省去内环转矩与磁链PI调节器的参数整定环节。基于TD3的SVM-DTC策略能够实现对BDFM转矩与磁链的高品质控制,具有优异的动态响应能力与鲁棒性。
[Key word]
[Abstract]
[Objective] In the space vector modulation direct torque control (SVM-DTC) of brushless doubly-fed machine (BDFM), inner-loop torque and flux proportional-integral (PI) controllers are commonly used. However, fixed PI parameters are difficult to adapt to the system’s nonlinearity and frequently varying operating conditions, making it challenging to maintain optimal performance across all operating points. To address this problem, a deep reinforcement learning (DRL)-based SVM-DTC strategy is proposed to eliminate the dependence on manual parameter tuning and improve the system’s dynamic performance and robustness. [Methods] The twin delayed deep deterministic policy gradient (TD3) algorithm was adopted to construct and train an Actor-Critic neural network, which replaced the inner-loop torque and flux PI controllers. The state space consisted of estimated torque, estimated flux, torque error, flux error, integral of torque error, integral of flux error, speed reference, and speed feedback. The action space was defined as the d-q axis components of the control winding reference voltage vector. Meanwhile, a reward function was designed to integrate torque ripple suppression, and flux deviation penalty, guiding the TD3-agent to learn optimal voltage decisions in a continuous action space. [Results] Simulations were carried out under typical operating conditions such as speed step change and sudden load variation to compare the performance of the proposed DRL-based SVM-DTC strategy and the traditional SVM-DTC strategy. The simulation results showed that the proposed method achieved basically consistent speed control performance with the traditional SVM-DTC strategy. [Conclusion] The well-trained Actor network can directly generate the control winding reference voltages according to the system’s real-time states, obviating the parameter tuning process of the inner-loop torque and flux PI controllers. The TD3-based SVM-DTC strategy can achieve high-quality control of torque and flux for BDFM, exhibiting excellent dynamic response and robustness.
[中图分类号]
[基金项目]
国家自然科学基金项目(62173151)