[关键词]
[摘要]
【目的】针对微电网群协同调度中可再生能源出力和负荷需求不确定性强、运行约束复杂以及决策变量连续等问题,提出一种基于改进双延迟深度确定性策略梯度(TD3)算法的微电网群经济优化方法。【方法】首先,构建包含风电、光伏、微型燃气轮机、储能系统、电动汽车充电负荷及主网交互的三微电网群模型,并将其经济优化调度过程建模为马尔可夫决策过程。其中,状态空间包括负荷需求、可再生能源出力、分时电价和储能荷电状态等信息,动作空间包括储能充放电功率、微型燃气轮机出力、微网间交互功率及主网交互功率。其次,综合考虑归一化运行成本、环境成本和约束惩罚项构造奖励函数,引导智能体学习经济且可行的调度策略。最后,在原始TD3算法基础上引入双经验池(DEP)机制和自适应指数衰减高斯探索噪声(AED-GEN)机制。DEP机制将经验样本划分为可行解样本和边界探索样本,以提高样本利用率并增强约束边界学习能力;AED-GEN机制用于平衡训练前期探索能力和后期收敛稳定性。【结果】以三微电网群为对象进行算例分析,结果表明,改进TD3算法在收敛稳定性、运行经济性和调度安全性方面均优于深度确定性策略梯度(DDPG)、原始TD3和SAC算法。改进TD3算法最后100个训练回合的奖励波动标准差为1.51,较SAC、原始TD3和DDPG分别降低35.74%、45.09%和54.24%;其日运行总成本为4 186.94元,较SAC、原始TD3和DDPG分别降低1.38%、6.31%和8.71%。同时,动作越限率在训练过程中快速下降,并在训练后期逐渐接近0。改进机制对比结果进一步表明,自适应探索噪声机制和双经验池机制均能提升算法性能。【结论】本文所提改进TD3方法能够在满足物理运行约束的前提下,发挥微电网群内多主体的互济能力,提升了系统协同调度的经济效益与安全性。
[Key word]
[Abstract]
[Objective] To address the strong uncertainty in renewable energy output and load demand, complex operational constraints, and continuous decision variables in the cooperative scheduling of multi-microgrid systems, this paper proposes an economic optimization method based on an improved twin delayed deep deterministic policy gradient (TD3) algorithm. [Methods] Firstly, a three-microgrid cluster model was constructed, incorporating wind power, photovoltaic generation, micro-turbines, energy storage systems, EV charging loads, and main grid interactions. Its economic optimal dispatch process was modeled as a Markov decision process. The state space encompassed load demand, renewable generation, time-of-use electricity prices, and battery state of charge, while the action space comprised energy storage charge/discharge power, micro-turbine output, power exchange between microgrids, and interaction power with the main grid. Next, a reward function integrating normalized operating costs, environmental costs, and constraint penalty terms was designed to guide the agent in learning economical and feasible scheduling strategies. Finally, building on the original TD3 algorithm, a dual experience pool (DEP) mechanism and an adaptive exponentially decaying Gaussian exploration noise (AED-GEN) mechanism were introduced. The DEP mechanism classified experience samples into feasible solution samples and boundary exploration samples, improving sample utilization efficiency and enhancing the learning capability for constraint boundaries. Meanwhile, the AED-GEN mechanism balanced exploration during early training and convergence stability in later stages. [Results] Case studies were conducted on a three-microgrid cluster. The results demonstrated that the improved TD3 algorithm outperformed the deep deterministic policy gradient (DDPG), original TD3, and SAC algorithms in terms of convergence stability, operational economy, and scheduling security. The standard deviation of reward fluctuation in the last 100 training episodes for the improved TD3 was 1.51, representing reductions of 35.74%, 45.09%, and 54.24% compared to SAC, original TD3, and DDPG, respectively. Its total daily operating cost was 4,186.94 yuan, which was 1.38%, 6.31%, and 8.71% lower than that of SAC, original TD3, and DDPG, respectively. Meanwhile, the action limit violation rate dropped rapidly during training and gradually approached zero in the later stages. Further comparisons of the improved mechanisms verified that both the adaptive exploration noise mechanism and the dual experience pool mechanism enhanced algorithm performance. [Conclusion] By satisfying the physical operational constraints, the proposed improved TD3 method leverages the mutual support capability among multiple entities within the microgrid cluster, thereby enhancing the economic efficiency and security of the system’s collaborative scheduling.
[中图分类号]
[基金项目]
黑龙江省重点研发计划(2024ZXJ01A04)