Purposes In this paper, considering the current distributed photovoltaic (PV) prosumers’ supply/use dual attributes, a multi-agent active and reactive coodinated voltage control method based on deep reinforcement learning is proposed, which combines the PV prosumers’ flexible active load response with the reactive power support capability of the active distribution network. Methods First, by analyzing the influence of PV access on node voltage and line currents, the active power reverse quota and overvoltage responsibility of distributed PV prosumers were calculated. Second, an incentive policy was constructed based on the quota and responsibility to guide PV prosumers’ flexible active load response, and a multi-agent voltage control model was designed for the active distribution network and multiple PV prosumers. Results The analyses of the IEEE 33-bus system demonstrate that the proposed active and reactive power coordinated voltage control method can effectively reduce voltage violations and power losses caused by excessive PV generation, and improve the revenue distribution of PV prosumers.
分布式光伏(distributed photovoltaic,DPV)因其利用效率高、对环境的负面影响小而广泛接入主动配电网(active distribution network, ADN)。国际能源署的报告中指出,2024年全球光伏发电量已超过2 000 TWh,同比增长约30%,其中装机容量最大的中国累计装机容量已达887 TW[1]。光伏设备在配电网中的分布式接入增加了配电网的不确定性,导致电压越限问题频发,加剧了安全运行风险[2]。分布式光伏并网的日益增多所带来的电压稳定性挑战已成为制约电力系统低碳转型的关键挑战之一。
传统的电压控制是通过控制配电网中的无功设备来实现的[3]。然而,随着新能源渗透率的逐步提高,仅靠无功设备进行电压控制难度越来越大。具体来说,有载分接开关(on-load tap changing, OTLC)和分组投切电容器(switched capacitors, SCs)等传统机械设备动作慢且容易达到运行上限,因此一般与实时控制方法一起应用[4-5]。实时控制设备中静止无功补偿器(static var compensator, SVC)和静止无功发生器(static var generator, SVG)的安装、运维成本问题依然存在。此外,在高阻抗比的配电网络中,相比于提升有功负荷,平抑光伏倒送引起的过电压需要消耗更多的无功功率。研究表明,有功功率控制也能有效改善电压分布[6]。因此,如何有效调动当前配电网中的有功负荷,并将其应用于有功无功协同的电压控制方法,是一个非常值得研究的课题[7-8]。
功率传输分布因子(power transfer distribution factor,PTDF)是一种利用线路参数获得节点注入功率与整个网络潮流之间的关系而简化配电网潮流计算的方法[23]。分布式光伏发电会在节点中注入有功功率,进一步影响潮流进而导致过电压问题,而PTDF能够反映有功功率在支路潮流中的分布以及其影响力。
强化学习的基础模型框架为马尔科夫决策过程(markov decision process,MDP)。而多智能体是多个相互之间存在耦合影响关系的智能体,通过在某个大环境中观测自身实时状态并学习不同动作策略,最终执行动作以寻求自身奖励最大化的过程。在这个过程中每个智能体都会与外部环境进行交互来实现强化学习过程,构成马尔可夫博弈过程(markov game process,MGP)。
本文中采用的算法为TD3(Twin Delayed Deep Deterministic policy gradient),其是在DDPG(Deep Deterministic Policy Gradient)算法上改进得到的一种在线异策式深度强化学习算法[25]。基于TD3算法的多智能体控制过程中,主动配电网智能体以及每个分布式光伏智能体有其各自的两个Actor网络和四个Critic网络。若算法中包含K个分布式光伏智能体,其中,分布式光伏智能体的经验回放池仅包括其自身的状态、动作、奖励以及策略动作后的状态,符合实际配电网运行过程中分布式光伏产消者的信息获取。主动配电网智能体的经验回放池中包括整个多智能体协同模型中分布式光伏智能体和主动配电网智能体的状态、动作、奖励以及策略动作后的状态。主动配电网智能体根据可观测节点的电压状态,对SVC设备输出与激励政策系数进行调整以最小化电压越限、网络损耗以及政策支出。分布式光伏产消者智能体根据所提的功率倒送配额与过电压责任,进行灵活有功负荷响应来调整自己的用电计划,并根据激励政策获得收益。通过所提供的激励政策设计,可以实现分布式光伏产消者的个人收益与配电网系统稳定之间的协调平衡。
ÇAMEI, CASANOVASM, MOLONEYJ.Electricity 2025:Analysis and Forecast to 2027[J].2025.
[2]
HAQUEM M, WOLFSP.A review of high PV penetrations in LV distribution networks:Present status,impacts and mitigation measures[J].Renewable and Sustainable Energy Reviews,2016,62:1195-1208.
[3]
CAOD, ZHAOJ, HUW,et al.Model-free voltage control of active distribution system with PVs using surrogate model-based deep reinforcement learning[J].Applied Energy,2022,306:117982.
NIS, CUIC G, YANGN,et al.Multi-time-scale online optimization for reactive power of distribution network based on deep reinforcement learning [J].Automation of Electric Power Systems,2021,45(10):77-85.
HUD E, PENGY G, WEIW,et al.Multi-timescale deep reinforcement learning for reactive power optimization of distribution network[J].Proceedings of the CSEE,2022,42(14):5034-5044.
[8]
TONKOSKIR, LOPESL A C, EL-FOULYT H M.Coordinated active power curtailment of grid connected PV inverters for overvoltage prevention[J].IEEE Transactions on sustainable energy,2010,2(2):139-147.
FUL D, JIAQ Q, SUNL L,et al.Bi-level voltage regulation strategy for active and reactive power of distribution network with hybridGrid forming/grid following photovoltaic inverters[J].Automation of Electric Power Systems,2025(12):171-183.
HONGL C, WUM H, ZHUJ,et al.Constraint-enhanced safe reinforcement learning based decision-making method for re/active power optimization in highly penetrated PV-storage-charging distribution network[J].Proceedings of the CSEE,2025,45(22):8764-8778.
MAY X, ZHUW Q, GUOY,et al.Cluster voltage control strategy of high permeability photovoltaic distribution network considering participation of energy storage[J/OL].Journal of Electrical Engineering,1-9[2025-04-25].
LIP, LIUJ Y, LIJ W,et al.A voltage control method for distribution networks considering photovoltaicand electric vehicle charging station coordination[J].Journal of Electric Power Science and Technology,2024,39(6):121-130.
[17]
HAYATM A, SHAHNIAF, SHAFIULLAHG M.Replacing flat rate feed-in tariffs for rooftop photovoltaic systems with a dynamic one to consider technical,environmental,social,and geographical factors[J].IEEE transactions on industrial informatics,2018,15(7):3831-3844.
[18]
HAASR, DUICN, AUERH,et al.The photovoltaic revolution is on:How it will change the electricity system in a lasting way[J].Energy,2023,265:126351.
FUJ, SUNY H, YINC X,et al.Low-carbon optimization of park integrated energy system considering dynamic incentive mechanisms[J].Power System Technology,2025,49(11):4638-4648.
KOUL F, WUM, LIY,et al.Optimization and control method of distributed active and reactive power in active distribution network[J].Proceedings of the CSEE,2020,40(6):1856-1865.
HUW H, CAOD, HUANGQ,et al.Application of deep reinforcement learning in optimal operation of distribution network[J].Automation of Electric Power Systems,2023,47(14):174-191.
[27]
CAOD, ZHAOJ, HUW,et al.Data-driven multi-agent deep reinforcement learning for distribution system decentralized voltage control with high penetration of PVs[J].IEEE Transactions on Smart Grid,2021,12(5):4137-4150.
[28]
ZHANGY, WANGX, WANGJ,et al.Deep reinforcement learning based volt-var optimization in smart distribution systems[J].IEEE Transactions on Smart Grid,2020,12(1):361-371.
DENGQ T, HUD E, CAIT T,et al.Reactive power optimization strategy of distribution network based on multiagent deep reinforcement learning[J].Advanced Technology of Electrical Engineering and Energy,2022,41(2):10-20.
[31]
LIUH, WUW.Online multi-agent reinforcement learning for decentralized inverter-based volt-var control[J].IEEE Transactions on Smart Grid,2021,12(4):2980-2990.
[32]
YUP, ZHANGH, HUZ,et al.Voltage control of distribution grid with district cooling systems based on scenario-classified reinforcement learning[J].Applied Energy,2025,377:124415.
[33]
HUD, YEZ, GAOY,et al.Multi-agent deep reinforcement learning for voltage control with coordinated active and reactive power optimization[J].IEEE Transactions on Smart Grid,2022,13(6):4873-4886.
[34]
CHENGX, OVERBYET J.PTDF-based power system equivalents[J].IEEE Transactions on Power Systems,2005,20(4):1868-1876.
MAQ, DENGC H, LONGZ J.A real-time calculation method of voltage sensitivity based on white-box dendritic net[J].Proceedings of the CSEE,2023,44(19):7503-7514.
[37]
FUJIMOTOS, HOOFH, MEGERD.Addressing function approximation error in actor-critic methods[C]//International conference on machine learning.PMLR,2018:1587-1596.