Abstract:Deep reinforcement learning (DRL) has been recognized as a promising alternative for emergency controls recently, for its rapid decision-making and strong policy searching capabilities. However, when applied to complex large-scale power systems, two prominent challenges emerge>① high-dimensional discrete-continuous hybrid action space poses great adaptation difficulty for regular DRL methods, as they are designed primarily for purely discrete or continuous action spaces; and ② unguided exploration within vast action spaces leads to inefficient convergence performance. To mitigate these limitations, this paper develops a knowledge-guided hybrid DRL-based method for transient stability emergency control. The proposed method adopts a novel hybrid policy to represent the hybrid action domain, and employs a novel derivative-free DRL for optimization, which eliminates the training instability issues associated with jointly optimizing gradients for different action types. Then, the prior domain knowledge is utilized to formulate mask rules and incorporate them into DRL training through trainable action mask (TAM) technique to guide exploration. Moreover, the proposed method is further integrated into a hierarchical DRL framework to alleviate computational complexity and enhance scalability. Case studies on the IEEE 300-bus system show that the proposed method improves solution quality by 52.9% and 42.8% compared with purely discrete and conventional hybrid action DRL methods, and achieves an 80.0% improvement in convergence efficiency compared with the method without knowledge guidance.