Research & Papers

arXiv paper uses LLM-style training to optimize DER coordination

New SRL framework pre-trains on data, fine-tunes in two steps for grid efficiency.

Deep Dive

As renewable energy sources like solar and wind proliferate, the coordination of distributed energy resources (DERs) becomes critical for grid stability and decarbonization. Traditional optimization methods struggle with the inherent uncertainty and modelling complexity of DERs, while standard reinforcement learning (RL) suffers from sample inefficiency and sub-optimality when trained from scratch. In a new paper presented at PSCC2026, Haoyuan Deng from the University of Edinburgh and colleagues propose a Supervised Reinforcement Learning (SRL) framework that borrows the pre-training/fine-tuning paradigm from large language models.

The SRL framework first pre-trains a policy on demonstration data in a supervised-learning fashion, then fine-tunes it with RL using a novel two-step process: offline fine-tuning to enhance policy performance, followed by online fine-tuning to adapt to real-world dynamics. Experiments demonstrate that RL implementations based on this framework significantly beat all benchmarks—including standard RL and optimization-based methods—achieving high cost efficiency even when the demonstration data is of low quality. This work could enable utilities to better harness DER flexibility for grid balancing, reducing reliance on fossil fuel peaker plants and lowering operational costs.

Key Points
  • SRL pre-trains on demonstration data (supervised), then fine-tunes via RL in two phases: offline and online.
  • Outperforms all benchmarks significantly, even with low-quality demonstrations.
  • Inspired by large language model training (pre-train → fine-tune) applied to power grid coordination.

Why It Matters

More efficient DER coordination accelerates decarbonization and reduces grid operational costs for utilities.

📬 Get the top 10 AI stories daily