Research & Papers

MeRLa meta-learns reward shaping for RLHF, boosting LLaMA-3-8B to 90.8% win rate

New framework cuts training instability by 41% while outperforming PPO, DPO, GRPO, and DAPO...

Deep Dive

Reinforcement Learning from Human Feedback (RLHF) relies on static, task-agnostic reward models that often provide sparse learning signals, limiting alignment quality. Researchers introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function across auxiliary tasks before RLHF training. This produces a composite reward that preserves policy optimality while delivering dense, task-specific signals. The meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence, with theoretical guarantees for policy invariance and analysis of representation drift.

Experiments on LLaMA-3-8B across four benchmarks show MeRLa consistently outperforms PPO, DPO, GRPO, and DAPO. It achieves a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability. The framework also retains its benefits when combined with process-based and rubric-based enhanced rewards. By addressing incentive misalignment and providing robust shaping, MeRLa offers a scalable path to more efficient and stable LLM alignment.

Key Points
  • MeRLa achieves a 90.8% length-controlled win rate on AlpacaEval 2.0 and 9.14 on MT-Bench using LLaMA-3-8B.
  • Training instability is reduced by 41% compared to standard RLHF methods like PPO, DPO, GRPO, and DAPO.
  • The meta-learned shaping function provides task-specific learning signals while preserving policy optimality through potential-based conservation.

Why It Matters

MeRLa makes RLHF alignment more efficient and stable, enabling safer and more capable LLMs with less training overhead.

📬 Get the top 10 AI stories daily