Research & Papers

Deep RL beats traditional models for portfolio optimization across market crises

New MORP-DRL framework reduces downside risk by up to 30% during COVID and post-COVID regimes...

Deep Dive

Traditional portfolio optimization often relies on static models that fail to capture sequential decision-making, tail risk, and market frictions like transaction costs. To address this, researchers from IIT Kanpur and associated institutions introduce MORP-DRL (Multi-Objective Reliability-based Portfolio Optimization via Deep Reinforcement Learning). The framework uses Proximal Policy Optimization (PPO) to jointly maximize expected return and minimize downside risk using three complementary risk measures: variance, Conditional Value-at-Risk (CVaR), and Entropic Value-at-Risk (EVaR). Asset returns are modeled with GARCH(1,1) for volatility clustering, Extreme Value Theory for heavy tails, and a t-copula for dependence structure. Quasi-Monte Carlo simulation generates realistic scenarios, while the optimization incorporates practical constraints like portfolio bounds and transaction costs.

Experiments on ten global equity indices—including S&P 500, Nikkei 225, and FTSE 100—spanning pre-COVID (2017-2019), COVID (2020-2021), and post-COVID (2022-2026) regimes show that MORP-DRL consistently achieves competitive risk-return profiles. During market stress (COVID period), it reduces tail risk significantly compared to the NSGA-II baseline, while also scaling efficiently to portfolios with hundreds of assets. The paper demonstrates that reinforcement learning can dynamically adapt to changing market conditions, offering a robust alternative to static multi-objective optimization. This work bridges deep RL and financial engineering, with potential applications in robo-advisory and algorithmic trading.

Key Points
  • Uses PPO with transaction costs and portfolio bounds, outperforming NSGA-II in downside risk
  • Tested on 10 global equity indices across pre-COVID, COVID, and post-COVID market regimes
  • Combines GARCH(1,1), Extreme Value Theory, and t-copula to model heavy-tailed returns and dependencies

Why It Matters

AI-driven portfolio optimization adapts to market regimes, reducing tail risk and improving scalability for institutional investors.

📬 Get the top 10 AI stories daily