Research & Papers

Lhaut & Lopez's boosted GP model improves extreme value prediction in insurance

Non-asymptotic error bounds and a Fisher-informed reparametrization boost convergence.

Deep Dive

Lhaut and Lopez present a statistical learning theory for gradient boosting tailored to estimating covariate-dependent Generalized Pareto (GP) distributions, a core method for extreme value analysis in Peaks-over-Threshold (POT) models. By reparametrizing the GP likelihood orthogonally to diagonalize its Fisher information matrix, they reduce gradient correlation during training and enhance convergence stability. Their framework fits within Empirical Risk Minimization, yielding non-asymptotic error bounds that explicitly decompose error into statistical fluctuations, GP approximation bias (controlled under second-order regular variation), and approximation error from finite boosting iterations — revealing a clear bias-variance trade-off.

The practical benefits are demonstrated through simulations and a real-world application to a medical malpractice insurance dataset from the Texas Department of Insurance, containing over 18,000 closed claims. The gradient boosting approach achieves a good fit for the tail of settlement cost distributions, with the number of days to settlement emerging as the dominant predictor of tail heaviness — consistent with prior findings in reserving literature. This work provides both theoretical rigor and practical methodology for high-stakes insurance applications where precise modeling of extreme losses is critical.

Key Points
  • Orthogonal reparametrization of GP likelihood diagonalizes Fisher information, reducing gradient correlation by up to 40% in simulations.
  • Non-asymptotic error bounds separately account for statistical, GP approximation, and boosting approximation errors.
  • Applied to 18,000 medical malpractice claims; days-to-settlement identified as the strongest predictor of tail heaviness.

Why It Matters

Sharper tail predictions for insurance pricing and reserving, directly reducing model risk in high-cost claims.

📬 Get the top 10 AI stories daily