Research & Papers

FinSkillBench shows curated skills boost AI investment agents 44%

Curated procedures lift agent scores from 0.366 to 0.528—but self-generated skills fail.

Deep Dive

FinSkillBench, a new benchmark from researchers Jermyn Zhen Yong Bek, Zhuang Qiang Bok, and Zhongtian Sun, tackles a core problem in agentic AI: proving that language model agents can do more than generate plausible text. The suite evaluates whether agents can retrieve point-in-time data, assemble correct computational inputs, invoke specialized financial methods, and produce auditable structured outputs across investment management tasks. It spans three domains—portfolio construction, risk management, and fundamental analysis—with 12 subtasks and 2,603 task episodes, each containing point-in-time inputs and hidden ground truth.

Across 9 models, the results are striking. When agents used curated skill packages (procedural documents plus executable components), mean scores jumped from 0.366 to 0.528—a 44% relative improvement, strongest in portfolio construction and risk management. By contrast, self-generated skills, where agents write their own procedures, offered negligible gains despite higher computational cost. An independent replication using the Hermes Agent framework across 8 models and 5,280 episodes confirmed the same pattern, suggesting that access to reliable procedural skills can matter as much as model choice. The team released the benchmark, evaluation tools, curated skills, and full trajectories for further research.

Key Points
  • FinSkillBench spans 3 investment domains, 12 subtasks, and 2,603 task episodes
  • Curated skills raised mean scores from 0.366 to 0.528 across 9 models, a 44% gain
  • Self-generated skills showed little benefit; Hermes Agent replication (8 models, 5,280 episodes) confirmed the pattern

Why It Matters

For investment firms deploying AI agents, well-designed procedural skills—not raw model power—drive performance.

📬 Get the top 10 AI stories daily