AI Safety

GPT-5.2 outperforms GPT-4o in team assessment for cybersecurity exercises

Clustering and LLMs automate grading of team problem-solving in tabletop exercises.

Deep Dive

Researchers compared clustering and LLMs (GPT-4o, GPT-5.2) for assessing 81 participants across two countries in tabletop exercises. Clustering proved valid and reliable with low computational cost, grouping teams by similar approaches. GPT-5.2 showed considerably lower error than GPT-4o against instructor rubrics. Both methods are now integrated into the open-source INJECT platform.

Key Points
  • Clustering method groups teams by similar task approaches, enabling fast, targeted feedback with low computational overhead.
  • GPT-5.2 achieved considerably lower error than GPT-4o when scoring team communication against instructor rubrics.
  • Both methods are now integrated into the open-source INJECT platform for scalable, automated team assessment.

Why It Matters

Automated team assessment in cybersecurity exercises saves instructors time and provides faster, more scalable feedback to learners.

📬 Get the top 10 AI stories daily