Developer Tools

AI test generation survey: no approach meets all six quality dimensions

21 studies analyzed across three eras — hallucination and traceability remain unsolved.

Deep Dive

Software testing remains one of the most time-consuming and expensive activities in development. Generating test cases directly from natural language requirements could shift validation earlier in the pipeline, but the ambiguity of human language has made full automation difficult. Recent advances in AI, NLP, and large language models (LLMs) have made this goal more accessible, yet introduce new risks such as hallucination, reduced traceability, and inconsistent evaluation. To map the landscape, Folorunsho and Reza conducted a systematic review following Kitchenham and Charters’ guidelines, searching major databases for studies spanning 2000–2025. After strict filtering, they identified 21 primary studies and organized the literature into three evolutionary eras of AI-based test generation.

Their analysis reveals a critical gap: no existing approach simultaneously satisfies all six quality dimensions — automation, ambiguity handling, domain applicability, traceability, evaluation thoroughness, and hallucination control. Most studies excel in one or two areas but fall short on others, particularly traceability and hallucination mitigation. The survey contributes a three-era evolutionary synthesis, a six-criteria gap analysis, and four actionable research guidelines targeting hallucination, traceability, complexity sensitivity, and compliance. For practitioners, the findings underscore that while AI-driven test generation is promising, production-ready solutions must still overcome fundamental reliability and transparency challenges.

Key Points
  • Reviewed 21 primary studies from 2000–2025 using Kitchenham and Charters' systematic review methodology.
  • Three evolutionary eras identified, but no current approach fulfills all six quality dimensions simultaneously.
  • Four research guidelines proposed targeting hallucination, traceability, complexity sensitivity, and compliance.

Why It Matters

For dev teams relying on automated test generation, this roadmap shows where current AI tools fall short and where to invest next.

📬 Get the top 10 AI stories daily