Developer Tools

Kitchenham et al.'s GUEST guidelines for safe GenAI-assisted literature reviews

GenAI can't replace human reviewers, but here's how to use it safely.

Deep Dive

A team of five software engineering researchers—Barbara Kitchenham, Sebastián Pizard, Lech Madeyski, Ronnie de Souza Santos, Martin Shepperd, and David Budgen—has published preliminary guidelines for using and evaluating generative AI (GenAI) and large language models (LLMs) in systematic literature reviews (SLRs). The 58-page preprint, available on arXiv, synthesizes a rapid review of existing guidelines, thought experiments, and the authors' own experience conducting SLRs. The result is a set of 11 tables and 2 figures that form the GUEST framework (GenAI Use and Evaluation in SLR Tasks), covering planning, conduct, and reporting phases. The authors stress that while GenAI can summarize text and speed up repetitive steps like screening or data extraction, it cannot yet meet the rigour, reliability, and transparency required for unsupervised systematic studies. Human oversight remains mandatory.

GUEST provides concrete recommendations for researchers who either want to use GenAI to assist their own SLRs or to evaluate how well a given tool performs on SLR tasks. The guidelines highlight common pitfalls—such as hallucinated citations, lack of reproducibility, and insufficient transparency—and propose checklists to mitigate them. For instance, any GenAI-assisted SLR should document which models were used, which prompts were applied, and how outputs were verified. The paper also suggests using GenAI for complementary validation: after a human completes a complex analysis, the model can check for consistency. The authors conclude that with GUEST, software engineering researchers can produce trustworthy SLRs that leverage AI efficiency without sacrificing scientific integrity, and they invite the community to test and refine the recommendations.

Key Points
  • GUEST framework spans 58 pages with 11 tables covering planning, conduct, and reporting phases
  • Authors explicitly state GenAI cannot run unsupervised systematic literature reviews—human oversight is mandatory
  • Recommendations include documenting model names, prompt versions, and verification steps to ensure reproducibility

Why It Matters

Gives researchers a practical, evidence-based playbook to harness GenAI without compromising review quality or trust.

📬 Get the top 10 AI stories daily