Research & Papers

New Survey Reveals Hallucination Gap in AI for OSINT and Cyber Investigations

AI promises to revolutionize open-source intelligence but fails to measure its own hallucinations.

Deep Dive

The rapid growth of digital information has made manual OSINT analysis insufficient, pushing large language models (LLMs) and agentic AI systems to the forefront. This survey systematically reviews 74 studies and makes four key contributions. First, it establishes agentic AI as a distinct analytical category, organizing the literature via an 11-category taxonomy covering LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it identifies a critical hallucination-validation gap: although over twenty studies acknowledge hallucination as a major reliability concern, only one OSINT-specific RAG-based system empirically measures it—under non-reproducible conditions. Related studies evaluate general-domain reasoning, not OSINT.

Third, the survey maps existing research to the OSINT lifecycle, revealing strong support for collection and analysis but limited coverage of verification, reporting, dissemination, and decision support. Fourth, it derives a ten-point research agenda addressing evaluation, benchmarking, adversarial robustness, dark-web coverage, multimodal intelligence, and governance. The authors conclude that a human-AI co-pilot model—where LLMs assist collection and triage while analysts retain verification and decision-making—represents the most defensible near-term deployment architecture for real-world cyber investigations.

Key Points
  • Only 1 out of 20+ studies discussing hallucination in OSINT empirically measures it in a RAG-based system, under non-reproducible conditions.
  • 11-category taxonomy organizes the field from LLM foundations to domain adaptation and risk.
  • Research heavily focuses on collection and analysis but neglects verification, reporting, dissemination, and decision support.

Why It Matters

For cybersecurity professionals, AI-assisted OSINT must address the reliability gap before it can be trusted for critical decisions.

📬 Get the top 10 AI stories daily