New paper defines how to engineer trustworthy agentic AI for critical systems
Five dimensions of trustworthiness mapped across power grids, autonomous vehicles, and more
A new survey paper from Omar Al-Refai, Ibrahim Shahbaz, and colleagues (submitted to IEEE) addresses a critical gap in agentic AI research: how to verify, audit, and trust autonomous systems that make decisions with real-world consequences. Instead of evaluating agents purely by task capability, the authors propose a trustworthiness model organized around five cross-cutting dimensions—safety and constraint satisfaction, robustness and reliability, transparency and interpretability, accountability and auditability, and privacy and security. They map these dimensions onto an agentic assurance workflow that spans from perception through audit, covering architectures, threats, concrete trust mechanisms, and quantitative metrics.
The researchers apply their framework to four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks. Across these domains they identify recurring design patterns, shared failure modes, and domain-specific gaps. The paper concludes that agentic AI trustworthiness is a single, cross-cutting problem and outlines a path toward a reusable assurance framework—analogous to the graded certification regimes used in mature safety-critical fields like aviation or nuclear engineering. This work provides both a taxonomy for engineers building agentic systems and a roadmap for regulators and standards bodies.
- Five-dimension trustworthiness model: safety, robustness, transparency, accountability, and privacy/security
- Applied to four critical domains: power systems, autonomous vehicles/UAVs, high-performance computing, and communication networks
- Proposes a cross-domain assurance framework similar to graded certification in aerospace and nuclear engineering
Why It Matters
Provides engineers and regulators a concrete framework to certify agentic AI for high-stakes, real-world deployment.