Agent Frameworks

Pezego-HITL: new LLM architecture achieves 94% policy alignment in Ghanaian agriculture

Cut latency by 55% while maintaining expert oversight—AI for smallholder farming just got safer.

Deep Dive

Pezego-HITL is a novel large language model architecture developed by researchers from multiple institutions to address the high-stakes challenge of crop protection advice in smallholder agriculture. Unlike generic LLM deployments, Pezego-HITL formalizes safety-compliance, helpfulness, latency, and expert supervision workload as an adaptive compute allocation problem. The accompanying P-EVAL (Policy-grounded Expert-calibrated VALidation protocol) framework provides a unified way to evaluate policy-constrained decision support. On a simulated database of 1,240 field queries, Pezego-HITL achieved a Policy Alignment Rate (PAR) of 0.94 and an Agronomic Utility Rate (AUR) of 0.95, with gold-standard human expert agreement at Cohen's κ = 0.77. The memory-routed architecture reused 59.6% of cached responses, reducing the 95th percentile latency by 55%—from 28.6 seconds down to 12.9 seconds.

The architecture's generalisability was demonstrated by swapping in the open-source Qwen3.5-9B-DeepSeek-V4-Flash model, which still reached a PAR of 0.86 with a 54.5% latency reduction (to 10.2 seconds). To evaluate real-world integration, the team administered detailed questionnaires to 30 Ghanaian Extension Services Officers and 36 smallholder farmers. The results showed that policy-grounded structured retrieval-augmented generation combined with validated-memory routing makes the safety-utility-latency trade-offs explicit and manageable. Pezego-HITL offers a scalable template for trustworthy AI-driven extension services in smallholder farming systems, directly addressing the need for context-aware, safe, and fast decision support.

Key Points
  • Introduces P-EVAL, a policy-grounded evaluation protocol that measures safety, utility, latency, and expert workload simultaneously.
  • Memory-routed RAG architecture achieves 0.94 Policy Alignment Rate and 55% lower P95 latency (12.9s vs 28.6s) on 1,240 field queries.
  • Validated in Ghana with 30 extension officers and 36 farmers; also tested on open-source Qwen3.5-9B-DeepSeek-V4-Flash with 0.86 PAR.

Why It Matters

Explicitly balancing safety, utility, and latency brings trustworthy AI decision-support to smallholder farming at scale.

📬 Get the top 10 AI stories daily