Decentralized AI learns without rewards, 90% become experts dynamically
A new paper shows agents can self-organize expertise without any central reward signal.
Florin Neagu's new paper, "Decentralised Consensus Learning Networks: SME Rotation Without Centralised Reward," challenges the dominance of centralized reward signals in modern AI. Instead of relying on an external definition of correct knowledge, the framework enables agents to update beliefs via weighted social consensus. Trust is allocated based on peer consistency rather than ground truth, and subject-matter expert (SME) status is assigned dynamically as a top-percentile competence rank. This creates a system where expertise naturally rotates through the network without any central orchestrator.
Testing across 84 simulation runs with agent counts from 30 to 10,000 and multiple graph topologies, Neagu found that SME rotation is robust, persistent, and scale-invariant. In Phase 1, 90-100% of agents attained SME status, with most turnover occurring after belief convergence. Phases 2 and 3 introduced vector beliefs and revealed five distinct dynamical regimes. At high dimensionality (D=150-200), the network reached stable partial consensus while expertise became concentrated in a single agent—an emergent property driven by belief complexity, not noise. This work opens a path toward truly decentralized AI systems that self-organize expertise without a central reward function.
- New decentralized framework uses peer consensus instead of centralized rewards to allocate expertise.
- In 84 simulations with up to 10,000 agents, 90-100% achieved SME status dynamically.
- At high belief dimensionality (D=150-200), expertise concentrates in one agent — an emergent property of complex consensus spaces.
Why It Matters
Decentralized AI could eliminate single-point-of-failure reward systems, enabling robust, self-organizing multi-agent networks.