Data sharing in robotics: When pooling data hurts competition and profits
New theory reveals that pooling deployment data can backfire under competition.
A new paper from arXiv (arXiv:2607.00168) formalizes the strategic tension in data sharing for 'learning-by-deploying' industries—where each product deployed generates real-world data that directly improves future versions. Unlike classic 'learning-by-doing' (firms get better through production experience), the new framework focuses on data generated by units in the field (e.g., robot fleet telemetry, autonomous vehicle sensor logs). Authors Yunjin Tong and Luca-Andrei Manea build a two-period model with symmetric firms making irreversible capacity choices, where data feeds a shared learning curve that boosts future productivity.
Crucially, the model reveals that under downstream Cournot competition, pooling data across firms can depress prices enough to make the private value of sharing negative—firms may be better off keeping data siloed. The paper identifies a sustainability threshold governed by the elasticity of industry demand, offering a theoretical basis for when mandatory data sharing promotes or harms overall welfare. The results have direct implications for robotics, autonomous vehicles, and other AI-driven sectors, suggesting that regulators must carefully calibrate data-pooling mandates to industry demand conditions.
- Firms in learning-by-deploying industries (e.g., robotics, autonomous vehicles) generate data that improves future models
- Under Cournot competition, pooling deployment data can reduce profits by depressing market prices
- Sustainability of data sharing depends on industry demand elasticity—higher elasticity makes sharing more viable
Why It Matters
Provides a new economic framework for regulators deciding data-sharing policies in robotics, AI, and autonomous systems.