Research & Papers

Researchers solve data arbitrage with Blackwell dominance pricing model

New paper prevents buyers from cheaply reconstructing expensive datasets or models.

Deep Dive

Data markets sell two main products—datasets and machine learning models—both replicable at negligible cost. Sellers typically version them through query access and noisy releases, but this immediately creates an arbitrage problem: a buyer can purchase several cheap, less informative queries and combine them to reconstruct a more informative (more expensive) product, bypassing the intended pricing. Existing work either ignored arbitrage-freeness or treated buyer values as exogenous. Wu et al. bridge the gap by studying a seller's optimal pricing problem where buyers value data via Bayesian decision making, and they impose arbitrage-freeness constraints.

The authors first formulate the general arbitrage-free information selling problem, showing it is NP-hard. They then present a branch-and-bound algorithm based on McCormick relaxations to solve it. For the special case of threshold utilities (buyers only value experiments that are sufficiently informative), they prove that arbitrage-freeness can be characterized by Blackwell dominance—a powerful ordering of information structures. This unifies prior separate results for query pricing (Dehghani et al. 2017) and model pricing (Chen et al. 2019). Finally, the paper characterizes revenue-maximizing pricing schemes under restricted query and model menus, offering practical guidance for data marketplaces.

Key Points
  • Unifies query and model pricing under a single Blackwell-dominance condition for arbitrage-freeness.
  • Proves the general arbitrage-free pricing problem is NP-hard, then provides a branch-and-bound algorithm.
  • Characterizes revenue-optimal pricing for restricted menus when buyers have threshold utilities.

Why It Matters

Plugging data arbitrage loopholes could reshape how AI models and datasets are priced commercially.

📬 Get the top 10 AI stories daily