Research & Papers

New 256-model family for clustering skewed matrix data

Clustering high-dimensional skewed data just got 256 times more efficient.

Deep Dive

Clustering high-dimensional skewed data, such as image matrices, often suffers from over-parameterization—too many parameters relative to data points. Existing bilinear factor analyzers help reduce dimensionality, but further constraints across clusters can yield even more parsimonious models. In a new arXiv preprint, researchers Jacob Moore and Michael P.B. Gallaugher introduce a comprehensive family of 256 such models built for skewed matrix variate data, specifically leveraging the skew t distribution for robust handling of asymmetry and heavy tails.

The core contribution is a structured approach to constraining parameters across clusters, producing a spectrum of models from highly flexible (many parameters) to highly constrained (few). They detail an AECM (Alternating Expectation-Conditional Maximization) algorithm for efficient parameter estimation. Through extensive simulations and real-world tests on the MNIST handwritten digits and Olivetti faces datasets, the method demonstrates competitive clustering performance while dramatically reducing the number of free parameters. This makes it particularly useful for applications where data is limited or computational resources are scarce.

Key Points
  • Family of 256 parsimonious models for mixtures of skewed matrix variate bilinear factor analyzers
  • Uses skew t distribution to model asymmetry and heavy tails in high-dimensional data
  • Tested on MNIST and Olivetti faces datasets, showing effective clustering with reduced parameters

Why It Matters

A practical toolbox for clustering image and sensor data where matrices are skewed and parameters are few.

📬 Get the top 10 AI stories daily