Research & Papers

STRAND method turns topological data into survival analysis for ML

New approach treats persistence diagrams as time-to-event data, enabling both testing and vectorisation

Deep Dive

Topological data analysis (TDA) has long struggled with two problems: persistence diagrams are not naturally vectorised for machine learning, and statistical tests for comparing diagrams have evolved separately from those vectorisation methods. A new paper from researchers Juliette Murris, Bernadette Stolz, and Karsten Borgwardt bridges that gap with STRAND (Survival Topological Representation ANalysis of Diagrams). The key insight: treat each topological feature's persistence value (birth-to-death distance) as a fully observed time-to-event, exactly like survival analysis in biostatistics. The persistence survival function S(t) = P(p > t) becomes a single, interpretable object that simultaneously powers hypothesis testing, effect size estimation, and feature vector construction.

STRAND delivers three concrete contributions from this unified representation. First, a non-parametric two-sample test that maintains calibrated Type I error rates while achieving high statistical power even with a small number of diagrams—critical for expensive experimental settings like fMRI. Second, it provides interpretable effect sizes, letting researchers understand not just whether two sets of diagrams differ, but how much. Third, it produces a 1-Wasserstein-stable feature vector ready for downstream ML tasks. The authors validate calibration and power on synthetic manifolds with controlled topology, then demonstrate competitive vectorisation across 14 graph and 3D point cloud benchmarks. To show real-world utility, they apply STRAND to study functional brain connectivity in fMRI data. The method claims to be the first to offer hypothesis testing and vectorisation from a single coherent representation.

Key Points
  • Treats persistence diagrams as survival data, defining a persistence survival function S(t) = P(p > t)
  • Provides a calibrated non-parametric two-sample test with high power from few diagrams
  • Achieves competitive vectorisation on 14 graph and 3D point cloud benchmarks

Why It Matters

Unifies statistical testing and ML vectorisation for topological data, simplifying analysis in fields like neuroscience.

📬 Get the top 10 AI stories daily