Research & Papers

Virginia Tech's CriterionSI lets you drag images to teach LLMs your clustering criteria

No predefined categories needed: just drag images to shape AI clustering on the fly.

Deep Dive

A team of Virginia Tech researchers (Yang Liu, Xuxin Tang, Jiahao Xu, Chris North) has published CriterionSI (Criterion-guided Semantic Interaction), a method that grounds large language models in spatial user interactions for iterative image clustering. Traditional dimension reduction and semantic interaction techniques require pre-defined embeddings or explicit criteria, which limits exploratory analysis. CriterionSI breaks this constraint by letting users drag individual images on a 2D projection to indicate their clustering intent. An LLM infers the underlying semantic dimension (e.g., action, location, mood) from the pattern of drags, and then combines this inferred criterion with local drag movements to guide a global reprojection of the entire dataset. This allows the clustering layout to evolve naturally as the user refines their intent through sequential interactions.

Evaluation using simulated user drags demonstrates that CriterionSI can discover and progressively refine a target clustering criterion, producing layouts that align with the user's unspoken mental model. The method replaces fixed prior assumptions with human-provided feedback loops, making clustering more intuitive for non-experts. Code and data are available on GitHub. While still a research prototype, CriterionSI points toward a new paradigm for human-AI collaboration in visual analytics—where LLMs interpret ambiguous spatial gestures to adaptively reshape information spaces without requiring explicit formulaic inputs.

Key Points
  • Uses LLMs to infer clustering criterion (e.g., action, location, mood) from sequential user drags on a 2D image layout
  • Combines inferred criterion with local drag movements to steer a global reprojection of the entire dataset
  • Simulation-based evaluation shows CriterionSI progressively aligns layouts with target criteria, outperforming fixed embedding methods

Why It Matters

Makes AI image clustering intuitive and fluid, adapting to user intent in real time without predefined categories.

📬 Get the top 10 AI stories daily