Image & Video

New MLLM system boosts forensic image retrieval for tattoos and face sketches

Multimodal fusion improves retrieval precision by combining text and visual embeddings.

Deep Dive

A team of researchers led by Ricardo González-Gazapo has introduced a unified retrieval framework designed to bridge the modality gap in forensic image analysis. The system tackles four critical forensic tasks: matching a tattoo from a query image, retrieving tattoos based on verbal witness descriptions, finding tattoos from hand-drawn sketches, and identifying faces from forensic sketches. By leveraging a multimodal large language model (MLLM), the framework automatically generates structured textual descriptions for both query images and gallery images. These descriptions are then embedded using sentence-transformer models for text-based comparison, while visual features are extracted using state-of-the-art encoders tailored to each task.

The key innovation lies in a fusion strategy that combines text- and image-based similarity scores. Across experiments, the multimodal fusion consistently improved retrieval precision and robustness compared to using either modality alone. This advantage is most pronounced in challenging real-world scenarios—such as partial tattoos, rough sketches, or incomplete witness statements—where visual information is limited or noisy. The paper demonstrates how modern MLLMs can operationalize tasks that traditionally relied on manual expert analysis, offering a scalable and automated pipeline for forensic investigations involving tattoos, facial composites, and witness descriptions.

Key Points
  • Framework supports four forensic tasks: tattoo retrieval from image, text description, sketch, and face retrieval from sketch.
  • Uses MLLM to auto-generate structured text descriptions for queries and gallery items, then sentence-transformer embeddings for comparison.
  • Multimodal fusion of visual and text embeddings boosts precision, especially when evidence is partial or noisy (e.g., sketches, fragmented witness statements).

Why It Matters

Automates manual forensic analysis of tattoos and facial composites, enabling faster, more reliable evidence matching in criminal investigations.

📬 Get the top 10 AI stories daily