MemeBuddy turns memes into dialog audio for blind users
New system converts static memes into two-speaker audio conversations, boosting engagement for blind users.
Online memes are a cornerstone of modern digital culture, but they remain largely inaccessible to blind and visually impaired users. Prior research attempted to bridge this gap with auto-generated descriptive captions, which often fall flat by missing the humor, narrative structure, and cultural nuances that make memes engaging. MemeBuddy, a new system from researchers Chirag Bhansali, Vikas Ashok, and Hae-Na Lee, tackles this problem by transforming image memes into structured dialog audio. It models each meme as a conversation between two speakers, extracting visible text and using a multimodal large language model (LLM) to implicitly infer context — such as recognizing common meme templates, character roles, and cultural references. The result is a multi-turn audio representation that conveys intent, timing, and implicit meaning through natural conversational interaction.
To validate their approach, the researchers conducted a user study with 14 blind participants, comparing MemeBuddy's dialog-style audio against traditional caption-style descriptions. Results consistently showed that dialog representations significantly improved engagement and user satisfaction, while comprehension stayed comparable to captions. This suggests that converting static visual humor into dynamic, character-driven audio can make meme experiences more natural and enjoyable for blind users. While the sample size is modest, the findings point to a promising direction for making visual internet culture — from memes to comics and advertisements — accessible through personalized, narratively rich audio formats. If scaled, MemeBuddy could transform how blind users participate in the humor and communication that defines modern online communities.
- MemeBuddy uses a multimodal LLM to detect meme templates and cultural references, generating role-based dialog between two speakers.
- In a study with 14 blind participants, dialog-style audio outperformed caption-style descriptions in engagement and satisfaction.
- Comprehension rates were comparable between dialog and caption formats, showing no loss in understanding despite richer audio.
Why It Matters
MemeBuddy makes online humor accessible to blind users by converting visual memes into engaging narrative audio.