Image & Video

PDSC framework pairs VLMs with latent diffusion for personalized image transmission

Receiver history shapes semantic tokens, boosting personalization while cutting bandwidth for wireless image delivery.

Deep Dive

Semantic communication (SC) promises to slash bandwidth by transmitting meaning, not raw pixels. But most SC systems are user-agnostic, ignoring what the receiver actually cares about. A new paper from researchers Nan Li, Li Zhou, Haijun Wang, Jun Xiong, Haitao Zhao, and Jibo Wei introduces PDSC (Personalized Digital Semantic Communication), which tailors image transmission to individual receivers. The framework combines a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. The encoder analyzes both the source image and the receiver's historical interactions to produce personalized semantic tokens, which are then vector-quantized into discrete indices and packed into a compact fixed-length bitstream. This bitstream is fully compatible with digital transmission systems, making it practical for real wireless channels. On the receiving end, the LDM decoder reconstructs an image conditioned on those tokens, effectively generating a version of the image that emphasizes the receiver's known preferences.

The authors also formalize the problem as a capacity-constrained personalized semantic rate-distortion optimization, introducing a distortion metric that jointly measures source-semantic fidelity and user-preference alignment. Experiments demonstrate that PDSC achieves superior source-semantic consistency and personalization compared to state-of-the-art SC baselines, including CDDM and MoS, especially under strict bandwidth limits. By integrating VLMs for semantic understanding and diffusion models for generation, PDSC points toward a future where wireless image delivery is not just efficient but deeply user-aware. The work was accepted by IEEE GLOBECOM 2026, a top venue for communications research, signaling strong peer validation.

Key Points
  • VLM-based encoder extracts personalized semantic tokens from both the source image and the receiver's historical interactions
  • Vector-quantized tokens form a compact fixed-length bitstream, enabling digital transmission
  • Outperforms CDDM and MoS in source-semantic consistency and personalization under bandwidth-limited conditions

Why It Matters

Makes wireless image transmission bandwidth-efficient and user-adaptive, a key step toward practical 6G semantic communications.

📬 Get the top 10 AI stories daily