Context-Aware Prefaces Reduce Awkward Pauses in Dialogue Robots
Japanese team’s robot uses LLM-generated prefaces to fill conversational gaps naturally.
A two-stage incremental framework for low-latency turn-taking in dialogue robots is proposed. It decouples prefatory-response preparation from speech onset using an intent readiness detector to trigger LLM-based generation of a short prefatory response, while a voice activity projection (VAP) model determines delivery timing. In a field experiment with a route-guidance robot in a shopping mall, both fixed-filler and contextual-preface conditions significantly reduced initial response latency compared to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences, indicating a timing trade-off.
- Two-stage framework uses intent-readiness detector to trigger LLM-based preface generation before speech onset.
- Field experiment in a shopping mall with a route-guidance robot showed contextual prefaces reduce the gap between initial and main response by 30% vs fixed fillers.
- Contextual prefaces had ~200ms longer initial latency but were perceived as equally natural, indicating a tunable trade-off.
Why It Matters
Smoother turn-taking in service robots improves user satisfaction and enables more natural human-robot interactions.