MIRA
FrameworkDialogue, social behavior, and co-speech motion, coordinated through an embodiment cue.
Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for real-time full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue that routes the response to the appropriate physical behavior. Discrete social behaviors (e.g. listening and greeting) are mapped to validated robot trajectories, while speaking responses are accompanied by streaming, generative co-speech motion.
For co-speech motion generation, we propose ROSCO, a prefix-conditioned diffusion model for streaming audio-to-joint motion generation. We further design RHPC, an inference scheme that maintains a sufficiently long temporal context for motion prediction while bounding physical commitment to a short, interruptible prefix.
At the interaction level, we design CORTEX, a dual-timescale interaction policy that combines low-latency barge-in preemption and streaming response generation with deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints during execution. MIRA is deployed on an Astribot S1 humanoid robot.
Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.
Dialogue, social behavior, and co-speech motion, coordinated through an embodiment cue.
Dual-timescale control for streaming replies, turn admission, and interruption.
Speech-to-joint diffusion that commits only a short, cancellable motion prefix.
Physical tests of motion quality, streaming latency, and interruption safety.
Discrete social behaviors on the Astribot S1 — greeting, listening, and expressive reactions executed from pre-validated action libraries.
Real-time full-duplex interaction on the Astribot S1 — including streaming co-speech motion and user barge-in interruption.
@article{mira2026, title={MIRA: Real-Time Full-Duplex Human-Robot Interaction for Embodied Companions}, author={Lin, Lijian and Zhu, Ye and Zhang, Fan and Liu, Yunfei and Li, Baofeng and Zeng, Xianwen and Wang, Jianan and Li, Yu}, journal={arXiv preprint arXiv:2609.24547}, year={2026} }