Beyond Captioning: Weaving Visual and Social Dialogue for Robotic Receptionists
Combining Visual and Social Dialogue for Human-Robot Interaction
The paper presents a multimodal conversational AI prototype designed for a hospital receptionist robot. It integrates the award-winning Alana social bot with a visual perception system to handle task-based navigation, medical check-ins, and visually-grounded dialogue alongside open-domain social interaction.
TL;DR
Researchers from Heriot-Watt University have unveiled a multimodal prototype that transforms robots from simple tools into socially-aware assistants. By combining the Alana social bot with visual perception and a Petri-Net-based interaction planner, this system allows a hospital receptionist robot to handle everything from medical check-ins and "Where is my bag?" queries to discussing the latest news and coronavirus quizzes.
Background: The Gap in Situated Interaction
In the world of AI, vision and language are often treated as separate silos or linked through static tasks like "Image Captioning." However, in a real-world setting like a hospital waiting room, a robot needs to be situated. It must understand the physical space (Visual Dialogue) while maintaining the social fabric of the interaction (Social Dialogue).
The SPRING project addresses this by creating a Socially Assistive Robot (SAR) that doesn't just answer questions but acts as a proactive, empathetic participant in a shared environment.
The "Brain" of the System: Multi-Bot Architecture
The core innovation lies in the system's modularity. Rather than a single monolithic model, it uses an ensemble of specialized "bots" managed by a priority-based Dialogue Manager.
1. The Dialogue System (The Green Blocks)
Built on the Alana v2 framework, the system includes:
- Task/Direction Bots: Handles hospital-specific logic like navigation and check-ins.
- Visual Dialogue Bot: Connects language to the physical world.
- Social/Quiz Bots: Prevents the interaction from feeling "robotic" by providing entertainment and information.
2. Social Interaction Planner (The Blue Blocks)
This module acts as the conductor. It solves the "Multi-Threaded Dialogue" problem. If a user asks a question that requires looking at the room, the planner executes a Petri-Net Plan (PNP). This allows the robot to manage state, wait for sensor data, and resume conversation without losing the context of the interaction.
Figure 1: The architecture shows the interplay between the Social Interaction Planner, the Alana-based Dialogue System, and the ROS Vision Action Server.
Seeing the World: Visual Grounding
The robot uses Detectron2 for scene segmentation, which is then translated into a Scene Graph. While the current prototype uses manually refined scene graphs, it enables the robot to answer spatial queries like "Is there a seat available near the door?" or "I left my jacket on the chair, can you see it?"
Figure 2: The web-based demonstration interface showing a user checking in while the robot maintains the dialogue state.
Critical Insight: Why This Matters
Most Task-Oriented Dialogue (TOD) systems fail because they are too rigid—one "off-script" comment and the system breaks. By nesting domain-specific bots within an open-domain social framework (Alana), this research provides a "safety net." If the robot doesn't understand a specific medical check-in intent, it can fall back to social conversation to maintain rapport while it attempts to resolve the task error.
Conclusion and Future Outlook
This work represents a "first step" toward a truly autonomous hospital assistant. The researchers are moving toward:
- Automatic Scene Graph Generation: Removing the need for manual scene rules.
- Multi-party Interaction: Enabling the robot to handle a group of people in a waiting room, not just a 1-on-1 chat.
As these systems move from web interfaces to physical platforms like the ARI robot, the boundary between "computer vision" and "human conversation" will continue to blur, leading to robots that truly understand both the words we say and the world we inhabit.
