DialoStack: task-oriented spoken dialogue for ROS 2, with the LLM kept out of the control flow

I’ve been working on getting a robot to run structured spoken conversations, things like a clinic intake (collect name, age, reason for the visit, allergies…), explaining a topic and checking the person understood it, or running a short quiz. Wiring an LLM straight into the speech loop gets you a nice demo quickly, but in my tests it was hard to rely on: it drifted off task, re-asked things it already had, or never decided the conversation was over.

So I built DialoStack, a set of ROS 2 packages where the dialogue flow is plain Python (states, turn limits, confirmation, cancellation) and the LLM is only used to interpret what the user said and to phrase the next reply.

It started as my bachelor’s thesis, and I presented the work at WAF 2026, the International Workshop on Physical Agents. I’ve put a lot of care into it and I don’t want it to end up as another thesis repo that nobody touches after graduation, so I’m releasing it as open source (MIT) and plan to keep working on it. Feedback from people doing HRI would help a lot.

A slot-filling task in the web monitor: the frame on the right is filled turn by turn and confirmed before the action returns.

Demo on a real NAO (4 min): Dropbox

How it’s used

Everything goes through one action. You describe the task, and if you don’t pass a schema the engine generates one:

ros2 action send_goal --feedback /dialog/execute_task \
  ros2_dialog_interfaces/action/DialogTask \
  "{task_description: 'Take a coffee order: drink type, size and customer name',
    dialog_mode: 'slot_filling', domain: 'coffee shop'}"

The robot asks for whatever is missing, handles corrections (“actually, make it a large”), confirms, and returns the filled frame as JSON in the result. You can also pass your own schema and a partially filled frame to resume an interrupted dialogue. There is a BehaviorTree.CPP leaf node that wraps the same action.

What’s in the repo

  • ros2_dialog_manager: the FSM and three strategies (slot filling, explanation, quiz), with a registry for adding new ones
  • ros2_llm_manager: LLM inference as an action with per-dialogue sessions; Gemini or Ollama behind the same interface, so it can run fully local
  • speech_io: faster-whisper + Silero VAD for STT, Piper for streaming TTS; the mic is muted while the robot speaks
  • vision_io (optional): facial emotion and lip activity, published as topics and added to the prompt context
  • robots/nao: talking gestures and eye-LED feedback, in RViz and on the real robot (via nao_lola)
  • evaluation/: ROS-free benchmarks against hand-annotated data (slot extraction P/R/F1, intent confusion matrix, quiz grading, simulated end-to-end dialogues)

The deterministic core has 173 unit tests that run with plain pytest.

Status and limitations

  • ROS 2 Jazzy on Ubuntu 24.04, built from source. No binaries yet.
  • You don’t need a robot: a microphone and speakers are enough.
  • Local models via Ollama work and keep all data on the machine, but small ones are noticeably less accurate than Gemini.
  • The only embodiment layer so far is NAO. The core just publishes /is_speaking and /user_vad, so mapping those to another robot should be straightforward, but I haven’t done it.

Related work

llama_ros and whisper_ros provide LLM and STT inference as ROS 2 nodes, and ROS4HRI defines common HRI interfaces. DialoStack sits a level above those: it’s the dialogue management layer that decides what to ask next and when a task is done. In principle the STT and LLM nodes could be replaced by those packages, though I haven’t tried it yet.

Feedback I’m looking for

  • Does a single DialogTask action with a free-text task description feel like the right interface, or would you rather define dialogues declaratively (YAML, BT)?
  • If you work with other social robots (Pepper, TIAGo, …), what would you need from the embodiment layer?
  • Dialogue types you’d use that aren’t covered by slot filling / explanation / quiz.

A few issues are tagged good first issue if anyone wants to contribute (an OpenAI-compatible provider, more TTS voices, tests for the quiz and explanation strategies). Bug reports are just as welcome: I’d rather hear what breaks than have it fail quietly on someone’s robot.

Repo: GitHub - aquintan4/DialoStack: Task-oriented spoken dialogue for ROS 2 robots. Deterministic dialogue strategies (slot filling, explanation, quiz) powered by pluggable LLMs (Gemini, Ollama), with full speech and vision I/O. · GitHub
Paper (WAF 2026, “Deterministic Control and Confined LLM Reasoning in Task-Oriented Dialogue for Service Robots”, p. 20): https://waf26.unex.es/wp-content/uploads/2026/08/WAF2026_Proceedings_Final.pdf

1 Like