# Voice recognition module for ROS2 foxy

**URL:** <https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902>\
**Category:** ROS General\
**Tags:** ros2, foxy\
**Created:** [February 11, 2021, 9:57am UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902 "2021-02-11T09:57:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ajaykumaar\_S](https://sea2.discourse-cdn.com/flex022/user_avatar/discourse.openrobotics.org/ajaykumaar_s/32/6321_2.png) [@Ajaykumaar\_S](https://discourse.openrobotics.org/u/Ajaykumaar_S)\
**Post date:** [February 11, 2021, 9:57am UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/1 "2021-02-11T09:57:58Z")

</div>

Hi,  
Are there any packages for voice recognition in ROS2 Foxy?  
Preferably to work offline like the PocketSphinx.

Thanks,  
Ajay

---

<div class="post-metadata">

**Author:** ![dignakov](https://sea2.discourse-cdn.com/flex022/user_avatar/discourse.openrobotics.org/dignakov/32/5053_2.png) [@dignakov](https://discourse.openrobotics.org/u/dignakov)\
**Post date:** [February 11, 2021, 3:06pm UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/2 "2021-02-11T15:06:22Z")

</div>

I’m not sure about something ROS2 specific, but there’s Mozilla Deep Speech, which should run offline:

[https://github.com/mozilla/DeepSpeech](https://github.com/mozilla/DeepSpeech)

---

<div class="post-metadata">

**Author:** ![BrettRD](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@BrettRD](https://discourse.openrobotics.org/u/BrettRD)\
**Post date:** [February 22, 2021, 9:13am UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/3 "2021-02-22T09:13:32Z")

</div>

Audio support in ROS is pretty thin on the ground, but there are a few packages that will bind gstreamer pipelines

This ROS1 package looks like it would be pretty easy to port to ROS2:

> **[pocketsphinx - ROS Wiki](https://wiki.ros.org/pocketsphinx)**

Otherwise, a riff on gscam might be doable, the pocketsphinx pipeline is pretty simple.  
`gst-launch-1.0 autoaudiosrc ! audioconvert ! audioresample ! pocketsphinx ! fdsink fd=1`

I’d love to plug my own ros-gstreamer package, but it doesn’t support string payloads yet.

---

<div class="post-metadata">

**Author:** ![Cam\_Buscaron](https://sea2.discourse-cdn.com/flex022/user_avatar/discourse.openrobotics.org/cam_buscaron/32/6761_2.png) [@Cam\_Buscaron](https://discourse.openrobotics.org/u/Cam_Buscaron)\
**Post date:** [August 31, 2021, 6:22pm UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/4 "2021-08-31T18:22:46Z")

</div>

Hi @Ajaykumaar_S,

Not offline. However, at AWS we have build a number of ROS 2 packages to help integrate audio recognition and speech generation with AWS cloud services like Amazon Lex and Polly. You can find them here:

[https://github.com/aws-robotics/tts-ros2](https://github.com/aws-robotics/tts-ros2)

[https://github.com/aws-robotics/lex-ros2](https://github.com/aws-robotics/lex-ros2)

Regards,  
Cam

---

<div class="post-metadata">

**Author:** ![BrettRD](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@BrettRD](https://discourse.openrobotics.org/u/BrettRD)\
**Post date:** [September 8, 2021, 2:07am UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/5 "2021-09-08T02:07:59Z")

</div>

I just got pocketsphinx running with ROS2 and discovered it doesn’t like my accent.  
Fortunately pocketsphinx is not the only speech-to-text package that has gstreamer bindings

[https://github.com/BrettRD/ros-gst-bridge](https://github.com/BrettRD/ros-gst-bridge)

It’s easiest to launch from bash, but that repo also has tools to launch it from roslaunch.

`gst-launch-1.0 --gst-plugin-path=install/gst_bridge/lib/gst_bridge/ autoaudiosrc ! audioconvert ! audioresample ! pocketsphinx ! queue ! rostextsink`

---

<div class="post-metadata">

**Author:** ![danber](https://avatars.discourse-cdn.com/v4/letter/d/e5b9ba/32.png) [@danber](https://discourse.openrobotics.org/u/danber)\
**Post date:** [December 13, 2021, 2:10pm UTC](https://discourse.openrobotics.org/t/voice-recognition-module-for-ros2-foxy/18902/6 "2021-12-13T14:10:06Z")

</div>

Another option would be the voice assistant _Jaco_:

> **[Jaco-Assistant / Jaco-Master · GitLab](https://gitlab.com/Jaco-Assistant/Jaco-Master)**
>
> Master nodes for Jaco Assistant. Main repository of the Jaco Assistant project.

It can run completely offline on most Linux computers and even on a RaspberryPi and supports multiple languages. For a robot project of mine I created an interface skill that allows to control the assistant over the ROS2 topics. You can find it as demo skill [here](https://gitlab.com/DANBER/BabbelAndIhrs).

If you’re only interested in Speech-To-Text without NLU or TTS, the STT module [Scribosermo](https://gitlab.com/Jaco-Assistant/Scribosermo) can also be used as standalone model.
