Voice-powered terminal agent. Ask questions out loud, get answers. Your audio and your questions never leave your machine.
Three steps, all on your machine. No API keys, and no network request apart from a one-time download of the speech model.
decibri captures microphone audio as raw PCM. Cross-platform, prebuilt binaries.
whisper.cpp converts speech to text locally. GPU-accelerated when available.
Your local LLM generates the answer. Any Ollama model. Default: llama3.2.
A voice interface that respects your machine, your data, and your workflow.
Speech recognition and the language model run on your hardware. After a one-time download of the speech model, voxagent talks only to Ollama on your machine. No API keys. No accounts.
Windows x64, macOS 14 or later on Apple Silicon, and Linux x64. Prebuilt native binaries, so nothing is compiled at install.
Works with any Ollama model. Swap models with --model mistral or use the default llama3.2.
Your audio is held in memory and passed straight to the speech engine on your machine. It is never sent anywhere, and never written to disk unless you turn on debug mode. Your questions go only to Ollama on your machine.
Apache 2.0 licensed. Inspect the code, contribute, or fork it. Built in the open on GitHub.
Run with --debug to see audio levels and raw whisper output. It also saves each recording to debug-capture.wav in the current directory, which stays until you delete it.
Three commands. No configuration needed.
Windows x64, macOS 14 or later on Apple Silicon, or Linux x64. No cloud accounts.
Built on proven open-source foundations.