voxagent

Voice-powered terminal agent. Ask questions out loud, get answers. Your audio and your questions never leave your machine.

$ npm install -g voxagent Copied!
voxagent
$ voxagent
voxagent v0.2.0 - voice-powered terminal
Checking Ollama connection...
Ollama connected.
Loading llama3.2...
llama3.2 ready.
Checking whisper model...
Loading whisper model...
Whisper model ready.
Press ENTER to speak, or Ctrl+C to quit.
Recording. It stops when you stop speaking, or press ENTER to stop now.
Captured 4.7s of audio (150824 bytes)
Transcribing...
You: What is the capital of Australia?
Thinking...
The capital of Australia is Canberra.
Press ENTER to speak, or Ctrl+C to quit.
Recording. It stops when you stop speaking, or press ENTER to stop now.
Captured 5.1s of audio (163624 bytes)
Transcribing...
You: How many days in a leap year?
Thinking...
There are 366 days in a leap year.
Press ENTER to speak, or Ctrl+C to quit.

How it works

Three steps, all on your machine. No API keys, and no network request apart from a one-time download of the speech model.

Step 01
🎙️

You speak

decibri captures microphone audio as raw PCM. Cross-platform, prebuilt binaries.

→
Step 02
📝

Whisper transcribes

whisper.cpp converts speech to text locally. GPU-accelerated when available.

→
Step 03
🤖

Ollama responds

Your local LLM generates the answer. Any Ollama model. Default: llama3.2.

Built for developers

A voice interface that respects your machine, your data, and your workflow.

⬡ Runs locally

Speech recognition and the language model run on your hardware. After a one-time download of the speech model, voxagent talks only to Ollama on your machine. No API keys. No accounts.

⬡ Cross-platform

Windows x64, macOS 14 or later on Apple Silicon, and Linux x64. Prebuilt native binaries, so nothing is compiled at install.

⬡ Any model

Works with any Ollama model. Swap models with --model mistral or use the default llama3.2.

⬡ Privacy first

Your audio is held in memory and passed straight to the speech engine on your machine. It is never sent anywhere, and never written to disk unless you turn on debug mode. Your questions go only to Ollama on your machine.

⬡ Open source

Apache 2.0 licensed. Inspect the code, contribute, or fork it. Built in the open on GitHub.

⬡ Debug mode

Run with --debug to see audio levels and raw whisper output. It also saves each recording to debug-capture.wav in the current directory, which stays until you delete it.

Quick start

Three commands. No configuration needed.

# Install voxagent
$ npm install -g voxagent
# Make sure Ollama has a model
$ ollama pull llama3.2
# Run
$ voxagent

Requirements

Windows x64, macOS 14 or later on Apple Silicon, or Linux x64. No cloud accounts.

✓  Node.js 18.3 or later, except Node.js 19
✓  Ollama running locally
✓  A microphone
On Windows, voxagent needs the Microsoft Visual C++ Redistributable. On Linux, it needs glibc 2.38 or later, such as Ubuntu 24.04 or Debian 13, and the packages libvulkan1, libgomp1 and libasound2t64. Intel Macs, and ARM64 on Windows and Linux, are not supported. The first run downloads the whisper model, about 150 MB, to .voxagent/models in your home directory.

Powered by

Built on proven open-source foundations.