erm: Bring Voice, Media and Operational Knowledge to Your Desktop
Keep a small command from becoming another context switch
Pause the music. Find a track. Ask a question about an indexed document. These are modest interactions, but they repeatedly pull attention away from the task on screen. erm brings them into a supervised Erlang desktop workspace with voice input, media control, optional local model assistance and speech output.
The clearest first audience is a Linux developer or operator willing to run the local audio and GTK dependencies. Start with a few observable commands on your own workstation. That is enough to evaluate whether the interaction helps without treating the whole desktop as an autonomous agent.
Understand the parts of a voice interaction
A useful voice interface has several jobs. It captures speech, decides whether the user is addressing it, turns the request into an allowed action, performs that action and reports the outcome. Separating those jobs makes a failure easier to locate.
| Component | Role in the interaction |
|---|---|
| Capture and Whisper integration | Turn microphone audio into text |
| Voice coordinator and intent parser | Identify a command and manage its progress |
| MPV adapter and playlist | Apply media controls to the player |
| Optional Ollama model | Help interpret unfamiliar phrases or answer a question |
| Piper speech service | Speak the result using a configured local voice |
| GTK/Lens interface | Provide a visible desktop interface alongside voice |
For a workstation user, the first useful test is simple: say a command and observe the correct effect. A transcript alone is only one stage of that interaction. The table also explains why a working recogniser can coexist with a missing playback or speech-output dependency.
Common commands have a direct path
The voice-intent module recognises play, pause, stop, next, previous, volume, track selection and player visibility before asking a model to classify an unknown phrase. Volume input is bounded. Known negated commands are refused. The media adapter sends typed actions to the existing MPV process owner and uses the playlist state for track selection and progression.
The voice coordinator handles transcript boundaries, command windows, deduplication and asynchronous work. Model planning and media actions have timeouts; they do not run as blocking work inside the coordinator's message handler. The operating benefit is that command state can be inspected and failures can be reported through the same service interface.
Try a small command set such as βBob, pause musicβ, βBob, next songβ and βBob, volume 30β after configuring the trigger. These are pilot utterances, not measured recognition results. Check actual playback and volume after each command, including repetitions and silence around the wake phrase.
Model assistance has a defined action vocabulary
For phrases outside the direct parser, the code requests a bounded JSON intent and validates it against built-in or operator-configured actions. Model output does not choose arbitrary Erlang module/function pairs. Custom callbacks come from trusted configuration.
That is a useful extension boundary for a workstation tool. A maintainer can add a named action with an explicit implementation and test its effect. The callback still carries whatever authority its implementation has; a voice phrase or speaker match should not be treated as approval for payments, secret access or other sensitive operations.
The configured Ollama path can answer a general question or assist with an intent. Local processing depends on the actual configuration and installed models. The code also supports pulling a missing configured model, so setup and downloads are separate from an offline-use claim.
Ask indexed knowledge, with the retrieval boundary visible
The βask ECAIβ route retrieves up to three sources from the configured disk corpus and sends bounded excerpts to the answer model. Missing corpus configuration and empty retrieval have explicit error paths.
This is a concrete connection between erm and ECAI, but it uses
ecai_ollama_rag:retrieve_sources/3. It does not call the corpus/principal
private bridge. Use an appropriate public or non-sensitive corpus for this
pilot. Private voice access needs an explicit authenticated integration with
the private retrieval boundary.
Speech and service health are part of the interface
The TTS service exposes native Piper speech, voice selection, volume controls, repeat, cancellation and diagnostics. Model installation is handled by an asynchronous service with manifest-driven file checks and immutable cache directories. Listening suppression is used around speech output to reduce feedback into the recogniser.
The native voice coordinator also handles enrolment, speaker scores and
model preparation. In the referenced implementation it can request missing
configured models through erm_model_pull, report progress and retry failed
preparation. The model service checks file sizes and SHA-256 digests before
using its cache. Initial downloads require connectivity.
Native capture and speech sources live in the repository's apps/erm/c_src/
directory. They still need a compatible local build, installed models and
audio devices. No microphone or playback test was run for this article.
Evaluate speaker matching with recordings, other speakers and replay attempts
before assigning it a security role; a speaker score alone does not establish
that a live authorised person is issuing the command.
In a running configured release, these inspection calls help separate a command problem from a missing model or disconnected speech backend:
erm_voice:status().
erm_voice:healthcheck().
erm_tts:diagnostics().
Readiness is not an end-to-end microphone test. Verify capture, recognition, dispatch, playback and the spoken response as one real interaction.
A desktop that can grow around the operator
Lens provides an optional GTK interface with Nostr feed, media and configured wallet integration components. Its supervision and show operation return useful startup errors, allowing it to be evaluated separately from voice. Wallet adapters, mainnet permission and confirmation behaviour need their own setup and tests. No payment or wallet action is part of the suggested voice pilot.
The first useful result is a workstation that performs a small set of commands reliably and makes its failures visible. Measure accidental triggers, missed commands and time to recover a disconnected backend on your own hardware.
Explore the implementation
- Voice coordination and status
- Direct commands, model assistance and ECAI retrieval
- Native voice lifecycle and model preparation
- Model installation and cache checks
- Speech controls and diagnostics
- Native audio, speech and GTK sources
- Voice and desktop test fixtures
The isolated fixtures exercise service behavior with controlled dependencies. Use real audio and your desktop session to assess recognition, playback, speech output and recovery on the machine where erm will run.
