erm: Bring Voice, Media and Operational Knowledge to Your Desktop

Keep a small command from becoming another context switch

Pause the music. Find a track. Ask a question about an indexed document. These are modest interactions, but they repeatedly pull attention away from the task on screen. erm brings them into a supervised Erlang desktop workspace with voice input, media control, optional local model assistance and speech output.

The clearest first audience is a Linux developer or operator willing to run the local audio and GTK dependencies. Start with a few observable commands on your own workstation. That is enough to evaluate whether the interaction helps without treating the whole desktop as an autonomous agent.

Understand the parts of a voice interaction

A useful voice interface has several jobs. It captures speech, decides whether the user is addressing it, turns the request into an allowed action, performs that action and reports the outcome. Separating those jobs makes a failure easier to locate.

Component Role in the interaction
Capture and Whisper integration Turn microphone audio into text
Voice coordinator and intent parser Identify a command and manage its progress
MPV adapter and playlist Apply media controls to the player
Optional Ollama model Help interpret unfamiliar phrases or answer a question
Piper speech service Speak the result using a configured local voice
GTK/Lens interface Provide a visible desktop interface alongside voice

For a workstation user, the first useful test is simple: say a command and observe the correct effect. A transcript alone is only one stage of that interaction. The table also explains why a working recogniser can coexist with a missing playback or speech-output dependency.

Common commands have a direct path

The voice-intent module recognises play, pause, stop, next, previous, volume, track selection and player visibility before asking a model to classify an unknown phrase. Volume input is bounded. Known negated commands are refused. The media adapter sends typed actions to the existing MPV process owner and uses the playlist state for track selection and progression.

The voice coordinator handles transcript boundaries, command windows, deduplication and asynchronous work. Model planning and media actions have timeouts; they do not run as blocking work inside the coordinator's message handler. The operating benefit is that command state can be inspected and failures can be reported through the same service interface.

Try a small command set such as β€œBob, pause music”, β€œBob, next song” and β€œBob, volume 30” after configuring the trigger. These are pilot utterances, not measured recognition results. Check actual playback and volume after each command, including repetitions and silence around the wake phrase.

Model assistance has a defined action vocabulary

For phrases outside the direct parser, the code requests a bounded JSON intent and validates it against built-in or operator-configured actions. Model output does not choose arbitrary Erlang module/function pairs. Custom callbacks come from trusted configuration.

That is a useful extension boundary for a workstation tool. A maintainer can add a named action with an explicit implementation and test its effect. The callback still carries whatever authority its implementation has; a voice phrase or speaker match should not be treated as approval for payments, secret access or other sensitive operations.

The configured Ollama path can answer a general question or assist with an intent. Local processing depends on the actual configuration and installed models. The code also supports pulling a missing configured model, so setup and downloads are separate from an offline-use claim.

Ask indexed knowledge, with the retrieval boundary visible

The β€œask ECAI” route retrieves up to three sources from the configured disk corpus and sends bounded excerpts to the answer model. Missing corpus configuration and empty retrieval have explicit error paths.

This is a concrete connection between erm and ECAI, but it uses ecai_ollama_rag:retrieve_sources/3. It does not call the corpus/principal private bridge. Use an appropriate public or non-sensitive corpus for this pilot. Private voice access needs an explicit authenticated integration with the private retrieval boundary.

Speech and service health are part of the interface

The TTS service exposes native Piper speech, voice selection, volume controls, repeat, cancellation and diagnostics. Model installation is handled by an asynchronous service with manifest-driven file checks and immutable cache directories. Listening suppression is used around speech output to reduce feedback into the recogniser.

The native voice coordinator also handles enrolment, speaker scores and model preparation. In the referenced implementation it can request missing configured models through erm_model_pull, report progress and retry failed preparation. The model service checks file sizes and SHA-256 digests before using its cache. Initial downloads require connectivity.

Native capture and speech sources live in the repository's apps/erm/c_src/ directory. They still need a compatible local build, installed models and audio devices. No microphone or playback test was run for this article. Evaluate speaker matching with recordings, other speakers and replay attempts before assigning it a security role; a speaker score alone does not establish that a live authorised person is issuing the command.

In a running configured release, these inspection calls help separate a command problem from a missing model or disconnected speech backend:

erm_voice:status().
erm_voice:healthcheck().
erm_tts:diagnostics().

Readiness is not an end-to-end microphone test. Verify capture, recognition, dispatch, playback and the spoken response as one real interaction.

A desktop that can grow around the operator

Lens provides an optional GTK interface with Nostr feed, media and configured wallet integration components. Its supervision and show operation return useful startup errors, allowing it to be evaluated separately from voice. Wallet adapters, mainnet permission and confirmation behaviour need their own setup and tests. No payment or wallet action is part of the suggested voice pilot.

The first useful result is a workstation that performs a small set of commands reliably and makes its failures visible. Measure accidental triggers, missed commands and time to recover a disconnected backend on your own hardware.

Explore the implementation