WakeHuBERT wake words
Ready wake-word models for OpenVoiceOS, trained with wakeforge. Each model is a small GRU classifier that reads features from WakeHuBERT tiny, a 0.64M-parameter speech feature extractor distilled from HuBERT-base. A model scores the last 1.5 s of audio (75 feature frames) and outputs a logit; its sigmoid is the probability that the window holds the wake word.
Try them in your browser: WakeHuBERT wake-words Space runs every model here on your microphone or an uploaded file, with a live score and the calibrated threshold. Everything runs locally in the browser.
The repository holds one ONNX file per word under models/, and models.json, which lists each model's word,
featurizer, calibration, default threshold, SHA-256 and measured results. Every model carries the same facts in
its own ONNX metadata (wake_word, pretrained_featurizer, default_threshold, window_frames, license,
training_data, and on calibrated models calibrated, calib_a and calib_b).
| model | word | featurizer | calibrated | default threshold |
|---|---|---|---|---|
wakehubert_jarvis |
jarvis | wakehubert-int8 | yes | 0.57 |
wakehubert_alexa |
alexa | wakehubert-int8 | yes | 0.44 |
wakehubert_hey_jarvis |
hey jarvis | wakehubert-int8 | yes | 0.16 |
wakehubert_hey_marvin |
hey marvin | wakehubert-int8 | yes | 0.34 |
wakehubert_home_assistant |
home assistant | wakehubert-int8 | yes | 0.19 |
wakehubert_okay_nabu |
okay nabu | wakehubert-int8 | yes | 0.45 |
wakehubert_hello_nabu |
hello nabu | wakehubert-int8 | yes | 0.49 |
wakehubert_hey_chatterbox |
hey chatterbox | wakehubert-int8 | yes | 0.13 |
wakehubert_hey_floyd |
hey floyd | wakehubert-int8 | yes | 0.43 |
wakehubert_hey_rhasspy |
hey rhasspy | wakehubert-int8 | yes | 0.40 |
wakehubert_hey_robin |
hey robin | wakehubert-int8 | yes | 0.05 |
wakehubert_marvin |
marvin | wakehubert-int8 | yes | 0.06 |
wakehubert_sheila |
sheila | wakehubert-int8 | yes | 0.47 |
wakehubert_stop |
stop | wakehubert-int8 | yes | 0.14 |
wakehubert_android |
android | wakehubert-int8 | yes | 0.42 |
wakehubert_hey_computer |
hey computer | wakehubert-int8 | yes | 0.31 |
wakehubert_hey_k9 |
hey k9 | wakehubert-int8 | yes | 0.06 |
wakehubert_hey_scout |
hey scout | wakehubert-int8 | yes | 0.36 |
wakehubert_wake_up |
wake up | wakehubert-int8 | yes | 0.36 |
wakehubert_hey_ziggy |
hey ziggy | wakehubert-int8 | yes | 0.27 |
wakehubert_hey_potato |
hey potato | wakehubert-int8 | yes | 0.15 |
wakehubert_hey_stemcom |
hey stemcom | wakehubert-int8 | yes | 0.28 |
wakehubert_computer |
computer | wakehubert (float32) | no | 0.99 |
wakehubert_hey_mycroft |
hey mycroft | wakehubert-int8 | yes | 0.18 |
wakehubert_despierta |
despierta (Spanish) | wakehubert-int8 | yes | 0.05 |
wakehubert_aufwachen |
aufwachen (German) | wakehubert-int8 | yes | 0.43 |
wakehubert_wakker_worden |
wakker worden (Dutch) | wakehubert-int8 | yes | 0.16 |
wakehubert_sveglia |
sveglia (Italian) | wakehubert-int8 | yes | 0.08 |
Use with OpenVoiceOS
The models run in ovos-ww-plugin-wakeforge, which
downloads them from this repository on first use, at a pinned revision, checks each file against its sha256 in
models.json, and keeps them in the Hugging Face cache; the featurizer comes from
TigreGotico/wakehubert-tiny the same way. Install the plugin
and name the model in mycroft.conf:
pip install --pre ovos-ww-plugin-wakeforge
{
"listener": {
"wake_word": "hey_jarvis"
},
"hotwords": {
"hey_jarvis": {
"module": "ovos-ww-plugin-wakeforge",
"model": "wakehubert_hey_jarvis",
"listen": true
}
}
}
A model file from this repository also loads by path: set model to the local .onnx file. The plugin reads
the featurizer and the default threshold from the model's metadata.
Use in your own code
The models need only numpy, onnxruntime and huggingface_hub; nothing from OpenVoiceOS. Each model is a small classifier that reads WakeHuBERT-tiny features, so you run two ONNX files: the featurizer, then the wake-word model.
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
word = "jarvis"
head = ort.InferenceSession(hf_hub_download("OpenVoiceOS/wakehubert-wakewords", f"models/wakehubert_{word}.onnx"))
meta = head.get_modelmeta().custom_metadata_map
featurizer_file = "wakehubert_int8.onnx" if meta["pretrained_featurizer"] == "wakehubert-int8" else "wakehubert.onnx"
feat = ort.InferenceSession(hf_hub_download("TigreGotico/wakehubert-tiny", featurizer_file))
threshold = float(meta["default_threshold"])
WINDOW = 24000 # 1.5 s of 16 kHz audio: the window the models were trained on (75 feature frames)
BLOCK = 1280 # score every 80 ms
DEBOUNCE_BLOCKS = 25 # ignore 2 s after a detection
class WakeWordDetector:
def __init__(self):
self.buf = np.zeros(WINDOW, np.float32)
self.cooldown = 0
def push(self, chunk):
"""chunk: 1280 float32 samples at 16 kHz in -1..1. Returns (score, detected)."""
self.buf = np.concatenate([self.buf, chunk.astype(np.float32)])[-WINDOW:]
features = feat.run(None, {"waveform": self.buf[None]})[0] # [1, 75, 128]
logit = head.run(None, {"features": features})[0][0]
score = float(1.0 / (1.0 + np.exp(-logit)))
self.cooldown = max(0, self.cooldown - 1)
detected = score >= threshold and self.cooldown == 0
if detected:
self.cooldown = DEBOUNCE_BLOCKS
return score, detected
From a microphone, for example with sounddevice:
import sounddevice as sd
detector = WakeWordDetector()
with sd.InputStream(samplerate=16000, channels=1, dtype="float32", blocksize=BLOCK) as stream:
while True:
chunk, _ = stream.read(BLOCK)
score, detected = detector.push(chunk[:, 0])
if detected:
print(f"{word} detected (score {score:.2f})")
Each block featurizes the last 1.5 s on its own, exactly as the models were trained and scored. The featurizer is causal, the per-block cost is about 1–2 ms on one CPU core for the int8 featurizer, and one featurizer run can feed any number of wake-word models: run feat once per block and pass the same features to each model. Replace threshold to change sensitivity (see below). Checked against the plugin on real recordings: the snippet fires on the same clips.
Choosing a threshold
Set "threshold" in the hotword config to trade missed wake words against false activations. A lower value
fires more readily and falsely more often; a higher value misses more wake words and fires falsely less often.
On a calibrated model the number means the same thing for every word. Its score is mapped so that a threshold of 0.5 gives about one false activation per hour on held-out speech and noise. The default threshold is the point that maximises F2, which weighs recall above precision. Raising the threshold toward 0.8 or 0.9 trades recall for fewer false activations, and lowering it does the opposite.
The uncalibrated model (wakehubert_computer) outputs a probability too, but its score is not mapped to a
false-activation rate. Its threshold is not a calibrated knob, and useful values sit close to 1.
Results
Measured through the plugin at the default threshold and at 0.8. Recall is the share of test clips detected; false activations are counted per hour of negative audio.
| model | recall at default | false activations/h at default | recall at 0.8 | false activations/h at 0.8 | recall test set |
|---|---|---|---|---|---|
wakehubert_jarvis |
98.4% | 0.69 | 93.8% | 0.15 | 384 Picovoice recordings of real speakers |
wakehubert_alexa |
98.4% | 0.60 | 91.4% | 0.28 | 315 Picovoice recordings of real speakers |
wakehubert_hey_jarvis |
94.5% | 0.09 | 86.2% | 0.02 | 384 clips in held-out synthetic voices |
wakehubert_hey_marvin |
97.4% | 0.95 | 90.7% | 0.24 | 386 clips in held-out synthetic voices |
wakehubert_home_assistant |
91.1% | 0.30 | 84.2% | 0.15 | 380 clips in held-out synthetic voices |
wakehubert_okay_nabu |
92.0% | 0.26 | 75.4% | 0.02 | 386 clips in held-out synthetic voices |
wakehubert_hello_nabu |
80.9% | 0.39 | 65.2% | 0.02 | 382 clips in held-out synthetic voices |
wakehubert_hey_chatterbox |
82.8% | 0.09 | 61.2% | 0.00 | 116 OVOS community recordings of real speakers |
wakehubert_hey_floyd |
91.7% | 0.47 | 82.3% | 0.02 | 96 OVOS community recordings of real speakers |
wakehubert_hey_rhasspy |
100.0% | 0.60 | 97.6% | 0.11 | 374 clips in held-out synthetic voices |
wakehubert_hey_robin |
99.5% | 0.39 | 97.9% | 0.06 | 380 clips in held-out synthetic voices |
wakehubert_marvin |
73.3% | 0.77 | 71.8% | 0.67 | 195 Speech Commands test recordings of real speakers |
wakehubert_sheila |
88.7% | 2.08 | 84.4% | 0.54 | 212 Speech Commands test recordings of real speakers |
wakehubert_stop |
86.6% | 2.06 | 76.9% | 0.45 | 411 Speech Commands test recordings of real speakers |
wakehubert_android |
98.7% | 0.95 | 96.4% | 0.11 | 390 clips in held-out synthetic voices |
wakehubert_hey_computer |
96.4% | 0.47 | 93.3% | 0.04 | 390 clips in held-out synthetic voices |
wakehubert_hey_k9 |
99.4% | 0.19 | 93.5% | 0.04 | 338 clips in held-out synthetic voices |
wakehubert_hey_scout |
95.9% | 0.09 | 93.8% | 0.04 | 390 clips in held-out synthetic voices |
wakehubert_wake_up |
98.1% | 1.10 | 96.8% | 0.19 | 378 clips in held-out synthetic voices |
wakehubert_hey_ziggy |
93.6% | 0.64 | 89.3% | 0.24 | 374 clips of OmniVoice and held-out edge-tts voices |
wakehubert_hey_potato |
88.9% | 0.34 | 80.6% | 0.02 | 360 clips in held-out edge-tts voices only, without OmniVoice test clips |
wakehubert_hey_stemcom |
89.0% | 0.21 | 83.1% | 0.06 | 337 clips of OmniVoice and held-out edge-tts voices |
wakehubert_hey_mycroft |
97.4% | 1.57 | 84.8% | 0.24 | 285 clips in held-out edge-tts voices and 379 Kokoro clips converted to unseen speakers |
wakehubert_despierta |
95.4% | 0.39 | 92.4% | 0.16 | 370 Spanish OmniVoice clips with unseen seeds |
wakehubert_aufwachen |
99.2% | 0.16 | 94.6% | 0.03 | 390 German OmniVoice clips with unseen seeds |
wakehubert_wakker_worden |
98.2% | 0.58 | 94.4% | 0.06 | 390 Dutch OmniVoice clips with unseen seeds |
wakehubert_sveglia |
98.7% | 0.19 | 96.0% | 0.03 | 222 Italian OmniVoice clips with unseen seeds |
The false activations are counted over 46.5 h of negative audio (speech, non-speech and household audio). The held-out synthetic voices are text-to-speech voices that no training clip uses, converted to the voices of speakers who appear in no training clip. For the localised wake-up models (the words with a language in brackets in the table above), the false activations are counted over the 31.1 h of speech and non-speech audio without the household audio, and recall is measured on clips of OmniVoice voices drawn from seeds that no training clip uses. No figures are published for the uncalibrated model.
Training data
The calibrated models were trained on synthetic speech only, with six edge-tts voices held out of training for
testing. The positives are an edge-tts voice grid, further edge-tts and Google Translate TTS voices and OmniVoice
clips, and for wakehubert_jarvis, wakehubert_hey_chatterbox, wakehubert_hey_floyd, wakehubert_marvin,
wakehubert_sheila and wakehubert_stop also voice-converted copies of edge-tts clips. wakehubert_alexa trained instead on the
multi-engine, edge-tts and Piper (LibriTTS-R speakers) clips of its dataset and on OmniVoice clips. Each model's training_data metadata names its own sources. Most of these clips are published in the TigreGotico/synthetic-wakeword-<word>
datasets listed above. Speech from LibriSpeech train-clean-100 is mixed into training clips as background babble.
The negatives are the wakeforge negative list and half of an AudioSet noise sample; the other half is held out.
The localised wake-up models were trained on synthetic speech only: OmniVoice clips with no reference speaker, one seed per clip, kept when a speech recogniser transcript matched the phrase (or, for languages no recogniser covers, when duration, level and speech-activity checks passed), together with the clips of the phrase's TigreGotico/synthetic-wakeword-<word> dataset, which also holds the kept OmniVoice clips and their held-out test split. Speech from LibriSpeech train-clean-100 is mixed in as background babble, and the negatives are the same as above.
wakehubert_computer was trained on synthetic speech only, from TigreGotico/synthetic-wakeword-computer, with
negatives from TigreGotico/not-wake-words-speech-en and AudioSet-derived clips.
Calibration
A calibrated model has an affine map folded into its graph, applied to the GRU's logit before the sigmoid. The map
is fitted on a stream of held-out noise and speech that the model did not train on, so that a probability of 0.5
falls at about one false activation per hour of that stream. The default threshold is then the F2-optimal point on
the calibrated scale. The fitted slope and offset are in each model's calib_a and calib_b metadata.
Limitations
- Twenty of the twenty-seven calibrated models are scored on held-out synthetic voices, so recall on real speech can
be lower for those words.
wakehubert_jarvisandwakehubert_alexaare scored on the Picovoice recordings,wakehubert_hey_chatterboxandwakehubert_hey_floydon OVOS community recordings, andwakehubert_marvin,wakehubert_sheilaandwakehubert_stopon Speech Commands, all real speakers. wakehubert_sheilaandwakehubert_stopgive about two false activations per hour at their default threshold; raise it toward 0.8 for about one every two hours.wakehubert_marvinhas a steep calibration, so its recall and false-activation rate change little between its 0.06 default and 0.8.- The calibration is fitted on about 10 h of audio with few false activations in it (2 to 14 per model), so the map is extrapolated, and "0.5 is about one false activation per hour" is approximate.
- On real jarvis recordings,
wakehubert_jarvispeaks just above its 0.57 default, so a quiet or distant speaker has little margin. - Similar-sounding words trigger each other's model.
wakehubert_jarvisfires on "hey jarvis",wakehubert_marvinon "hey marvin", andwakehubert_hey_jarvison "hey chatterbox".wakehubert_hello_nabu,wakehubert_hey_marvinandwakehubert_okay_nabucan fire on each other's words,wakehubert_hey_marvinalso on "hey rhasspy", "hey robin" and "marvin",wakehubert_okay_nabuon "hey rhasspy",wakehubert_hey_robinon "hey marvin", "hey rhasspy" and "okay nabu",wakehubert_okay_nabusometimes on "hey k9",wakehubert_androidsometimes on "hey floyd",wakehubert_computerandwakehubert_hey_computeron each other's words, andwakehubert_sheilaon "computer", all at their default thresholds. Raise the threshold when two of these models run side by side. - The localised wake-up models also fire on inflections of the same stem:
wakehubert_despiertaon "despierto", "despiertas" and "depierta",wakehubert_svegliaon "sveglio" and "sveglie".wakehubert_svegliaalso fires on rhymes such as "meraviglia" and "bottiglia", andwakehubert_aufwachenon "aufmachen", all at their default thresholds. Raising the threshold toward 0.8 removes most of these. - The models are English except the localised wake-up models, whose language is in brackets in the table above.
Their negative audio is English speech and non-speech audio; on read speech in their own language (the FLEURS
development set) at the default threshold they gave:
wakehubert_despierta0.00 per hour over 1.4 h;wakehubert_aufwachen0.79 per hour over 1.3 h;wakehubert_wakker_worden0.00 per hour over 0.5 h;wakehubert_sveglia1.30 per hour over 1.5 h.
License
Apache-2.0. The featurizer, TigreGotico/wakehubert-tiny, is
Apache-2.0. The synthetic-wakeword-* and not-wake-words-speech-en datasets are CC BY 4.0; LibriSpeech is
CC BY 4.0; AudioSet labels are CC BY 4.0 and its audio comes from YouTube videos under their uploaders' terms. The
voice-conversion and voice-cloning folders of the synthetic-wakeword-* datasets take their voices from Mozilla
Common Voice contributors, through the MLCommons Multilingual Spoken Words Corpus (CC BY 4.0).
Model tree for OpenVoiceOS/wakehubert-wakewords
Base model
facebook/hubert-base-ls960