TeensyVAD: Tiny Voice Activity Detectors for Telephony

1VoxLogic — TeensyVAD by VoxLogic · pankaj@voxlogic.ai

Abstract

We present TeensyVAD (TeensyVAD by VoxLogic), a family of voice activity detectors sized 20,449–99,593 parameters — small enough to read end-to-end in one sitting, implemented in pure NumPy with hand-written backpropagation, and native to 8 kHz telephony audio (no resampling anywhere). Across five training generations we study what matters at this scale: (1) labels — distilling frame targets from Silero VAD tightens onset/offset boundaries by 2–3× over synthetic construction labels; (2) data — capacity only pays when data scales with it (families are flat 20k→100k on 1M frames, but improve monotonically on the 37.9M frames of the full 100 h LibriSpeech train-clean-100); (3) quantization — quantization-aware training recovers accuracy over post-training quantization at identical size, yielding a 28 KB int8 model that outperforms its 87 KB float32 parent; and (4) license — the newest generation, teensy-v5, retrains the v4 recipe on MUSAN noise (CC BY 4.0) instead of the NC-licensed share of ESC-50, making the weights commercial-safe with the architecture unchanged. On human-labelled real-world audio — the TEN VAD public set and AMI SDM meetings — v5 is the family's best: the 80k model reaches 0.8877 ROC-AUC on TEN (past v4's 0.880 and FlashVAD v0.1's published 0.882, which uses 2.3× the parameters), and the 40k model posts AMI frame F1 0.8853 vs Silero's stock 0.714 — the best room-audio F1 of any generation at half v4-80k's size (the 1.77M-parameter teacher's default threshold misses 44% of speech in real rooms). Offline batched inference runs at 0.064 µs/frame with ONNX Runtime int8, unchanged from v4 at 63–64 µs per 20 ms streaming chunk — the fastest figure we are aware of among open VADs in this mode, enabled by the stateless-MLP design (recurrent VADs cannot batch through time). Everything — code, weights, evaluation harness, and this site's charts — is reproducible from the linked repositories.

Real-world benchmarks

All systems evaluated on the same 10 ms frame grid with human labels. Operating points calibrated on held-out AMI dev meetings for every detector (WebRTC aggressiveness swept 0–3; Silero and Energy at stock settings). AUC computed on raw probabilities where available. Test sets: the TEN VAD public set (30 recordings) and 8 AMI SDM meetings (distant mic); neither was used in training or calibration.

Real-world comparison chart

Best size per family vs public baselines. Silero ranks best by AUC (87× larger model); teensy-v5 wins the calibrated operating point on AMI and passes FlashVAD's published TEN AUC at a fraction of the size.

modelparamsKB TEN F1*TEN AUC AMI F1AMI AUCµs / 20 ms
teensy-v7-GRU96 — SHIPPING CHAMPION41,809169 0.89920.89340.91530.918237.9
teensy-v7-tt (transformer)120,753486 0.87380.74030.92240.903434.4
teensy-v8-GRU-192×3582k2,285 0.85030.78430.91910.9227138.8
teensy-v8-transformer d128×4570k2,229 0.87260.72550.92340.896984.7
teensy-v8-a3 MLP (float)549k2,157 0.91220.88440.89490.888075.0
teensy-v8-a3 int8 QAT — v8's best549k554 0.91330.89490.88520.8889281.9**
teensy-v6-a2 (250 ms)49,249204 0.90810.88700.88220.872666
teensy-v6-a1 (250 ms)24,44196 0.90810.88340.88420.868364
teensy-v120,44987 0.8770.8480.8870.83564
teensy-v220,44987 0.8900.8680.8800.84863
teensy-v3-80k80,373321 0.8940.8770.8820.86166
teensy-v5 (20k)20,44987 0.89530.87600.88360.857964
teensy-v5-40k39,609162 0.89630.88100.88530.862064
teensy-v5-80k80,373321 0.90160.88770.88450.862263
teensy-v5-100k99,593396 0.90080.88650.88230.859663
teensy-v4 (20k)20,44987 0.8920.8710.8840.86163
teensy-v4-40k39,609161 0.8920.8750.8830.86265
teensy-v4-80k80,373321 0.8960.8800.8800.86266
teensy-v4-100k99,593396 0.8920.8750.8820.86164
teensy-v4-qat (int8)20,44928 0.8940.8760.8840.86292
teensy-v4-40k-qat (int8)39,60947 0.8920.8770.8810.86299
teensy-v4-80k-qat (int8)80,37388 0.8940.8770.8820.863113
teensy-v4-100k-qat (int8)99,593107 0.8910.8760.8810.862120
Silero VAD (teacher)1,774,0002,200 0.9380.9520.7140.89489
WebRTC VAD (agg 0)~6k (C)~50 n/an/a0.8420.7602
Energy VAD (adaptive)—— —0.6700.5920.6587

*TEN F1 at best-F1 threshold (upper bound) — like-for-like with FlashVAD's published TEN numbers (F1 0.889 / AUC 0.882 / FAR 26.3%). AMI F1 at AMI-dev-calibrated thresholds. Speed = full streaming path, median µs per 20 ms telephony chunk, one core of an Apple M2 Pro. **unoptimized NumPy int8 runtime; artifact 4× smaller than float. v8 models are a research release — v7-GRU96 remains the shipping champion.

The v5 generation: commercial-safe, same accuracy

v5 family comparison chart

teensy-v5 is the current release: an identical architecture and training recipe to v4, with one change — training noise from MUSAN (CC BY 4.0) replaces the NC-licensed share of ESC-50, so the weights are fully commercial-safe. Nothing was traded for it: teensy-v5-80k sets a new family record of 0.8877 TEN AUC (v4-80k 0.880, FlashVAD v0.1 0.882), teensy-v5-40k posts the best room-audio F1 of any generation (AMI F1 0.8853) at half v4-80k's size, and speed is unchanged at 63–64 µs per 20 ms streaming chunk. Thresholds were calibrated on AMI dev; all sizes converged on thr_hi 0.10. Weights, float npz ×4, ONNX float ×4 and ONNX int8 ×4 (22 KB int8 for the 20k) are on the teensy-vad-v5 model card (CC BY 4.0 · © 2026 Pankaj Doharey / Metacritical).

New: teensy-v7 — the tiny recurrent generation

teensy-v7 vs baselines

teensy-v7 is the current release — and the family's first recurrent model. The v6 capacity ablation showed more parameters don't help at 100 ms of context; v7's answer is that memory was the missing ingredient: a single-layer GRU with 96 units (41,809 parameters, 169 KB npz), distilled from Silero, replaces the stateless MLP. It beats the entire v5/v6 MLP family on all four real-world metrics — TEN F1 0.8992, TEN AUC 0.8934, AMI F1 0.9153, AMI AUC 0.9182 — and at 37.9 µs per 20 ms chunk it is the fastest model in family history (vs 63–75 µs for the MLPs). It also beats Silero on both AMI room metrics (F1 0.9153 vs 0.7136, AUC 0.9182 vs 0.8938); Silero keeps the clean near-mic crown (TEN AUC 0.9519). Weights are CC BY 4.0 on MUSAN-safe data.

New: teensy-v8 — the 500k scaling ablation

teensy-v8 vs baselines

teensy-v8 is a research release — four ~500k models (3-layer GRU-192, causal transformer d128×4, context-MLP a3 float and int8 QAT) trained on 660 h of prior-balanced LibriSpeech+MUSAN mixtures with Silero distillation and QAT fine-tuning, to test whether scale beats the 42k v7 champion. The answer is negative but informative: no v8 model dominates v7-GRU96 — deep nets win AMI rooms (F1 0.919–0.923) but collapse on clean near-mic ranking (TEN AUC 0.73–0.78 vs 0.89); the context-MLP scaled best, and QAT improved its TEN ranking (0.8844→0.8949) at family-best TEN F1 0.9133 with a 4× smaller int8 artifact. v7-GRU96 remains the shipping champion; the honest caveat is that only ~1.8 training passes over 660 h leaves under-training as a confound.

New: teensy-v6 — context beats capacity

teensy-v6 vs baselines

teensy-v6 widens temporal context from 100 ms to 250 ms instead of adding parameters — and every v6 point at ≤49k params sits above the entire measured v5 100 ms capacity curve (which peaks at 80k, TEN AUC 0.8877, and falls past it: 0.8865 → 0.8827 → 0.8849 at 100k–200k, latency 64→75 µs). v6-a2 (49,249 params) matches the v5-80k peak on TEN AUC (0.8870) with 60% fewer parameters and the family's best AMI AUC (0.8726); v6-a1 (24k) sets a family-record TEN F1 0.9081. Same MUSAN-safe CC BY 4.0 data as v5 — artifacts on the teensy-vad-v6 model card.

What scales: data, not just parameters

Capacity scaling chart

The v1/v2 families (1M training frames) are flat from 20k to 100k parameters — larger nets only overfit. The v3 family (10.7M frames) finds a sweet spot at ~80k. The v4 family, trained on the full ~100 h of LibriSpeech train-clean-100 (28,539 utterances, 37.9M teacher-labelled frames, memmap/float16 plumbing for a 60 GB-scale design matrix), keeps improving to the largest size tested:

family (frames)20k40k80k100k
v1 (1M, construction labels)0.9110.9090.9100.910
v2 (1M, distilled)0.9100.9120.9120.911
v3 (10.7M, distilled)0.9110.9160.9170.916
v4 (37.9M = 100 h, distilled) 0.9140.9180.9200.921

(validation frame F1, Silero-teacher labels)

Quantization: QAT makes int8 free

v4 family comparison chart

Post-training int8 quantization is essentially free on these models (ΔAUC ≈ 0.000). Quantization-aware training (fake-quant forward pass with straight-through-estimator gradients; a unit test pins simulation ≡ deployed inference) holds that line across the whole size range — int8 matches float32 within ±0.003 real-world AUC at every size, in files 2.7–3.7× smaller. Size-by-size (TEN AUC, float → int8): 20k 0.871→0.876 (int8 beats float), 40k 0.875→0.877, 80k 0.880→0.877, 100k 0.875→0.876. Two standouts: teensy-v4-80k-qat — 88 KB, TEN AUC 0.877, the family's best AMI AUC (0.863), and teensy-v4-qat — 28 KB, TEN AUC 0.876, beats its 87 KB float32 parent. Notably, int8-vs-float is not monotonic in size: quantization sensitivity depends on the learned weight distribution, not parameter count. In pure NumPy, int8 is a size play (no int8 BLAS kernels); real int8 speed lives in ONNX Runtime.

Speed, stated honestly

Speed vs accuracy chart
scenarioteensyvadFlashVAD (published)WebRTC
model inference, single frame 6.5 µs (ONNX f32)11.4 µs (native Accelerate)2 µs (C ext)
full streaming path, as shipped 63–66 µs / 20 ms (NumPy) ~23 µs / 20 ms (compiled)2 µs
offline, batched (20k frames) 0.064 µs/frame (ONNX int8) n/a — recurrent, cannot batchn/a

We make one speed claim, precisely: the fastest batched (offline) inference we are aware of among open VADs — 0.064 µs/frame, roughly 15,000× faster than real time. This is architectural, not magic: a stateless MLP over a fixed 10-frame window treats an entire recording as one matrix multiply, while recurrent VADs (Silero, FlashVAD, TEN) must run hop-by-hop through time. In streaming mode, WebRTC's C extension remains the floor (2 µs) and FlashVAD's compiled frontend beats our NumPy one — though our ONNX model alone is faster per single frame than FlashVAD's native figure. At telephony budgets every system here uses < 0.6% of a core; accuracy-per-KB, not speed, is the real axis.

Method in one paragraph

Audio is mono 8 kHz (the native rate of PSTN, G.711 and Asterisk's slin — no upsampling to a 16 kHz core). Each 25 ms frame every 10 ms is Hann-windowed, FFT'd to 256 points, summed through 20 mel triangles over 80–3800 Hz, log-compressed, and band-mean-subtracted — making features invariant to line gain, so the model sees spectral shape only. First-order deltas concatenate to 40 dims/frame; 10 frames (100 ms) flatten to the 400-dim input of a 3-layer MLP (48/24 hidden, 20,449 params) with sigmoid output. Training targets come from Silero VAD run over synthetically mixed training audio (LibriSpeech speech + ESC-50 noise + synthetic 7-talker babble + real AMI room ambience, −5–20 dB SNR, 40% G.711 µ-law round-trip). Decisions run through hysteresis + 250 ms hangover to emit speech_start / speech_end events. The full feature spec is reimplementable from the v1 model card.

Limitations

English read speech dominates training; music was absent (reads as "activity"); 100 ms context is short for unvoiced fricatives in noise. Silero — the 87× larger teacher — still ranks best by AUC on both real sets; our wins are the calibrated operating point on real rooms, accuracy-per-KB, and batched throughput. TEN-set F1 values marked * are threshold-tuned on that set and are upper bounds (FlashVAD's published numbers share this caveat). Speed figures are from an Apple M2 Pro; Celeron-class CPU estimates are engineering extrapolations, not measurements.

Citation

@software{teensyvad2026,
  title  = {TeensyVAD: Tiny Voice Activity Detectors for Telephony},
  author = {Doharey, Pankaj (VoxLogic)},
  year   = {2026},
  url    = {https://huggingface.co/Teensy},
  note   = {TeensyVAD by VoxLogic — families v1--v5; 20k--100k parameters; 8 kHz native; contact: pankaj@voxlogic.ai}
}