This is not an officially supported Google product.
This interactive demo and the accompanying open-source library reproduce the Personal VAD 1.0 and 2.0 architectures from scratch using public 8-language Multilingual LibriSpeech (MLS) audio (`de`, `en`, `es`, `fr`, `it`, `nl`, `pl`, `pt`). Instead of the original papers' proprietary 4.88M-parameter 3-layer GE2E LSTM speaker encoder, enrollment 256-D d-vectors are extracted using our self-contained open-set regularized LDA+PCA speaker subspace extractor (`speaker_subspace.npz`, trained from scratch on 98 MLS training speakers; `0.9521` ROC-AUC on 35 unseen test speakers).
2. Side-by-Side Frame Posterior Comparison
1. Multi-Speaker Waveform & Target Speaker Gate Mask
Waveform
Target Gate Active
2. Standard VAD Baseline (B2 Conformer) — Triggers on All Speakers
Speech (All)
Silence (ns)
3. Personal VAD 1.0 (ET + Weighted Pairwise Loss, 2-Layer LSTM, 130K params)
Target Speech (tss)
Non-Target (ntss)
Silence (ns)
4. Personal VAD 2.0 (E3/E5 Conformer + FiLM + Speaker Pre-Net)
Target Speech (tss)
Non-Target (ntss)
Silence (ns)