Personal VAD (1.0 & 2.0): Speaker-Conditioned Voice Activity Detection

Interactive Frame-Level Target-Speaker Voice Activity Detection & Audio Gating Demo

This is not an officially supported Google product. This interactive demo and the accompanying open-source library reproduce the Personal VAD 1.0 and 2.0 architectures from scratch using public 8-language Multilingual LibriSpeech (MLS) audio (`de`, `en`, `es`, `fr`, `it`, `nl`, `pl`, `pt`). Instead of the original papers' proprietary 4.88M-parameter 3-layer GE2E LSTM speaker encoder, enrollment 256-D d-vectors are extracted using our self-contained open-set regularized LDA+PCA speaker subspace extractor (`speaker_subspace.npz`, trained from scratch on 98 MLS training speakers; `0.9521` ROC-AUC on 35 unseen test speakers).

1. Input & Conditioning Controls

Enrollment-Less Mode (e_target = 0, PVAD 2.0 Unified E5)
Model Target Speech Non-Target Silence

2. Side-by-Side Frame Posterior Comparison

1. Multi-Speaker Waveform & Target Speaker Gate Mask
Waveform Target Gate Active
2. Standard VAD Baseline (B2 Conformer) — Triggers on All Speakers
Speech (All) Silence (ns)
3. Personal VAD 1.0 (ET + Weighted Pairwise Loss, 2-Layer LSTM, 130K params)
Target Speech (tss) Non-Target (ntss) Silence (ns)
4. Personal VAD 2.0 (E3/E5 Conformer + FiLM + Speaker Pre-Net)
Target Speech (tss) Non-Target (ntss) Silence (ns)