NeurIPS 2026 Submission · Project Page

Efficient Streaming Audio-Visual Target Speaker Extraction for Real-World Acoustic Scenes

Anonymous Authors
Double-blind submission page · Audio-Visual Target Speaker Extraction · Streaming Inference
7.7 dBRealSSA SI-SNRi, Falcon
16.4 dBLRS3-2Mix SI-SNRi, Falcon
0.5MParameters, Falcon
51.3 msGPU latency, Falcon

Abstract

Despite rapid progress in audio-visual target speaker extraction (AVTSE), deploying it as a streaming system remains hard for two reasons: high-performance causal models incur large parameter counts and long inference latency, whereas lightweight alternatives recover capacity by repeatedly applying the same separator and thereby reintroduce computation.

We propose Falcon, an efficient causal AVTSE framework whose AudioBlock performs multi-scale time-frequency separation in a single forward pass and is trained with an encoder-space speaker contrastive loss that suppresses near-field leakage at no inference cost and is architecture-agnostic across AVTSE backbones.

We also release RealSSA, a realistic simulated-scene AVTSE benchmark with 2 to 6 dynamic speakers across four scene types, paired with a real-recorded evaluation set. Falcon achieves state-of-the-art extraction on RealSSA, LRS2-2Mix, and LRS3-2Mix while reducing parameters, computational cost, and GPU inference latency over the AV-TFGridNet baseline by 91.9%, 20.2×, and 11.1×, respectively.

Streaming AVTSE Single-pass Multi-scale Separator RealSSA Benchmark Real-world Acoustic Scenes

Key Contributions

The page is designed to make the workload visible: a new streaming model, a new realistic benchmark, and comprehensive evaluation across simulated and recorded scenes.

Single-pass streaming separation

Falcon replaces iterative separator cycles with an encoder-decoder AudioBlock that captures multi-scale time-frequency structure in one forward pass.

Speaker contrastive training

An encoder-space contrastive loss pulls estimates toward the clean target while pushing them away from near-field interferers, without adding inference cost.

RealSSA benchmark

RealSSA covers dynamic speaker counts, near-/far-field interference, reverberation, environmental noise, sudden noise, and real-recorded evaluation.

RealSSA Dataset

RealSSA is built around dynamic speaker counts, explicit distance control, and systematic acoustic variation across realistic scenes.

Scene Coverage

Four scene types simulate real deployment settings: small meeting rooms, medium indoor spaces, large public spaces, and outdoor environments.

Acoustic Factors

RealSSA jointly varies near-field speakers, far-field speech, RT60, environmental noise SIR, and sudden noise events.

SceneSpace SizeRT60Near SpeakersFar SpeakersEnv. Noise SIRSudden Noise
Small meeting4–6 × 3–5 × 2.5–3 m0.3–0.6 s1–5, 0.5–1.5 mnone5–15 dB30%, 1 event
Medium indoor8–15 × 8–12 × 3–4 m0.5–0.9 s1–3, 1.0–2.5 m10–15, 3–15 m5–10 dB40%, 1–2 events
Large public20–40 × 15–30 × 4–8 m1.0–2.0 s1–4, 1.0–3.0 m10–15, 5–20 m2–8 dB50%, 1–3 events
Outdoor50 × 50 × 20 m0.05–0.1 s1–4, 1.0–4.0 m10–15, 5–25 m-2–5 dB40%, 1–3 events

Methodology

Falcon receives mixture audio and target-speaker mouth video, extracts audio-visual representations, and outputs the target speech stream under causal constraints.

Overall Pipeline

Falcon adopts a standard target-speaker extraction framework: audio and visual encoders extract modality-specific features, a fusion module combines them, and a single-pass separator with mask generator recovers the target speech. Unlike iterative separators that recycle one compact block across multiple cycles, Falcon accumulates multi-scale time-frequency capacity within a single forward pass.

Overall pipeline of Falcon
Figure 1. Overall pipeline of Falcon for streaming audio-visual target speaker extraction.
Mixture AudioStreaming waveform input under dynamic interference.
Mouth VideoVisual target-speaker cue for permutation-free extraction.
Falcon AudioBlockSingle-pass multi-scale time-frequency modeling.
Target SpeechLow-latency extracted speech for downstream ASR.

Single-Pass Multi-Scale Separator

The separator follows an encoder-decoder design. The encoder progressively downsamples the fused feature across 3 stages to extract multi-scale representations and computes a global context at the bottleneck. This context is injected into skip features at each scale via gated modulation. The decoder restores resolution from coarse to fine, applying time-frequency modeling at every level.

Single-pass multi-scale separator
Figure 2. Single-pass multi-scale separator with frequency-wise, temporal, and joint time-frequency refinement.

Time-Frequency Refinement Block (TFRB)

Each TFRB applies three stages of modeling in sequence: (1) a bidirectional LSTM along the frequency axis for spectral pattern capture, (2) a unidirectional LSTM along the time axis to maintain causal streaming, and (3) masked 2D self-attention over the joint time-frequency plane. This design enables fine-grained spectro-temporal refinement at each decoder level while preserving the causal constraint.

Time-Frequency Refinement Block
Figure 3. Time-Frequency Refinement Block (TFRB). Each block applies frequency-wise BiLSTM, causal temporal LSTM, and masked 2D self-attention in sequence.

Audio-Visual Demo Samples

Select a scene and switch among different systems. The video timeline controls the selected separated audio, enabling synchronized audio-visual comparison.

Small Meeting Room Close-range target speaker extraction with near-field interference.
Choose a system
Now Playing: Base
Synchronized with video

Play or seek the video. The selected separated audio will follow the same timestamp.

Medium Indoor Indoor competing speakers and reverberation.
Choose a system
Now Playing: Base
Synchronized with video

The visible video is fixed, while the synchronized separated audio can be switched among systems.

Large Public Space High reverberation and dynamic interference.
Choose a system
Now Playing: Base
Synchronized with video

This sample highlights robustness under high reverberation and sudden interference events.

Outdoor Outdoor background noise and moving interference.
Choose a system
Now Playing: Base
Synchronized with video

This sample demonstrates streaming AVTSE under outdoor noise with moving interference.

Note. The original video audio is muted. The sound you hear comes from the selected separated audio track and is synchronized with the video timeline.

Experimental Results

Main separation quality and efficiency comparisons. Bold indicates the best result; underlined values indicate the second-best result.

MethodRealSSA-RealRealSSA-SimLRS2-2MixLRS3-2MixVoxCeleb2-2Mix
DNSMOS ↑SI-SNRi ↑SDRi ↑SI-SNRi ↑SDRi ↑SI-SNRi ↑SDRi ↑SI-SNRi ↑SDRi ↑
AV-ConvTasNet2.5801.22.69.810.312.512.86.98.2
AV-Sepformer2.8135.56.310.611.113.413.76.98.3
AV-TF-GridNet3.0307.68.112.713.114.314.85.87.7
RTFSNet-123.0664.95.712.412.814.014.48.410.7
CTCNet2.8405.46.28.08.89.410.22.54.7
AV-DPRNN2.9015.46.210.210.711.912.37.89.0
AV-Mossformer22.9126.06.713.213.916.216.67.79.1
Swift-Net-123.2767.37.813.914.116.016.39.410.8
Dolphin2.9606.77.314.014.614.314.99.211.1
Falcon (Ours)3.2767.78.214.114.716.416.710.511.8
MethodParams (M) ↓MACs (G) ↓CPU Latency (ms) ↓GPU Latency (ms) ↓RAM (MB) ↓GPU Inf. (MB) ↓
AV-TF-GridNet6.2213.85646.3567.8136.31015.8
AV-Mossformer256.7224.216670.7114.9330.9539.8
Swift-Net-120.545.34163.8172.1482.3411.7
Dolphin5.618.12411.087.9560.2390.3
Falcon (Ours)0.510.61361.251.3125.3364.7
Efficiency comparison among Falcon and the four best-performing causal AVTSE baselines from the main results table. Lower is better for all metrics.
MethodTrained on VoxCeleb2Trained on RealSSA (Ours)
NISQA ↑DNSMOS ↑NISQA ↑DNSMOS ↑
AV-TF-GridNet2.4952.7572.4133.030
AV-Mossformer21.9952.7152.2562.912
Swift-Net-122.4813.0643.4843.276
Dolphin1.5932.6082.5902.960
Falcon (Ours)2.4312.8783.6153.276
Evaluation on RealSSA-Real when models were trained on VoxCeleb2 versus RealSSA. Training on RealSSA improved perceptual quality (NISQA and DNSMOS) across all tested architectures.

Ablation Studies

Ablations verify the role of downsampling depth, recurrent unit design, encoder-space speaker contrastive loss, and negative-weighting schedule.

Downsampling StagesSI-SNRi ↑SDRi ↑
17.17.7
27.47.9
3 (Ours)7.78.2
Recurrent UnitSI-SNRi ↑SDRi ↑Params (M) ↓MACs (G) ↓
RNN6.27.10.081.90
SRU7.47.90.152.71
GRU7.37.90.204.42
LSTM (Ours)7.78.20.255.69
ModelLossSI-SNRi ↑SDRi ↑
FalconWithout7.37.8
Falcon+ Ours7.78.2
SwiftNet-12Baseline7.37.8
SwiftNet-12+ Ours7.58.0
DolphinBaseline6.77.3
Dolphin+ Ours7.07.6
τ StrategySI-SNRi ↑SDRi ↑
Fixed τ = 07.58.1
Fixed τ = 17.47.9
Annealing τ: 0 → 27.58.0
Annealing τ: 0 → 1 (Ours)7.78.2

Citation

For double-blind submission, keep author fields anonymous. Replace after acceptance or when releasing publicly.

@inproceedings{falcon2026streamingavtse,
  title     = {Efficient Streaming Audio-Visual Target Speaker Extraction for Real-World Acoustic Scenes},
  author    = {Anonymous Authors},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}