Reserved English test set · real recorded noise

Real noise.
Paired truth.
Listen.

Forty held-out English clips with naturally noisy inputs and aggressively regenerated Auphonic speech-isolated, bandwidth-extended references, enhanced by the KVAE-Audio latent-flow model.

Checkpoint Frozen KVAE 64-dim continuous Solver Euler-1 Corruption None added
Overall MOS recovery
IPTV MOS recovery
Podcast MOS recovery
Held-out panel4020 IPTV · 20 podcast

Episode-disjoint test

No white noise this time.

The noisy column comes from real aligned recordings passed through the same frozen KVAE-Audio posterior-mean encoder and halo-aware decoder used by the model. No synthetic corruption or target leakage is added. The clean targets were regenerated with the more aggressive Auphonic speech-isolation and bandwidth-extension configuration, then receive the exact same gain as the paired noisy input. Enhancement uses one Euler model evaluation. Quality is estimated with Microsoft DNSMOS P.835 OVRL on the exact MP3s below. Recovery is 100 × (enhanced MOS − noisy MOS) ÷ (clean MOS − noisy MOS), without clamping; values above 100% mean the estimator rated enhanced audio above its paired clean reference. Per-clip recovery is marked n/a when the clean/noisy MOS gap is under 0.1.

Ground truthAligned clean reference
Real noisyRecording → frozen codec
EnhancedEuler-1 → frozen KVAE decoder

Loading the real-noise evaluation panel…