Episode-disjoint test
No white noise this time.
The noisy column comes from real aligned recordings passed through the same frozen KVAE-Audio posterior-mean encoder and halo-aware decoder used by the model. No synthetic corruption or target leakage is added. The clean targets were regenerated with the more aggressive Auphonic speech-isolation and bandwidth-extension configuration, then receive the exact same gain as the paired noisy input. Enhancement uses one Euler model evaluation. Quality is estimated with Microsoft DNSMOS P.835 OVRL on the exact MP3s below. Recovery is 100 × (enhanced MOS − noisy MOS) ÷ (clean MOS − noisy MOS), without clamping; values above 100% mean the estimator rated enhanced audio above its paired clean reference. Per-clip recovery is marked n/a when the clean/noisy MOS gap is under 0.1.
Loading the real-noise evaluation panel…