Episode-disjoint test
Two restorers. One panel.
Super Sonique’s 141M-parameter DiT model is trained from scratch on 500 hours of Studio clean-voice targets with conditional flow matching. Both restorers receive the original noisy recording. The “Noisy (codec)” player is its frozen KVAE reconstruction, matching the previous comparison’s playback and MOS baseline. Super Sonique uses Euler-8 with CFG 1.5, noise temperature 0.80 and the preferred sampling preset. Ground truth comes from the matched Studio evaluation audio. Each row uses shared playback scaling. DNSMOS P.835 OVRL scores the exact MP3s, and recovery is 100 × (restored MOS − noisy MOS) / (ground-truth MOS − noisy MOS). Scores are estimates; recovery can exceed 100% or be negative, and is shown as n/a when the ground-truth advantage is below 0.1 MOS.
Loading the real-noise evaluation panel…