Reserved English test set · real recorded noise

Real noise.
Paired truth.
Listen.

Forty held-out English clips with naturally noisy inputs and aligned Studio evaluation references, compared across Super Sonique’s Studio CFM DiT model and single-speaker Sidon.

Checkpoint Frozen KVAE 64-dim continuous Super Sonique Euler-8 · CFG 1.5 · temperature 0.80 Corruption None added
Super Sonique MOS recovery
Sidon MOS recovery
IPTV · Sonique / Sidon
Podcast · Sonique / Sidon

Episode-disjoint test

Two restorers. One panel.

Super Sonique’s 141M-parameter DiT model is trained from scratch on 500 hours of Studio clean-voice targets with conditional flow matching. Both restorers receive the original noisy recording. The “Noisy (codec)” player is its frozen KVAE reconstruction, matching the previous comparison’s playback and MOS baseline. Super Sonique uses Euler-8 with CFG 1.5, noise temperature 0.80 and the preferred sampling preset. Ground truth comes from the matched Studio evaluation audio. Each row uses shared playback scaling. DNSMOS P.835 OVRL scores the exact MP3s, and recovery is 100 × (restored MOS − noisy MOS) / (ground-truth MOS − noisy MOS). Scores are estimates; recovery can exceed 100% or be negative, and is shown as n/a when the ground-truth advantage is below 0.1 MOS.

Ground truthStudio evaluation reference
Noisy (codec)Real recording → frozen KVAE reconstruction
Super SoniqueStudio CFM DiT · Euler-8 · CFG 1.5
Sidon v0.1Single-speaker restoration model

Loading the real-noise evaluation panel…