Noise-robust bandwidth expansion aims to reconstruct high-fidelity wideband speech from noisy low-resolution inputs. While flow matching has shown strong performance in speech generation, accurately recovering clean speech from noisy inputs remains challenging due to the ambiguity of velocity estimation under noise. In this work, we propose VeRe-Flow, a clean-guided flow matching framework that introduces multi-level clean supervision to guide the generative process toward clean speech. At the velocity level, we introduce velocity contrastive regularization, which attracts the predicted velocity toward the clean trajectory while repelling it from noisy trajectories. At the representation level, we incorporate representation alignment that aligns intermediate features with clean self-supervised learning representations. The results demonstrate that the proposed method achieves the lowest LSD and highest DNSMOS OVRL among all baselines, and the highest MOS among generative baselines.
This work has been accepted to Interspeech 2026.
Figure 1. Overview of the proposed system
Comparison of 8k Noisy Input, 16k Predicted Speech (NU-Wave2, FlowHigh, VeRe-Flow (Proposed)), and 16k Clean Ground Truth.
* NU-Wave2 and FLowHigh are retrained under the same NR-BWE setting as the proposed method.
* We recommend listening with headphones for the best experience.
| Sample | 8k Noisy Input | NU-Wave2 | FLowHigh | VeRe-Flow (Proposed) | 16k Clean GT |
|---|---|---|---|---|---|
| Speaker p232 | |||||
| Sample 1 | |||||
| Sample 2 | |||||
| Sample 3 | |||||
| Sample 4 | |||||
| Sample 5 | |||||
| Sample 6 | |||||
| Sample 7 | |||||
| Sample 8 | |||||
| Sample 9 | |||||
| Sample 10 | |||||
| Speaker p257 | |||||
| Sample 1 | |||||
| Sample 2 | |||||
| Sample 3 | |||||
| Sample 4 | |||||
| Sample 5 | |||||
| Sample 6 | |||||
| Sample 7 | |||||
| Sample 8 | |||||
| Sample 9 | |||||
| Sample 10 | |||||
[1] C. Valentini-Botinhao et al., "Noisy speech database for training speech enhancement algorithms and TTS models," 2017.