VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

Abstract

Noise-robust bandwidth expansion aims to reconstruct high-fidelity wideband speech from noisy low-resolution inputs. While flow matching has shown strong performance in speech generation, accurately recovering clean speech from noisy inputs remains challenging due to the ambiguity of velocity estimation under noise. In this work, we propose VeRe-Flow, a clean-guided flow matching framework that introduces multi-level clean supervision to guide the generative process toward clean speech. At the velocity level, we introduce velocity contrastive regularization, which attracts the predicted velocity toward the clean trajectory while repelling it from noisy trajectories. At the representation level, we incorporate representation alignment that aligns intermediate features with clean self-supervised learning representations. The results demonstrate that the proposed method achieves the lowest LSD and highest DNSMOS OVRL among all baselines, and the highest MOS among generative baselines.

This work has been accepted to Interspeech 2026.

Model Structure

Figure 1. Overview of the proposed system


Audio Samples (Valentini-Botinhao testset)

Comparison of 8k Noisy Input, 16k Predicted Speech (NU-Wave2, FlowHigh, VeRe-Flow (Proposed)), and 16k Clean Ground Truth.

* NU-Wave2 and FLowHigh are retrained under the same NR-BWE setting as the proposed method.

* We recommend listening with headphones for the best experience.

Sample 8k Noisy Input NU-Wave2 FLowHigh VeRe-Flow (Proposed) 16k Clean GT
Speaker p232
Sample 1
Sample 2
Sample 3
Sample 4
Sample 5
Sample 6
Sample 7
Sample 8
Sample 9
Sample 10
Speaker p257
Sample 1
Sample 2
Sample 3
Sample 4
Sample 5
Sample 6
Sample 7
Sample 8
Sample 9
Sample 10

References

[1] C. Valentini-Botinhao et al., "Noisy speech database for training speech enhancement algorithms and TTS models," 2017.