Joint Video-Audio Conditional Generation
for Restoring Degraded Historical Films
University of Science and Technology of China (USTC)
† Equal contribution ‡ Project leader ✉ Corresponding author
Background: a historical clip restored by OmniVR
OmniVR is the first joint audio-video generative restoration model for historical films — recovering visual structure, temporal motion, and acoustic detail under one coordinated objective, so picture and sound are healed together.
Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization—the first method to jointly address all three aspects.
Three designs that turn a text-to-audio-video generator into a joint old-film restorer while preserving its generative prior.
Simulates real old-film characteristics—blur, noise, flicker, hiss, clipping, dropout—from Internet-collected data, producing paired low-quality audio-video for training.
An architecture-preserving transition with prompt annealing recasts text-to-audio-video generation as audio-video-to-audio-video restoration, maximally retaining the generative prior.
First-frame image-to-video anchoring with loss reweighting and waveform supervision enables long-video extrapolation and high audio fidelity.
Drag the divider to wipe between the degraded input (left) and the OmniVR restoration (right). Then use the Hear the difference panel below to A/B the degraded vs restored soundtrack — hold either side to compare in place.
Press & hold either side to A/B in place · one track plays at a time
Frequency (low → high, bottom → top) over time · identical scale for both
Every pair is an independent slider. Drag any divider to compare input and restoration.
@article{lu2026omnivr,
title={OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films},
author={Lu, Xin and Fan, Zihao and Zhong, Mingchen and Huang, Jie and Fu, Xueyang and Zha, Zheng-Jun},
journal={arXiv preprint arXiv:2608.04224},
year={2026}
}