OmniVR

Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

Xin Lu, Zihao Fan, Jie Huang, Mingchen Zhong, Xueyang Fu, Zheng-Jun Zha

University of Science and Technology of China (USTC)

Equal contribution    Project leader    Corresponding author

Background: a historical clip restored by OmniVR

OmniVR is the first joint audio-video generative restoration model for historical films — recovering visual structure, temporal motion, and acoustic detail under one coordinated objective, so picture and sound are healed together.

0B
Audio-Video Backbone
Multimodal DiT
0 Modalities
Jointly Restored
Video + Audio
0 Axes
OmniVRBench
Visual / Audio / Temporal / Sync
0
Real Historical Clips
Benchmark set

Abstract

Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization—the first method to jointly address all three aspects.

Key Contributions

Three designs that turn a text-to-audio-video generator into a joint old-film restorer while preserving its generative prior.

Joint AV Degradation Pipeline

Simulates real old-film characteristics—blur, noise, flicker, hiss, clipping, dropout—from Internet-collected data, producing paired low-quality audio-video for training.

Prior-Preserving T2AV → AV2AV

An architecture-preserving transition with prompt annealing recasts text-to-audio-video generation as audio-video-to-audio-video restoration, maximally retaining the generative prior.

I2V Anchoring & Waveform Supervision

First-frame image-to-video anchoring with loss reweighting and waveform supervision enables long-video extrapolation and high audio fidelity.

Stream It: Before ↔ After

Drag the divider to wipe between the degraded input (left) and the OmniVR restoration (right). Then use the Hear the difference panel below to A/B the degraded vs restored soundtrack — hold either side to compare in place.

Degraded Input OmniVR
Hear the difference
Listen

Press & hold either side to A/B in place · one track plays at a time

Citation

@article{lu2026omnivr,
  title={OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films},
  author={Lu, Xin and Fan, Zihao and Zhong, Mingchen and Huang, Jie and Fu, Xueyang and Zha, Zheng-Jun},
  journal={arXiv preprint arXiv:2608.04224},
  year={2026}
}