ManifoldCache

Training-Free Diffusion Acceleration via Constraint Manifold Caching

ManifoldCache delivers acceleration for scientific diffusion models through: (a) a unified CMDM abstraction covering several models across five domains, including brain MRI, molecules, proteins, crystals, and 4D scenes spanning radically different ambient spaces from voxel grids to motif token graphs to crystal lattices to 4D video latents under one closed-form schedule; (b) true backbone-agnosticism across fundamentally different families, including volumetric 3D ConvNets, SE(3)/E(3)-equivariant networks, graph diffusion transformers, and multi-view video DiTs including discrete diffusion; (c) systematic failure analysis showing quantization causes OOM, pruning breaks constraints, fast ODE solvers drift off the manifold, and standard caching induces mode confusion; (d) a provably sharp safe-caching threshold T* with both bounded-error and guaranteed-confusion regimes confirmed by ablations; (e) depth-adaptive caching from Jacobian decomposition, where deeper blocks get larger strides with bounded and competitive VRAM overhead; and (f) completely training-free and data-free with zero calibration, zero retraining, working out-of-the-box on existing pretrained checkpoints while delivering consistent joint dominance on both speed and quality across all models.

3D Brain MRI Synthesis Comparison (Med-DDPM) (Drag to rotate, scroll to zoom)
Full inference FID: 1.21 (↓) · Time: 165 secs (↓)
ManifoldCache (Ours)
FID: 6.93 (↓)  ·  Time: 91 secs (↓)
DeepCache
FID: 20.14 (↓)  ·  Time: 143 secs (↓)
Lyra Multi-View Synthesis Comparison
Full inference Time: 5577 secs (↓)
ManifoldCache (Ours)
PSNR: 39.27 (↑)  ·  SSIM: 0.9817 (↑)  ·  LPIPS: 0.0564 (↓)  ·  Time: 3416 secs (↓)
PAB
PSNR: 36.85 (↑)  ·  SSIM: 0.9750 (↑)  ·  LPIPS: 0.0552 (↓)  ·  Time: 4744 secs (↓)
Med-DDPM Cache Stride Profile (ManifoldCache)
Stride Spectrum
Sorted by depth · deeper blocks receive heavier stride allocations
UNet Block Depth Stride