# elastic-forcing
**Repository Path**: liu-yueyi/elastic-forcing
## Basic Information
- **Project Name**: elastic-forcing
- **Description**: No description available
- **Primary Language**: Unknown
- **License**: Apache-2.0
- **Default Branch**: main
- **Homepage**: None
- **GVP Project**: No
## Statistics
- **Stars**: 2
- **Forks**: 0
- **Created**: 2026-09-28
- **Last Updated**: 2026-09-29
## Categories & Tags
**Categories**: Uncategorized
**Tags**: None
## README
# Elastic Forcing
### From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
Chi Zhang1,* · Yueyi Liu1,2,* · Haoyang Shi1,3,* · Ruichuan An4
Haoyu Li2 · Yuhang Wu1 · Sen Cui5 · Miao Liu1,†
1 College of AI, Tsinghua University
2 IAIR, Xi’an Jiaotong University · 3 Xianghui Academy, Fudan University
4 Peking University · 5 BAAI
* Equal contribution. † Corresponding author.
[Contact Chi Zhang](mailto:imzc.2004@gmail.com) · [Contact Miao Liu](mailto:miaoliu@mail.tsinghua.edu.cn)
[](https://raw.githubusercontent.com/yueyiliu08/Elastic-Forcing/main/docs/elastic-forcing.pdf?v=3399996653f7)
[](https://video-examples-m8r2v6.pages.dev/)
[](https://github.com/yueyiliu08/Elastic-Forcing)
[](https://huggingface.co/collections/liuyueyi-8/elastic-forcing-6aba0e62a4416b13ae7b13bf)

[1.3B Training & Inference](1.3B/README.md) · [14B Training & Inference](14B/README.md)
---
### Unfortunately, we cannot submit the code to GitHub at the moment due to an unknown issue. Several accounts have been banned, so we are investigating the cause. We are working hard to resolve it and will share the GitHub link as soon as possible.
### As a friendly reminder, please do not upload any part of this code to GitHub, as doing so may result in your account being banned.
Elastic Forcing post-trains few-step autoregressive video generators by matching generated and reference videos in frozen representation spaces. It combines a hybrid Nyström–Monte Carlo distribution estimator with gradient replay. Only the generator is optimized; diffusion score models are not needed during post-training.
Both released recipes use the same three-encoder MMD objective; no auxiliary frame-level distribution objective is enabled or included.
Read the implementation through the [paper-to-code guide](docs/paper-to-code.md), which maps the paper's notation to configuration fields, the training loop, the hybrid MMD objective, and gradient replay.
## Models
Each model directory contains its own implementation, configuration, dependency list, asset guide, and training/inference scripts. Run its commands from that directory.
| Release | Frozen feature encoders | Training | Inference | Guide |
|---|---|---|---|---|
| **1.3B** | DINOv3 + V-JEPA 2 + VideoMAEv2 | 4 GPUs, 150 updates | 4 causal steps per chunk | [Training and inference](1.3B/README.md) |
| **14B** | DINOv3 + V-JEPA 2 + VideoMAEv2 | 8 GPUs, 80 updates | 4 trained steps + terminal zero-step cleanup | [Training and inference](14B/README.md) |
Both configurations use a distribution batch of 256 generated videos, a differentiation batch of 128 videos, and a reference minibatch of 128 videos. They generate 81-frame, 832 × 480 videos at 16 FPS.
The 1.3B three-encoder configuration achieves **84.25 VBench Total** in the paper. The 14B release targets the **step-80 checkpoint**, retaining the first 80 updates of the original **200-step cosine learning-rate schedule**. This is the **three-encoder, 8,000-reference-video model used in the human and VLM evaluations**; see [evaluated settings](14B/docs/evaluated-settings.md).
```text
Elastic-Forcing/
├── 1.3B/ # Three-encoder training and inference
│ ├── code/
│ ├── configs/
│ ├── scripts/
│ └── README.md
├── 14B/ # Three-encoder training and step-80 inference
│ ├── code/
│ ├── configs/
│ ├── scripts/
│ └── README.md
└── docs/elastic-forcing.pdf
```
## Getting Started
```bash
git clone https://github.com/yueyiliu08/Elastic-Forcing.git
cd Elastic-Forcing
# Choose one model, then follow its README.
cd 14B
# Or: cd 1.3B
```
See the [1.3B asset guide](1.3B/docs/assets.md) or [14B asset guide](14B/docs/assets.md) for model and reference-data requirements. This GitHub release contains source code and the paper; trained generator weights, the 14B ODE initialization, and frozen reference assets are stored separately and are not bundled here. Download the released model checkpoints and training data from our [Hugging Face collection](https://huggingface.co/collections/liuyueyi-8/elastic-forcing-6aba0e62a4416b13ae7b13bf).
Validation is documented separately for [1.3B](1.3B/docs/validation.md) and [14B](14B/docs/validation.md).
The cleaned source-layout and CPU objective checks are described in [source validation](docs/readability-changes.md).
## Acknowledgments
This implementation builds on [Self-Forcing](https://github.com/guandeh17/Self-Forcing) and [Wan2.1](https://github.com/Wan-Video/Wan2.1). We also thank the authors of [DINOv3](https://github.com/facebookresearch/dinov3), [V-JEPA 2](https://github.com/facebookresearch/vjepa2), [VideoMAEv2](https://github.com/OpenGVLab/VideoMAEv2), and [VBench](https://github.com/Vchitect/VBench) for making their work available. Bundled source retains its upstream license notices; model weights and datasets follow their respective upstream terms.
We are also deeply grateful to the Krea AI team for generously providing the ODE-initialized checkpoint for
Wan2.1-T2V-14B. Their support allowed us to avoid reproducing an exceptionally computationally intensive
initialization procedure at this scale, saving a substantial amount of GPU time and making our large-scale
experiments possible.