Unified Character Video Editing for Live Streaming
1 University of Macau · 2 vivo BlueImage Lab · 3 GVC Lab, Great Bay University
- Real-Time
- 14.47 FPS
- Causal Streaming
- Instruction-Guided
- Long-Stream Consistency
- Appearance-Motion Decoupled
Gallery
Add red frame sunglasses to his face.
Abstract
EditaLive is a framework for real-time streaming character video editing. Starting from a pretrained image animation model that naturally decouples appearance from motion, it is repurposed for instruction-based human-centric video editing through reference-frame editing and video reconstruction on CharEdit-50K. The model is then adapted from offline bidirectional to causal streaming generation and compressed into a two-step sampler with aligned self-rollout distillation. Fixed RoPE and Align Forcing reduce training-inference discrepancies, while First-Frame Preserved Sparse Attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate state-of-the-art editing performance, faithful facial-expression preservation, and low-latency real-time streaming inference.

Method
EditaLive addresses streaming character video editing through three consecutive stages. It first reformulates character video editing as appearance editing conditioned on explicit motion signals, then adapts the offline bidirectional video model to causal streaming generation using cached historical context, and finally compresses the causal model into a two-step sampler with aligned self-rollout distillation. Together, these stages provide explicit motion preservation and stable long-term generation.

EditaLive reformulates character video editing by decoupling the appearance to be edited from the motion to be preserved. A reference frame represents character appearance, while a skeleton sequence and implicit facial representations capture body motion and facial dynamics. For reconstruction-based training, only the reference image is synthetically edited; the real video remains the reconstruction target, paired with a reverse instruction for recovering the original appearance.
Cross-Character Editing
Combine the appearance of a reference character with motion signals extracted from a different driving video.
Turn it into a C4D felt doll style.

Add red frame sunglasses to his face.

Comparison
On CharEdit-Bench-S, EditaLive ranks first or second across all eight quality metrics. Character consistency is evaluated with ID-SIM, AED, and APD; editing quality with TA, EQ, BC, and SR; overall video quality with Pick Score; and efficiency with end-to-end FPS and average inter-chunk latency.
Short-video evaluation
150 five-second videos, 81 frames each, at 480 × 832.
| Method | Character Consistency | VLM Evaluation | Video Quality | Efficiency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ID-SIM ↑ | AED ↓ | APD ↓ | TA ↑ | EQ ↑ | BC ↑ | SR ↑ | Pick Score ↑ | FPS ↑ | Latency ↓ | |
| LiveEdit* | - | - | - | 1.542 | 1.949 | 1.640 | 0.393 | 19.26 | 17.25 | 0.696 |
| SANA-Streaming | 0.379 | 0.629 | 0.381 | 2.200 | 1.777 | 1.582 | 0.428 | 19.46 | 31.20 | 0.772 |
| StreamDiffusionV2 | 0.060 | 0.885 | 0.485 | 1.316 | 1.869 | 0.473 | 0.027 | 19.92 | 20.96 | 0.191 |
| LucyEdit* | - | - | - | 1.604 | 1.797 | 2.121 | 0.326 | 18.92 | 3.083 | 26.27 |
| UniVideo | 0.592 | 0.605 | 0.447 | 2.377 | 2.182 | 1.866 | 0.533 | 19.54 | 0.119 | 680.9 |
| Ditto | 0.334 | 0.685 | 0.355 | 1.560 | 1.796 | 1.127 | 0.167 | 19.49 | 0.294 | 275.4 |
| Qwen. + Wan. | 0.491 | 0.541 | 0.132 | 2.792 | 2.544 | 1.897 | 0.713 | 19.55 | 1.274 | 63.55 |
| EditaLive | 0.550 | 0.499 | 0.124 | 2.796 | 2.609 | 2.024 | 0.720 | 19.61 | 14.47 | 0.829 |
APD values are multiplied by 10. TA, EQ, BC, and SR denote text alignment, edit quality, background consistency, and success rate. Pick Score evaluates overall video quality. * LiveEdit and LucyEdit do not support global style editing, so their character-consistency metrics are omitted. Efficiency is measured on a single NVIDIA H100 using 81-frame clips at 384 × 672.
Long-video evaluation
30 videos longer than one minute, paired with two prompts for 60 cases.
| Method | Character Consistency | VLM Evaluation | Video Quality | Temporal Consistency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ID-SIM ↑ | AED ↓ | APD ↓ | TA ↑ | EQ ↑ | BC ↑ | SR ↑ | Pick Score ↑ | CLIP ↑ | DINO ↑ | |
| LiveEdit* | - | - | - | 1.583 | 1.700 | 1.044 | 0.017 | 18.91 | 97.36 | 98.38 |
| SANA-Streaming | 0.233 | 0.725 | 0.402 | 2.272 | 1.938 | 1.121 | 0.233 | 19.62 | 97.28 | 98.25 |
| StreamDiffusionV2 | 0.049 | 0.876 | 0.329 | 1.431 | 2.064 | 0.638 | 0.100 | 19.70 | 98.42 | 99.17 |
| Qwen. + Wan. | 0.398 | 0.659 | 0.126 | 2.733 | 2.304 | 1.776 | 0.750 | 19.69 | 98.30 | 98.65 |
| EditaLive | 0.492 | 0.576 | 0.109 | 2.761 | 2.678 | 2.183 | 0.817 | 19.78 | 98.51 | 98.96 |
APD values are multiplied by 10. TA, EQ, BC, and SR denote text alignment, edit quality, background consistency, and success rate. Pick Score evaluates overall video quality. EditaLive obtains the strongest overall performance in the long-video setting, including the best TA, EQ, BC, SR, and Pick Score.
Ablation Study
Replacing Align Forcing with Self Forcing causes the largest performance degradation because it introduces a mismatch in KV-cache construction between rollout training and streaming inference. Removing Fixed RoPE exposes the model to unseen positional offsets and weakens reference conditioning, while FPSA filters redundant historical information and preserves first-frame features to mitigate appearance drift.

Aligned self-rollout distillation
| Variant | Character Consistency | VLM Evaluation | Video Quality | Temporal Consistency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ID-SIM ↑ | AED ↓ | APD ↓ | TA ↑ | EQ ↑ | BC ↑ | SR ↑ | Pick Score ↑ | CLIP ↑ | DINO ↑ | |
| w/o Align Forcing | 0.252 | 0.711 | 0.338 | 2.703 | 2.567 | 1.306 | 0.367 | 20.03 | 98.62 | 99.07 |
| w/o Fixed RoPE | 0.381 | 0.618 | 0.120 | 2.728 | 2.639 | 1.872 | 0.617 | 19.91 | 98.50 | 98.95 |
| w/o FPSA | 0.455 | 0.609 | 0.116 | 2.737 | 2.600 | 2.105 | 0.800 | 19.85 | 98.46 | 98.88 |
| w/o First-Frame Preservation | 0.466 | 0.604 | 0.114 | 2.725 | 2.622 | 2.139 | 0.833 | 19.84 | 98.42 | 98.89 |
| Full EditaLive | 0.492 | 0.576 | 0.109 | 2.761 | 2.678 | 2.183 | 0.817 | 19.78 | 98.51 | 98.96 |
Citation
If you find EditaLive useful for your research, welcome to cite our work using the following BibTeX:
@article{li2026editalive,
title = {EditaLive! Unified Character Video Editing for Live Streaming},
author = {Li, Zhiyuan and Pun, Chi-Man and Jiang, Peng-Tao and Li, Bo and Cun, Xiaodong},
journal = {arXiv preprint arXiv:2608.27123},
year = {2026}
}