Unified Character Video Editing for Live Streaming

Zhiyuan Li1,3Chi-Man Pun1,*Peng-Tao Jiang2,*Bo Li2Xiaodong Cun3,†

1 University of Macau · 2 vivo BlueImage Lab · 3 GVC Lab, Great Bay University

  • Real-Time
  • 14.47 FPS
  • Causal Streaming
  • Instruction-Guided
  • Long-Stream Consistency
  • Appearance-Motion Decoupled
A two-minute EditaLive showreel presenting source videos and edited results across short and long examples.

Abstract

EditaLive is a framework for real-time streaming character video editing. Starting from a pretrained image animation model that naturally decouples appearance from motion, it is repurposed for instruction-based human-centric video editing through reference-frame editing and video reconstruction on CharEdit-50K. The model is then adapted from offline bidirectional to causal streaming generation and compressed into a two-step sampler with aligned self-rollout distillation. Fixed RoPE and Align Forcing reduce training-inference discrepancies, while First-Frame Preserved Sparse Attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate state-of-the-art editing performance, faithful facial-expression preservation, and low-latency real-time streaming inference.

Comparison of general video-to-video editing, a cascaded editor-animation pipeline, and the appearance-motion decoupled EditaLive paradigm
Comparison of general video-to-video editing, a cascaded image-editing-and-animation pipeline, and EditaLive's appearance-motion decoupled formulation.

Method

EditaLive addresses streaming character video editing through three consecutive stages. It first reformulates character video editing as appearance editing conditioned on explicit motion signals, then adapts the offline bidirectional video model to causal streaming generation using cached historical context, and finally compresses the causal model into a two-step sampler with aligned self-rollout distillation. Together, these stages provide explicit motion preservation and stable long-term generation.

Three-stage EditaLive pipeline: appearance-motion decoupled editing, causal streaming adaptation, and aligned self-rollout distillation
Overview of the three-stage training pipeline: appearance-motion decoupled editing, causal streaming adaptation, and aligned self-rollout distillation.

EditaLive reformulates character video editing by decoupling the appearance to be edited from the motion to be preserved. A reference frame represents character appearance, while a skeleton sequence and implicit facial representations capture body motion and facial dynamics. For reconstruction-based training, only the reference image is synthetically edited; the real video remains the reconstruction target, paired with a reverse instruction for recovering the original appearance.

Cross-Character Editing

Combine the appearance of a reference character with motion signals extracted from a different driving video.

Example 1

Turn it into a C4D felt doll style.

Reference character for cross-character editing example 1
Reference
Driving video
EditaLive
Example 2

Add red frame sunglasses to his face.

Reference character for cross-character editing example 2
Reference
Driving video
EditaLive

Comparison

On CharEdit-Bench-S, EditaLive ranks first or second across all eight quality metrics. Character consistency is evaluated with ID-SIM, AED, and APD; editing quality with TA, EQ, BC, and SR; overall video quality with Pick Score; and efficiency with end-to-end FPS and average inter-chunk latency.

Single H100 @ 384 × 67214.47 FPS
Inter-chunk latency0.829 s
Short-video success0.720
Flash-VAED enabled16.4 FPS
CharEdit-Bench-S

Short-video evaluation

150 five-second videos, 81 frames each, at 480 × 832.

Selected quantitative comparisons on CharEdit-Bench-S
MethodCharacter ConsistencyVLM EvaluationVideo QualityEfficiency
ID-SIM ↑AED ↓APD ↓TA ↑EQ ↑BC ↑SR ↑Pick Score ↑FPS ↑Latency ↓
LiveEdit*---1.5421.9491.6400.39319.2617.250.696
SANA-Streaming0.3790.6290.3812.2001.7771.5820.42819.4631.200.772
StreamDiffusionV20.0600.8850.4851.3161.8690.4730.02719.9220.960.191
LucyEdit*---1.6041.7972.1210.32618.923.08326.27
UniVideo0.5920.6050.4472.3772.1821.8660.53319.540.119680.9
Ditto0.3340.6850.3551.5601.7961.1270.16719.490.294275.4
Qwen. + Wan.0.4910.5410.1322.7922.5441.8970.71319.551.27463.55
EditaLive0.5500.4990.1242.7962.6092.0240.72019.6114.470.829

APD values are multiplied by 10. TA, EQ, BC, and SR denote text alignment, edit quality, background consistency, and success rate. Pick Score evaluates overall video quality. * LiveEdit and LucyEdit do not support global style editing, so their character-consistency metrics are omitted. Efficiency is measured on a single NVIDIA H100 using 81-frame clips at 384 × 672.

CharEdit-Bench-L

Long-video evaluation

30 videos longer than one minute, paired with two prompts for 60 cases.

Selected quantitative comparisons on CharEdit-Bench-L
MethodCharacter ConsistencyVLM EvaluationVideo QualityTemporal Consistency
ID-SIM ↑AED ↓APD ↓TA ↑EQ ↑BC ↑SR ↑Pick Score ↑CLIP ↑DINO ↑
LiveEdit*---1.5831.7001.0440.01718.9197.3698.38
SANA-Streaming0.2330.7250.4022.2721.9381.1210.23319.6297.2898.25
StreamDiffusionV20.0490.8760.3291.4312.0640.6380.10019.7098.4299.17
Qwen. + Wan.0.3980.6590.1262.7332.3041.7760.75019.6998.3098.65
EditaLive0.4920.5760.1092.7612.6782.1830.81719.7898.5198.96

APD values are multiplied by 10. TA, EQ, BC, and SR denote text alignment, edit quality, background consistency, and success rate. Pick Score evaluates overall video quality. EditaLive obtains the strongest overall performance in the long-video setting, including the best TA, EQ, BC, SR, and Pick Score.

Ablation Study

Replacing Align Forcing with Self Forcing causes the largest performance degradation because it introduces a mismatch in KV-cache construction between rollout training and streaming inference. Removing Fixed RoPE exposes the model to unseen positional offsets and weakens reference conditioning, while FPSA filters redundant historical information and preserves first-frame features to mitigate appearance drift.

Visual ablation of causal streaming adaptation and aligned self-rollout distillation variants
Ablation study on causal streaming adaptation and aligned self-rollout distillation.
CharEdit-Bench-L

Aligned self-rollout distillation

Ablation study on aligned self-rollout distillation
VariantCharacter ConsistencyVLM EvaluationVideo QualityTemporal Consistency
ID-SIM ↑AED ↓APD ↓TA ↑EQ ↑BC ↑SR ↑Pick Score ↑CLIP ↑DINO ↑
w/o Align Forcing0.2520.7110.3382.7032.5671.3060.36720.0398.6299.07
w/o Fixed RoPE0.3810.6180.1202.7282.6391.8720.61719.9198.5098.95
w/o FPSA0.4550.6090.1162.7372.6002.1050.80019.8598.4698.88
w/o First-Frame Preservation0.4660.6040.1142.7252.6222.1390.83319.8498.4298.89
Full EditaLive0.4920.5760.1092.7612.6782.1830.81719.7898.5198.96

Citation

If you find EditaLive useful for your research, welcome to cite our work using the following BibTeX:

BibTeXarXiv:2608.27123
@article{li2026editalive,
  title   = {EditaLive! Unified Character Video Editing for Live Streaming},
  author  = {Li, Zhiyuan and Pun, Chi-Man and Jiang, Peng-Tao and Li, Bo and Cun, Xiaodong},
  journal = {arXiv preprint arXiv:2608.27123},
  year    = {2026}
}