
JD.com announced today that it has open-sourced its self-developed real-time streaming video editing model, JoyAI-Video-Edit. Users can modify people and scenes while watching a video, transforming video creation from “editing after obtaining the footage” into real-time interactive editing.
JD.com officially stated that, in the field of real-time streaming video editing, JoyAI-Video-Edit comprehensively outperforms representative models currently available in the industry across all metrics, reaching a world-leading level.
Real-time video editing must first solve the problem of speed. A video consists of consecutive frames, and if the model cannot process them fast enough to keep up with playback, stuttering and latency will occur. Based on the fully self-developed JoyAI-Video model, JoyAI-Video-Edit achieves an inference speed of 30 frames per second at 720P resolution, keeping pace with the playback of common film and television videos while maintaining high image clarity.
Video length is another hurdle for streaming editing. Previous streaming editing models could only process clips lasting several seconds or minutes, while JoyAI-Video-Edit supports stable streaming editing for videos of any length and can continue processing as the video plays. Users do not need to wait for the entire video to finish generating; they can modify scenes, replace objects, adjust characters, or change the visual style in real time during playback, achieving “creation while watching.”

For example, creators can have people in a video continuously change outfits, transform an ordinary street into an animated world during a livestream, or try out different furniture, wall surfaces, and lighting effects in home design.


To evaluate the model’s capabilities, the research team compared JoyAI-Video-Edit with representative international streaming video editing models, including SANA Streaming, LiveEdit, and Xmax-X2.0. The results showed that, across evaluation data covering typical video editing scenarios, JoyAI-Video-Edit delivered comprehensively superior overall results.

In the authoritative OpenVE-Bench general evaluation, the model surpassed all currently available streaming editing methods. It showed particularly strong advantages in tasks such as global style transfer, local replacement, local deletion, and subtitle editing, substantially outperforming comparable models in the industry.
Beyond video creation, JoyAI-Video-Edit also provides a new technical approach for large-scale synthesis of embodied intelligence data. Robot training requires large amounts of video showing objects being grasped, transported, and manipulated, but collecting real-robot data is costly, while dangerous or rare scenarios are difficult to collect repeatedly. Leveraging its controllable editing capabilities for scenes, objects, and visual content, JoyAI-Video-Edit can convert videos of human operations into operation footage for grippers or robotic arms. While preserving object positions, spatial relationships, and motion trajectories, it can replace scenes, objects, and robot forms, expanding a single video into more training samples.
The open-source links for JoyAI-Video-Edit are as follows:
https://huggingface.co/jdopensource/JoyAI-Video-Edit
https://github.com/jd-opensource/JoyAI-Video-Edit
