Toward More Controllable AI Vi... Note

Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix

Netflix researchers have developed two AI tools to assist video editors in creating promotional assets. The first tool, Vera, is a layered video diffusion model designed for content-preserving edits. Vera generates edits as separate layers, leaving untouched portions of the original footage unaltered, thus preserving identities and details. To train Vera, they created a specialized dataset and employed a Mixture-of-Transformers architecture for efficient output generation. Evaluations showed Vera significantly outperformed existing methods in content preservation with comparable video quality. The second tool, VOID, is a video object and interaction deletion framework. VOID addresses the challenge of physically implausible results when removing objects with significant interactions. It uses a two-pass inference pipeline, with the first pass generating a physically plausible counterfactual video. A second pass refines the output to prevent artifacts like object morphing. VOID was trained using synthetic data generated from simulations and real-world motion capture data. Experiments indicate that VOID better preserves consistent scene dynamics compared to previous approaches. These research explorations aim to advance AI video editing responsibly while empowering artists. Both Vera and VOID are detailed in publicly released research papers.
CdXz5zHNQW_YZ3adN1Eip.gif