EconPapers    
Economics at your fingertips  
 

OmniVideos: Multi-Factor Control for Versatile Diffused Video Generation

Rohini Matta and Ayan Rajput

International Journal of Scientific Research in Science and Technology, 2026, vol. 13, issue 3, 130-144

Abstract: While modern video diffusion models have achieved impressive aesthetic results, the lack of granular, high-precision control remains a significant hurdle for professional content creators. For AI-driven video production to be truly effective, three core elements must be harmonized: the composition of the scene, the ability to customize subjects while maintaining consistency across different views, and the precise direction of both camera and object movement. Currently, most technologies address these requirements in isolation, which often leads to a failure in preserving a subject's identity when camera angles or poses change. This lack of a cohesive architecture has historically made it difficult to generate versatile videos that are jointly controllable across multiple dimensions. To address these limitations, we present an integrated framework and a two-stage training system designed to unify scene composition, subject consistency, and motion dynamics. Our method utilizes a dual-condition motion module that processes the environment and the subject differently: it uses 3D tracking points to stabilize background scenes and downsampled RGB cues to anchor the foreground subjects. Furthermore, we introduce an inference-time ControlNet scale schedule, which ensures a fluid balance between strict structural control and high-quality visual realism. OmniVideos facilitates advanced creative workflows, such as the 3D-aware insertion of subjects into any environment and the direct manipulation of subjects from a static image. Our experimental evaluations show that OmniVideos significantly outperforms leading specialized models, such as Phantom [22] and DaS [11], particularly in its ability to maintain subject identity across multiple views, ensure 3D spatial consistency, and execute motion with high accuracy.

Keywords: Video Diffusion Architecture; Subject-Centric Video Synthesis; Integrated Motion Direction (search for similar items in EconPapers)
Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://ijsrst.com/home/article/view/IJSRST26133126 Abstract page (text/html)
https://ijsrst.com/home/article/download/IJSRST26133126/IJSRST26133126 Full text (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:etm:ijsrst:v13:y2026:i3:id:1581

DOI: 10.32628/IJSRST26133126

Access Statistics for this article

More articles in International Journal of Scientific Research in Science and Technology from Technoscience Academy
Bibliographic data for series maintained by Pankaj Sharma ().

 
Page updated 2026-07-27
Handle: RePEc:etm:ijsrst:v13:y2026:i3:id:1581