Summary of Key Points
The research team led by Academician Zhang Wenjun from Shanghai Jiao Tong University has developed the "World Narrative Model (WNM)." This model addresses the inefficiency issue in current AI video generation, which is characterized by a "lottery-like" approach where professional creators often need to try 20-50 times to produce a satisfactory shot. By utilizing a "controller + renderer" architecture, WNM transforms the process from a black-box probability sampling to a precisely controllable physical parameter planning and rendering method, significantly improving efficiency (reducing the number of edits per shot to less than three times and achieving a satisfaction rate of over 80%). The model has been piloted at a Shanghai-based AI micro-short drama production base with the aim of revolutionizing the underlying paradigms of film and television production. The goal is for Shanghai to evolve from a consumer of micro-short dramas to a leader in technology development and standard setting.
I. The Pain Points of "Lottery-Like" AI Video Generation
Professional creators using AI for video production face a situation similar to playing a lottery game. They provide a textual description (for example, "The protagonist is running in the rain, with the camera following from the side"), wait for several seconds to see the generated video, and if it doesn't meet their standards (such as the rain being too heavy to clearly show the face or the camera shaking), they have to modify the description and start over. According to research by Shanghai Jiao Tong University, it takes on average 20-50 attempts to produce a decent shot, with the success rate for high-quality shots being less than 50%.
Why is this so inefficient? Existing AI video models are essentially black boxes: the input text directly produces the output image, and the process behind it is unknown and uncontrollable. For instance, if a director wants the lighting to come from the left front, the model might randomly assign light from the right back, forcing the creator to start over, relying on chance.
II. The Core Innovation of WNM: Giving AI a "Professional Steering Wheel"
WNM divides the video generation process into two components:
1. Controller (the World Narrative Model itself): This component understands physics and performs planning. It converts the director's ideas (script, shot composition) into structured physical parameters—whether the scene is 3D or 2D, whether characters should walk or run, and the type of camera movement (zoom in/out, etc.). These parameters are like a detailed "script guide" for filming.
2. Renderer (existing video models): This component takes the parameters provided by the controller and renders the image without needing to guess what the director intended.
For example, if a director wants to adjust a character's gesture, they previously had to modify the description and generate a new shot; now, they can simply drag the relevant parameter in the controller, and the renderer immediately produces the updated image. This is like giving AI a professional steering wheel, allowing the director to precisely control every detail without relying on chance.
III. Differences Between WNM and Other Models
- Compared to Google Genie (a world model): Genie is more for gaming, allowing users to explore virtual worlds created by the model, but within fixed scenarios; WNM is designed for professional production, with all physical parameters being independently adjustable to meet film and television requirements.
- Compared to end-to-end models like Kling and Veo: These models are black boxes that process input text directly into output images without any intermediate adjustments. WNM, on the other hand, allows for step-by-step parameter planning before rendering.
In simple terms, Genie is a toy, end-to-end models are like blind boxes, while WNM is a professional tool.
IV. Two Challenges Faced by WNM
1. Data scarcity: Training the controller requires high-precision 3D data with physical annotations (e.g., the size, material, and movement trajectory of an object). This type of data is much rarer than what's available in standard internet videos. The team has developed an automated annotation pipeline to reduce manual labor costs.
2. Consistency over long durations: Generating videos longer than 5 minutes can lead to issues such as sudden changes in character appearance or disordered scene layouts. WNM addresses this by using the controller to maintain consistent physical states across frames, ensuring the video remains coherent from start to finish.
V. Implementation and Commercialization
- Shanghai Pilot Base: The model is being tested at a pilot base in Shanghai, with various commercialization strategies:
- Policy support: Shanghai has issued guidelines (AI Micro-Short Drama Shanghai 8) and established bases in districts like Xuhui, offering subsidies of up to 10 million yuan for independent research projects. WNM is the leading technology in this initiative, with computing power provided by JiuZhang Cloud Extreme. The goal is to integrate the model into short drama and film production processes to shorten production times.
- Commercialization models:
- Small and medium-sized teams can use it via SaaS subscriptions (monthly fees).
- Large film and television companies can deploy it on their own servers.
- Developers can use the API on a pay-per-use basis.
- Industry significance: While current AI video models focus on image quality, WNM emphasizes controllability as the key factor. The team hopes that WNM will set new technical standards for the film and television industry, transforming Shanghai from a consumer of micro-short dramas into a hub for technological innovation.
Conclusion
WNM marks a crucial step in the evolution of AI video generation, moving from mere capability to true control. It represents a fundamental shift in approach, similar to how large language models have progressed from guessing text to understanding semantics. The success of WNM will depend on its practical application in the pilot base and its ability to solve real-world challenges in the film and television industry.