Multimodal AI video generator
Combine images, video, audio, and text in one shot
Build a richer generation brief with multiple reference images, source clips, audio direction, and a precise prompt. Keep subjects, style, and motion aligned from one workspace.
Reference the whole creative brief
Up to 9 image references
Bring in character, product, environment, styling, and composition references.
Video references
Use source clips to communicate timing, movement, or performance direction.
Audio references and cues
Add rhythm, ambience, dialogue, or sound design to the generation brief.
Made with Cinelyo
See the workflow in motion
Example prompt
The Floating Library
One continuous 5-second magical-realist shot inside an immense old library flooded with still ankle-deep water. Hundreds of open books float gently in the air beneath a vaulted glass ceiling, and one luminous white koi swims through the air between them. From 0-3s, the camera glides forward just above the reflective water on a 28mm lens while the koi crosses from right to left and loose pages respond to its movement. From 3-5s, the camera tilts upward as sunlight breaks through the ceiling and the floating books slowly align into a spiral. Coherent large-scale architecture, elegant restrained magic, realistic water reflections, warm dust-filled light, stable book shapes, graceful motion. Audio: soft page rustle, water movement, one distant resonant bell. No cuts, no people, no readable text, no logos, no watermark.
A clear path from idea to clip
A workflow built for iteration
Assemble references
Upload the visual, motion, and audio materials that define the shot.
Connect the references
Use @image, @video, and @audio in the prompt to clarify each input.
Generate a coherent take
Review the output and adjust only the reference or motion that needs refinement.
Built for real work
What this format is good for
Consistent products across a complete campaign scene
Character performance guided by a source movement clip
Brand worlds that combine art direction and sound
Music-led social concepts and event visuals
Storyboard frames translated into moving scenes
Reference-heavy product and fashion direction
Questions before you create
Multimodal Video FAQ
How many references can I provide?
The generator supports up to 9 reference images, up to 3 video references, and up to 3 audio references in a multimodal request.
How do I tell the model which reference to use?
Reference uploaded materials in the prompt with @image, @video, or @audio so the creative direction is explicit.
Do all references need to be about the same scene?
No. You can use different references for identity, environment, motion, styling, and sound, as long as the prompt explains their roles.
