Multimodal AI video generator

Combine images, video, audio, and text in one shot

Build a richer generation brief with multiple reference images, source clips, audio direction, and a precise prompt. Keep subjects, style, and motion aligned from one workspace.

Reference the whole creative brief

Up to 9 image references

Bring in character, product, environment, styling, and composition references.

Video references

Use source clips to communicate timing, movement, or performance direction.

Audio references and cues

Add rhythm, ambience, dialogue, or sound design to the generation brief.

Made with Cinelyo

See the workflow in motion

5-second example

Example prompt

The Floating Library

One continuous 5-second magical-realist shot inside an immense old library flooded with still ankle-deep water. Hundreds of open books float gently in the air beneath a vaulted glass ceiling, and one luminous white koi swims through the air between them. From 0-3s, the camera glides forward just above the reflective water on a 28mm lens while the koi crosses from right to left and loose pages respond to its movement. From 3-5s, the camera tilts upward as sunlight breaks through the ceiling and the floating books slowly align into a spiral. Coherent large-scale architecture, elegant restrained magic, realistic water reflections, warm dust-filled light, stable book shapes, graceful motion. Audio: soft page rustle, water movement, one distant resonant bell. No cuts, no people, no readable text, no logos, no watermark.

Seedance 2.0 Fast720p16:95s
Use this prompt

A clear path from idea to clip

A workflow built for iteration

01

Assemble references

Upload the visual, motion, and audio materials that define the shot.

02

Connect the references

Use @image, @video, and @audio in the prompt to clarify each input.

03

Generate a coherent take

Review the output and adjust only the reference or motion that needs refinement.

Built for real work

What this format is good for

Consistent products across a complete campaign scene

Character performance guided by a source movement clip

Brand worlds that combine art direction and sound

Music-led social concepts and event visuals

Storyboard frames translated into moving scenes

Reference-heavy product and fashion direction

Questions before you create

Multimodal Video FAQ

How many references can I provide?

The generator supports up to 9 reference images, up to 3 video references, and up to 3 audio references in a multimodal request.

How do I tell the model which reference to use?

Reference uploaded materials in the prompt with @image, @video, or @audio so the creative direction is explicit.

Do all references need to be about the same scene?

No. You can use different references for identity, environment, motion, styling, and sound, as long as the prompt explains their roles.