A text prompt drives Suno to write and generate a full song, then per-scene image stills are generated with Flux and animated into short clips with Runway, which are composited against the finished track into a music video.
Properties
family
multi-shot
modality
text -> audio -> image -> video
whyMultiModel
Chains three distinct generative-media models across two modalities: Suno generates the music track itself (not just narration), Flux generates the scene artwork, and Runway generates video motion from those stills; a rendering step composites them. This differs from the baseline's beat-synced music-video pattern (which pairs a generated song with a beat-synced *editing* tool over existing footage) because here every visual asset, not just the audio, is freshly generated per scene.
steps
[object Object], [object Object], [object Object] +1 more
n8n + Suno API + Flux (BlackForest Labs/RapidAPI) + Runway ML + Creatomate
useCases
[AI music video generation][social lyric videos][artist promo clips from a single prompt]
pitfalls
per-scene image/video costs compound quickly across a full song length; visual continuity between scenes is not enforced by any identity/consistency model