Résumé
Un utilisateur de Reddit a utilisé Claude pour créer un clip musical complet en moins de 24 heures, pour un coût de 37,88 $ en crédits fal. Le projet a nécessité l'utilisation de l'API Claude pour automatiser les tâches de conception et de montage, ainsi que la résolution de problèmes techniques pour améliorer la qualité du rendu final.
Pourquoi c’est intéressant
Ce projet est intéressant car il montre les possibilités de création de contenu multimédia avec Claude, ainsi que la capacité de l'outil à apprendre et à s'adapter aux besoins de l'utilisateur.
Comment Claude est utilisé
Le projet utilise Claude pour créer un clip musical via l'application PortOS, en exploitant les capacités d'automatisation et de génération de contenu de Claude pour concevoir les décors, les personnages et les animations.
Idées dérivées
- 01
ClipMusik
Créer un outil de génération de clips musicaux pour les artistes indépendants, en utilisant Claude pour automatiser les tâches de conception et de montage.
- 02
MediaGen
Développer un système de génération de contenu multimédia pour les entreprises, en utilisant Claude pour créer des vidéos institutionnelles et des présentations dynamiques.
- 03
MusicVid
Concevoir un jeu de création de musique et de vidéos, où les joueurs peuvent utiliser Claude pour générer des clips musicaux et des animations en fonction de leurs choix et de leurs préférences.
Afficher le post originalMasquer le post original
Video: https://youtu.be/XQgSZGIOleg
On Sunday a tune popped into my head in the car. After a day of iterating in Suno (64 versions, lyrics rewritten by hand), I pointed Claude (Opus 5.5, High effort) at my open-source app, https://github.com/atomantic/PortOS. I asked it to make a real music video through the app's Music Video autopilot, and to fix anything in the pipeline that got in its way.
What it cost
• fal.ai: $37.88 (Minimax H3 for lip-sync and image-to-video)
• Claude: about 10 hours of runtime, 5% of a weekly 20x plan
• Images: Codex image gen for the character sheet, sets and keyframes
How it worked
• It designed the singer and the sets from my mood board, then stopped for my sign-off.
• It mapped the beats and lyrics: which lines get lip-sync, where the cuts and glitch hits land.
• It built an HTML/canvas animatic first (every frame is a pure function of song time) and used it as the storyboard.
• It generated keyframes, then lip-sync and image-to-video clips through fal.
• The final cut is that same HTML composition with the real takes dropped in: VHS/CRT treatment, a lab HUD, kinetic type.
What it learned (now built into the autopilot)
• Feed the lip-sync model the isolated vocal stem, not the full mix. With the mix it misheard sung words: "swarm" came out as "swan", "run" as "vun".
• Whole-song transcription drifts after silences, so lyrics get anchored to vocal-stem phrases.
• Motion prompts must pin the keyframe's lighting, and must not inherit mood-board captions (those triggered provider rejections).
• Wide shots with tiny faces lose identity when the model pushes in, and crowd shots need an explicit "every dancer is a clone of her" clause.
• Overlapping lead and backing vocals can close the lips on n/d sounds ("need" became "meeb"). Roll two takes and compare mouth crops.
Along the way it closed 4 open issues and merged 5 PRs in PortOS. I even changed the song's ending lyric after the video was finished, and only the last segment had to be re-rendered.
The full starter prompt is in this thread: https://x.com/antic/status/2105308911452790926