Expérimental~63 · IA en attente
New AI Slop standard?
r/ClaudeAIu/imalegonzalez1 octobre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
Full song Ive been experimenting with a pretty simple AI music video workflow and honestly the result surprised me. For the music, I started by collecting a few reference songs and ran them through a music analysis agent. The idea was to get a proper breakdown of the references (structure, production, energy, arrangem…
Afficher le post originalMasquer le post original
Full song https://youtu.be/RU7eRodudik
Ive been experimenting with a pretty simple AI music video workflow and honestly the result surprised me.
For the music, I started by collecting a few reference songs and ran them through a music analysis agent. The idea was to get a proper breakdown of the references (structure, production, energy, arrangement, mood, synth names, etc) and then have the agent connect the dots between them and turn that into a base prompt for SUNO.
I gave SUNO a lot of creative freedom with the lyrics. I only defined the general theme and the emotional direction I wanted. After around 3–4 iterations with SUNO v6, I had the track.
From there I moved everything local.
First I collected around 20 visual references, mostly from Pinterest, that matched the look and feel I had in mind and put them into a reference folder.
Then, inspired by a prompt shared by donaldjewkes on X, I built my own much longer production prompt around those references, with a few differences: I wanted to use Magnific through MCP for the generative image/video side, plus a few extra repos and tools for motion graphics and JavaScript rendering.
Before actually asking Claude to build the video, I spent some time ping-ponging ideas with it and defining the storyboard until it felt close to what I had in my head. Then adapt my initial prompt with the results.
Then I just sent it.
About 4 hours later, Claude produced the first full version.
That first pass used roughly 80% of my Claude Max session and around $50 worth of Magnific credits.
After that I only did 2 correction passes: replaced a few scenes that felt too cliché, asked for more dynamic camera work and typography, and made some smaller adjustments. That used around 40% of my next Claude session.
And this is the result.
There’s obviously a lot more that could be polished if I wanted to keep going frame by frame, but I think it’s already enough to show how far you can get now with surprisingly little manual production.
No doubt this is the new standard of slop.
We’re going to see a lot of this over the next few weeks.
Full initial prompt:
I've included Louder than alarms.wav, which is the final song. It's called "Louder Than Alarms". The lyrics are in Lyrics.txt The song is basically about building the next generation of AI and this weird pursuit of something we're not even sure we understand. There's excitement in it, but also dread. We're building the thing, pushing it forward, almost compulsively, while somewhere underneath there's this feeling that maybe we're crossing into something we won't be able to fully explain or control. I want you to take the whole thing end to end and make the music video. Don't just give me a storyboard or a treatment. I want you to actually think through the song, its structure, the lyrics, the pacing, the visual language, generate what needs to be generated, build the motion system around it and assemble the final piece using the exact audio track I gave you. Final format is 16:9. The visual language is already defined. I'm calling it Weak Signal. Before you start designing anything, go through every one of the 20 images inside the refes/ folder of the AI-Slop project. Those are the ground truth. Don't treat them as loose inspiration and then drift into generic AI-video aesthetics. Study them properly and figure out what actually makes them feel related. The basic idea is a near-monochrome nocturnal world where the human being keeps getting reconstructed by the medium itself. Bodies don't just appear on camera, they turn into dithering, scanlines, ASCII, point clouds, data, compression errors. Faces are rarely clean faces. A person is usually tiny against something much bigger than them. Architecture feels empty, wet, brutalist, foggy. Sometimes the thing looking back at us is an eye. Sometimes it's text. Sometimes it's the system itself. A few of the references are especially important because they define families of images I want to keep returning to. Images 8, 10, 12, 13, 16 and 17 are probably the core: bodies reconstructed through 1-bit dithering, ASCII, line scans, smearing, point clouds, VHS/datamosh, pixel sorting. References 1, 6 and 11 are about system messages and text becoming part of the world: things like "What have you done?" filling a terminal, "YOU ARE DREAMING.", "YOU HAVEN'T COME THIS FAR FOR NOTHING." The text isn't just graphic design over the footage. It should feel like something inside the world is addressing you. Then there's the watching-eye family in 2, 7 and 20: enormous eyes, engraved eyes, corrupted eyes, molten eyes. Surveillance, awareness, something waking up. References 5 and 9 are useful for the architecture: grey brutalist spaces where there's only one saturated source of light. 14, 15, 18 and 19 are more like data cartography, topographic contours, coordinates, strange maps, orbital diagrams, macro imagery. I think those are really useful as inner landscapes and transitions. And 4 and 7 have that intervened-classical-art thing which I love because it introduces a little satire and absurdity without turning the video into comedy. Old paintings or engravings interrupted by some contemporary technological object or impossible intervention. Reference 3, the figure against the wall of code, is almost too literal. It's useful as a starting point, but I absolutely don't want this to become "Matrix code = AI". That's exactly the generic shit I'm trying to avoid. Keep the palette extremely restrained. Something like 80–90% crushed black, graphite, dark blue, basically #07080C / #0B1420 territory. Every shot can have one real saturated accent and only one: signal red, Klein blue, neon magenta, UV violet, phosphor green, molten orange. If there's iridescent rainbow/oil-slick color, save it for an actual rupture. I want it to mean something when it happens. should almost always come from one motivated source: a screen, a sign, a doorway, a window, something like that. Fog, wet surfaces, reflections, no generic fill light making everything readable. Humans should mostly be silhouettes, backs of heads, distant bodies or bodies reconstructed through whatever visual process is happening. If we do see a face clearly, something should be wrong with it: the eyes are replaced, the mouth is intervened, the image is breaking apart. Nothing should look clean. Grain, sensor noise, cracked surfaces, scanlines, crosshatching, compression, dirty blacks. I want every frame to feel like it physically exists on some damaged substrate instead of being pristine digital concept art. Composition-wise I want a lot of negative space. Low angles. Tiny people against huge systems. Centered or slightly offset figures. Macros with really shallow focus. Typography is terminal monospace, brutal signage, HUD microtext, coordinates, system output. Short phrases. Second person works incredibly well because it makes it feel like whatever is inside this world is talking directly to us. Please actively avoid the usual AI visual slop: colorful cyberpunk with seven neon colors in the same frame, glossy sci-fi hallways, clean plastic 3D skin, HDR, generic lens flares, warm cinematic sunset lighting, front-facing beautiful people staring into the camera, Pixar-ish 3D characters, literal Matrix code rain. If it looks like a Midjourney prompt people have seen 500 times, throw it away. thing I do want you to lean into is the fact that some of the references themselves have obvious AI artifacts. The fake almost-readable text, hands that become lines, surfaces that don't fully resolve, forms disintegrating into particles. Don't treat those as mistakes to hide. Treat them as part of the grammar. The difference is that they need to feel intentional. The viewer should feel like the image is losing the ability to remain an image. I think degradation can actually become the narrative arc of the whole video. Start in something closer to a real nocturnal photographic or painterly world and progressively lose it. Maybe we move through dither, ASCII, scanlines, point clouds, datamosh, whatever makes sense according to the actual structure of the song. Don't follow that order blindly — listen to the track and decide where each stage belongs. But I really like the idea that by the climax the representation itself is breaking. That's probably where the iridescent rupture belongs. And then after everything peaks, maybe the whole world collapses back to 1-bit monochrome or a terminal. Please actually analyze the song before doing this. Map the sections, energy changes, transitions, important lyrics, drops, moments where visual density should rise or disappear. I don't want a sequence of cool disconnected AI shots. Everything needs to feel directed. can go through the full /asic folder too. There's other work I've already done in there and I want you to understand the visual and technical context instead of solving this in isolation. Use the mesh skill to look at the reference compendium I've collected. Use the video scoring skill too, especially to understand how the existing JavaScript music/motion workflows work from references. I don't think the song itself needs to be modified, but I want you to understand those systems so you can use them creatively. There is also ElevenLabs available for sound design. The main music track stays exactly as it is, but if subtle sound design helps transitions or the opening/ending — electrical hum, tape noise, signal interference, glitches, machine texture — you can use it. Feel free to use the internet. In fact, do. Look at current motion-design work, JavaScript/WebGL experiments, title sequences, music videos, K-pop editing, kinetic typography, interactive graphics, generative art, whatever is actually useful. I'm not attached to a specific stack. I care about the output. For generated imagery/video, use the Magnific MCP tools. Before spending anything, check account_balance, then use simulate_cost before batches. The budget is 30k Magnific credits. Don't waste credits brute-forcing garbage, but also don't be so conservative that the piece suffers. Before generating, read Magnific's guides and inspect the actual models that are currently available: guides_list, guides_get, images_models_list, images_models_settings, video_models_list, video_models_show. Don't assume you already know which model is best. Test. For video, I care about believable body motion, physics and character consistency. If Seedance or another model supports audio conditioning/lip sync and it's genuinely the best option, use it. Follow the instruction fields the tools return. Wait for generations properly with creations_wait, and when chaining operations use the creation IDs rather than grabbing random web URLs. Before building scenes, make the visual system reusable. Upload the strongest reference images and create a Weak Signal style asset in the Magnific library. Generate some test frames with it and actually compare them to the reference board. If they don't belong next to those references, iterate. Don't start producing 30 scenes from a mediocre style lock. I want a written style sheet too, because once you figure out exactly what makes the look work we should be able to use that consistently throughout the video. Then design the protagonist. I don't need some over-designed sci-fi hero. The silhouette is much more important than costume details because this character needs to survive being reduced into dither, ASCII, scanlines and point clouds. Most of the time we'll probably see them from behind, as a silhouette, or reconstructed through the medium. Make them instantly readable even at low resolution. Save the character as a reusable library asset so they stay consistent. Supporting people can exist too, but don't turn this into dancers performing choreography. I'm thinking isolated figures, crowds behaving like data, people standing completely still, maybe engineers/operators becoming indistinguishable from terminal output. And I want the eye / watcher to recur enough that it almost becomes another character. the same with the environments. Build a few spaces that belong to the same world rather than making a completely new universe every eight seconds: foggy roads with a neon sign disappearing in the distance, brutalist apartment blocks and stairwells, terminal/server rooms, strange topographic landscapes, maybe a dark empty lab, intervened classical paintings. Save the important sets as location assets so we can return to them and the video gains continuity. From there, generate the video bases. Think of them like footage we're going to rotoscope later, not the final aesthetic. The important things at this stage are motion, camera, blocking, performance, silhouettes, physics and continuity. Pass the style, character and location references, use keyframes when they matter and include the actual lyric context for the individual scene. The protagonist doesn't need to visibly sing every word. This should behave like a real music video. Sometimes we see the character delivering a key line, sometimes they're just moving through the world, sometimes the entire shot is an insert with no human being at all. For the lyrics where seeing someone perform the line would actually make it hit harder, try to get the sync right. Cut the exact relevant section of the song and feed that audio to the video model if it supports audio conditioning. If the best model can't do that, use video_speak or whatever workflow gets us closest. Verify it. Extract frames, analyze clips, compare timing with the waveform/lyrics. If the mouth is obviously wrong or the timing feels fake, regenerate it. Don't just accept the first thing because technically it completed. Then assemble the base edit with video_cut / video_concatenate as needed. But this base edit is not the finished video. The part I care about most is what comes after it: I want to use the generated footage essentially like live-action footage for rotoscoping. Build the final visual treatment in JavaScript. Reconstruct the shots from the footage and render them through the Weak Signal language so that the final video only shows the JavaScript-rendered interpretation. We should never see the raw AI-generated footage. That's really the key idea behind the whole thing: the body is reconstructed by the medium. Use the underlying footage to give us coherent physics, character movement and cinematography, then destroy and rebuild it visually using 1-bit dithering, ASCII, line scanning, point clouds, pixel sorting, engraving/hatching, topographic contours, whatever the shot calls for. The final image can be Canvas, WebGL, Three.js, shaders, Remotion, GSAP, custom pixel processing — use whatever gets the strongest result. The camera language should stay pretty restrained. Slow push-ins. Locked shots. Long holds. Let the glitches and material transformations carry the energy. Cuts can land hard on beats, but I don't want hyperactive random editing for the sake of retention. Frame rate can become another narrative device. Maybe the earlier reality is smoother around 24 fps, then degraded sections begin stepping at 12 fps or 8 fps. A neon sign shouldn't simply appear, it should struggle to ignite. Terminal text types, fails, repeats and scrolls. A body can dissolve into contour lines and those contour lines can become the next landscape. Pixel sorting can literally pull one scene into another. Point clouds can scatter and reform somewhere else. I also want the lyrics on screen throughout the video, but not as one repetitive lyric-video template. This style gives us a lot of ways to embed text into the world. Sometimes a lyric is a neon sign disappearing into fog. Sometimes it's terminal output repeated until it becomes texture. Sometimes it's a coordinate/HUD annotation floating over a topographic map. Sometimes it's huge brutal typography dominating the entire frame. And sometimes it's just a small subtitle because the image needs room to breathe. Use composition to plan for the text from the beginning. Don't generate a centered character filling the frame and then later realize there's nowhere to put the words. If a lyric needs to dominate, put the person small on the right and let the type own the left half of the image. Or invert it. Keep changing the hierarchy. The beginning especially needs to grab people fast. The lyrics should probably be more aggressive and visually present there. I want somebody scrolling Twitter to stop within the first seconds because they immediately understand that this isn't a random AI montage.There's also a zeitgeist layer I want in the piece because the target audience is basically San Francisco tech Twitter / the people currently living inside the AI acceleration discourse. Audit what's happening right now. Look at the recurring memes, anxieties and events people will recognize: models suddenly eating another supposedly-hard domain, math/physics hype, Navier–Stokes chatter if that's still relevant, the Shinji meme and all the language around it, benchmark obsession, agent swarms, recursive self-improvement discourse, the flood of AI slop itself, whatever the actual timeline is talking about right now. Don't turn it into a slideshow of tweets. Use those things as cultural texture. You can take screenshots/assets or recreate recognizable internet fragments in an internet-brutalist way, then process them through the same monochrome / one-accent / grain / degradation system so they feel like artifacts inside this universe. If a cultural reference breaks the visual world, don't use it. Study good music videos and motion pieces too. K-pop is useful here because the best videos are insanely good at directing attention, constantly refreshing visual patterns and knowing exactly when to introduce a new thing. I don't want the K-pop aesthetic. I want the psychology and directing discipline. The biggest thing I want you to avoid is assembling first and thinking later. That's when this kind of project turns into a bunch of individually cool shots that somehow feel awful together. Plan the compositions. Plan where the lyrics live. Plan how one scene transforms into the next. Establish recurring motifs. Give visual ideas enough time to register before replacing them. And iterate. Once the whole thing exists, watch it from beginning to end multiple times. Pull frames from different points and put them next to the original Weak Signal reference board. Ask yourself if every shot genuinely belongs there. Look for places where it turns into generic AI art, where the pacing dies, where the text is fighting the image, where the protagonist changes too much, where a transition feels arbitrary, where the degradation curve stops making sense. Fix those places. Regenerate when necessary. Rewrite the JavaScript effects when necessary. The first complete render isn't the deliverable just because it's complete. I've got Claude Max with basically full usage available and the Magnific budget is there to make the piece good. Use resources intelligently, but don't act artificially constrained. Think deeply before expensive generations, reuse assets where it helps continuity, test cheap before scaling, but push when something is worth pushing. I genuinely think the tools are capable enough now that this doesn't need to look like an "AI music video". That's the bar I'm interested in. I want it to feel like somebody had a very specific visual idea, directed it obsessively, and happened to use generative systems and code to make something that would have been almost impossible to produce any other way. The stretch goal is something where people stop asking what model made it and start asking how the fuck it was made.A few repos that may be useful as references for the animation/rendering side: https://github.com/JohnHeibel/ClaudeAnimationBase
https://github.com/iart-ai/webgl-animation-skills
https://github.com/Fats403/remotion-gsap Use them if they're useful, steal ideas from the workflows, or ignore them if you find something better. I'm not married to any of those implementations. I'm married to the quality of the final video.
