Ai Music Video Took 34 Hours And Many Rejected Takes
The column's own account suggests AI video tools now handle complex staging only when a person steers them, which puts the near-term pressure on directing skill rather than on the models alone.
Reporting from 1 source: ASCII.jp.
A column from ASCII.jp describes an attempt to build a full music video with GPT-6 Astra, released September 3, and the local video model MiniMax H3. The song, 'Design and Will,' came from Suno v6, which arrived September 9. The author reports 34 hours of production, repeated rejections of generated output, and a conclusion that a human still has to direct the AI closely to keep it on course.
The author of the ASCII.jp column says the song came first. 'Design and Will' was built with Suno v6, the music model released on September 9, and the lyrics were produced by having GPT-6 Astra summarize a conversation into verse. From there the column describes trying to extend that into a full music video, using MiniMax H3 for video generation in a local environment.
The stated result is 34 hours of work and repeated rejections, with output that sometimes drifted from the intended direction until the author stepped in to direct it. The column's own reading is that human control is what makes the complex staging possible, and that AI may push media further toward personalization.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.