SoulX-Singer Zero-Shot Vocal Synth Built With Claude Code
The piece documents a workflow that integrates composition algorithms with direct singing synthesis, reducing reliance on external generation steps.
Reporting from 1 source: GameBusiness.jp.
A developer building a personal vocal synthesizer reports progress on SoulX-Singer, a zero-shot singing synthesis tool that can sing from tens of seconds of vocal data. The project aims to let a composition algorithm generate MIDI and lyrics that directly drive a singing voice model, bypassing Suno and RVC.
The developer behind the LipSync Avatar project has been working on SoulX-Singer, a zero-shot vocal synthesizer that can reproduce a voice from just tens of seconds of singing data. The tool is being developed with Claude Code, an AI coding assistant.
The current pipeline still relies on RVC for voice conversion, but the goal is to have the composition algorithm output MIDI and lyrics that go straight into a singing synthesis model. That would let the developer hear changes to chords, melody, or lyrics immediately, rather than waiting for Suno to generate a full track.
OpenUtau is being evaluated as an entry point, since it can load MIDI and lyrics and supports neural singing models like DiffSinger.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.