Writer Builds Personal DiffSinger Vocal Synth From His Own Voice
The piece shows how open-source AI vocal synthesis now lets individuals create custom singing voices from scratch, bypassing the cost and data demands of commercial vocaloids.
Reporting from 1 source: GameBusiness.jp.
A writer for CloseBox and TechnoEdge documented building a personal DiffSinger vocal synthesis model, recording about an hour of his own singing over two days, then training it on an NVIDIA DGX Spark. The goal is to convert his recordings with an RVC voice changer to recreate his late wife's voice for new songs.
The writer's previous experiment used SoulX-Singer, a zero-shot synthesizer, to make his wife's voice sing in Japanese. That required only 10 to 20 seconds of data but remained unstable for full songs. DiffSinger, by contrast, demands about an hour of singing with each phoneme recorded at least 20 times, so he recorded his own voice over two days, then planned to convert it with an existing RVC model of his wife's voice.
He trained the model on a DGX Spark, reaching a preliminary checkpoint after several hours. The approach sidesteps the need for new recordings of his wife, whose singing exists only in a few tracks and an RVC model. The project builds on his 2013 UTAU-Synth voicebank, which used just three songs and took seven years to produce over 100 tracks.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.