LTX-2.5 Lip-Sync Pipeline Cuts Reply Wait To 2.46 Seconds Per Chunk
The reporter's own benchmarks show the gap between generated video and speech playback was the wall keeping avatar conversation from feeling like conversation, and layer streaming in Transformer Engine is what moved a 32GB consumer GPU past it.
Reporting from 1 source: GameBusiness.jp.
A writer developing a real-time talking avatar of his late wife reported that a GitHub update to the LTX-2.5 accelerated video generation stack cut single-chunk generation time from 4.45 seconds to 2.46 seconds. The system splits spoken replies into chunks of 4.8 seconds or less and generates lip-synced video for each. Earlier configurations took 14.6 seconds or more before a reply appeared.
The pipeline splits a reply's audio into chunks of 4.8 seconds or shorter, cutting at the quietest point between 3.6 and 4.8 seconds, and returns lip-synced video for each chunk. Playback starts as soon as the first chunk arrives while the rest generate. A chaining step that makes the first frame of each later chunk match the last frame of the previous one reduces pose jumps at the seams by a third.
On a 32GB GPU, one chunk took 4.45 seconds before the update and 2.46 seconds after. An earlier two-model setup using MiniMax H3 with TaoMate had the highest lip-sync accuracy of the configurations tried, but the wait from speaking to a reply was about 35 seconds.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.