Mouse Computer Shows 96GB VRAM Allocation For Local LLMs At TGS2026
The pitch treats memory allocation, not compute, as the practical gate on running the 70B to 80B class models developers want for long-context work, which is why a machine that turns main memory into VRAM is the demo rather than the software.
Reporting from 1 source: ASCII.jp.
At Tokyo Game Show 2026, Kiyoshi Shin of Barin Studio and Mouse Computer product manager Nami Hayashida presented a talk on running local LLMs in game development. The demo machine was Mouse Computer's DAIV CX, built on an AMD Ryzen AI Max+ processor with Radeon 8060S graphics, which can allocate up to 96GB of its 128GB unified memory as VRAM. Shin cited leak risk and API costs as reasons to run models locally.
The argument on stage was narrower than a case for AI in general. Shin framed the choice between cloud and local as one of control: unpublished planning documents, setting bibles, character backstories and source code cannot be sent to an outside server without legal and security review, and pay-per-use API billing plus rate limits discourage the repeated prompting that development work involves.
He conceded that a single PC runs slower than a data center, and pointed to Ollama and LM Studio as tools that make open models easy to run. The hardware on the booth is what he put behind the claim. Mouse Computer's DAIV CX allocates up to 96GB of its 128GB unified memory as VRAM; typical graphics cards hold 16GB to 24GB. That ceiling decides how large a model and how long a context can be held in memory.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.