Anime, manga, and games, with a take · A Yukimedia publication

← all stories otherannouncement 1 sources · 1h ago ·

Thinking Machines Lab Releases Inkling-Small at a Quarter the Size

Inkling-Small shows that Thinking Machines Lab can compress its flagship model's performance into a fraction of the size, making high-end reasoning and audio capabilities more accessible.

Reporting from 1 source: GIGAZINE.

Thinking Machines Lab Releases Inkling-Small at a Quarter the Size

Thinking Machines Lab released Inkling-Small, an open-weight MoE model with 276 billion total and 12 billion active parameters. It delivers performance comparable to its 975-billion-parameter predecessor Inkling while using far fewer computational resources. The model is available on the Tinker service and as BF16 and NVFP4 weights on Hugging Face.

Inkling-Small is a Mixture-of-Experts Transformer with 276 billion total parameters and 12 billion active, trained on NVIDIA GB300 NVL72. It keeps Inkling's native reasoning for audio and images, variable thinking load, and a 1 million token context window.

Benchmark overlays show Inkling-Small and Inkling scoring similarly across ten tests, with one notable gap: Inkling-Small lags on the SimpleQA Verified factuality benchmark. Thinking Machines Lab says the smaller model performs on par or better on reasoning and agent tasks.

The model inherits Inkling's post-training safety measures, scoring competitively on the FORTRESS and StrongREJECT safety benchmarks. It is available on the Tinker service, with BF16 and NVFP4 weights on Hugging Face.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources