Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 59m ago ·

AMD Releases Instella-MoE AI Model Trained on Its Own GPUs

AMD's release of Instella-MoE, trained entirely on its own hardware and software, shows the company's push to establish a fully open AI development stack that competes with comparable small models.

Reporting from 1 source: GIGAZINE.

AMD Releases Instella-MoE AI Model Trained on Its Own GPUs

AMD released Instella-MoE, a Mixture-of-Experts language model trained end-to-end on its Instinct MI300X and MI325X GPUs. The model has 16 billion total parameters and 2.8 billion active parameters. Six versions are available for free, and the training code and framework have been made public.

AMD has published Instella-MoE, a Mixture-of-Experts language model built and trained using its own Instinct MI300X and MI325X GPUs. The model activates only a subset of its 16 billion total parameters during inference, with 2.8 billion active.

Six versions are available, including a pretrained model on 7.1 trillion tokens, a mid-trained variant with enhanced math and coding skills, a base model with a 64k token context window, and fine-tuned versions using supervised fine-tuning, direct preference optimization, and reinforcement learning.

The training framework Primus and the MoE architecture FarSkip-Collective were used, and the codebase is public on GitHub. AMD's performance charts show the Think version scoring higher with fewer active parameters than Gemma-4-E4B-it.

  • Instella-MoE-16B-A3B-Pretrain: pretrained on 7.1 trillion tokens
  • Instella-MoE-16B-A3B-Midtrain: enhanced math, coding, and reasoning
  • Instella-MoE-16B-A3B-Base: context window extended to 64k tokens
  • Instella-MoE-16B-A3B-SFT: supervised fine-tuning applied
  • Instella-MoE-16B-A3B-DPO: direct preference optimization applied
  • Instella-MoE-16B-A3B-Think: reinforcement learning applied

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources