OpenTPU Project Runs Language Models On Kintex-7 FPGA Card
The stated aim is not to reproduce Google's TPU but to test how far AI agents can carry hardware design, including the design of the chips that run them, and the whole stack is published for inspection.
Reporting from 1 source: GIGAZINE.
An open-source project called openTPU has been released, built around a simple accelerator design that pairs a sequencer, DMA, matrix and vector units, and a quantization block. The developer wrote the circuits onto an FPGA card using AMD's Kintex-7 and ran LFM2.5-230M, Qwen3-0.6B, Qwen3.5-0.8B, Gemma 4, and Phi-4-mini, including the 34.7-billion-parameter Qwen3.5-35B-A3B. It ships under the Apache License 2.0.
The repository bundles the RTL, instruction set architecture, simulator, compiler, and profiler in one place, which is what makes the design inspectable rather than a black box. Its structure is deliberately plain: a sequencer issuing instructions in order, a DMA unit moving data to and from memory, a matrix operation unit, a vector unit for floating-point work, and a quantization stage. There are no caches and no complex instruction scheduling of the kind general CPUs and GPUs carry, and data movement is written out as instructions, so it stays possible to trace which process consumed how many clocks. The developer cites memory read speed for model data, not the arithmetic itself, as the current constraint.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.