Google Releases EmbeddingGemma 2 With 8K Context And 740M Parameters
The pitch is local: a multimodal embedder small enough to hold its working memory to roughly 191MB for text on a phone is aimed at on-device search and retrieval-augmented generation, where the vector index sits on the hardware instead of a server.
Reporting from 1 source: GIGAZINE.
Google announced EmbeddingGemma 2, an embedding model that maps text, code, images, video, and audio into one vector space. It has 740 million parameters, an 8K-token context window, and output vectors that developers can shorten from 768 dimensions to 512, 256, or 128. Google says text-only weights need about 191MB of active RAM on a Pixel 11 Pro.
Google says the model runs locally on mobile and desktop hardware, not only in a data center. On a Pixel 11 Pro, the text-only weights need a minimum of about 191MB of active RAM, and the full multimodal version about 567MB. For text-only work the model runs at 270 million parameters; adding a 170 million parameter vision encoder and a 300 million parameter audio encoder gives full multimodal support.
It is built on the Gemma 4 architecture and released under the Apache 2.0 license, which allows commercial use. Weights are available on Hugging Face and Kaggle, with the Gemini Enterprise Agent Platform Model Garden listed as a coming home.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.