Google introduced EmbeddingGemma 2, a multimodal embedding model with 740 million parameters designed for on-device inference. The new version expands beyond text to include code, images, video, and audio in a shared embedding space. It is released under the Apache 2.0 license and aims to enhance privacy-first search and retrieval applications directly on consumer hardware, according to blog.google.
The original EmbeddingGemma, launched last year, achieved over 20 million downloads as developers used it to build smarter on-device search tools and retrieval augmented generation (RAG) pipelines. EmbeddingGemma 2 builds on the Gemma 4 architecture to unify multiple data types in a single model, enabling tasks like finding a video clip from a voice memo or searching hours of audio recordings with a text query, all processed locally on devices.
EmbeddingGemma 2 addresses growing demand for privacy-conscious AI models that operate without cloud dependency. Its multimodal capabilities distinguish it from earlier text-only embedding models, offering developers a versatile tool for organizing and connecting diverse information types. The Apache 2.0 license facilitates commercial and open-source use, positioning EmbeddingGemma 2 as a competitive option in the expanding market for on-device AI and retrieval systems.
EmbeddingGemma 2 is now available for developers to integrate into applications, supporting enhanced multimodal search and retrieval on consumer devices while maintaining user privacy, per blog.google.