Google DeepMind has released EmbeddingGemma 2, a multimodal embedding model designed to run directly on consumer devices.

According to Google, the model enables users and developers to perform semantic search across text, image, video, and audio content, without needing an internet connection or external cloud processing.

Access deeper industry intelligence

Experience unmatched clarity with a single platform that combines unique data, AI, and human expertise.

Find out more

The newly launched EmbeddingGemma 2 is an open-weight model with 740 million parameters, released under an Apache 2.0 licence.

It can operate on a range of devices, with memory requirements of around 191MB for text-only usage and approximately 567MB for full multimodal performance on devices such as the Google Pixel 11 Pro.

EmbeddingGemma 2 covers text, vision and audio modalities using modular encoders, allowing users to only load the components required for each use case.

Google says the model delivers on-device vector search at low latency, enabling instant decision-making and zero-shot classification without fine-tuning or training data.

The Alphabet subsidiary announced the inclusion of EmbeddingGemma 2 in several showcase applications.

The Google AI Edge Gallery app has been updated with the “Instant Media Search” and “Video Moments Finder” features. These allow users to search local media using natural language or sample images, and to identify key moments in videos with descriptive text prompts.

The app processes queries and content on the device, storing embeddings in a local database and updating search results live as input is entered.

A separate application, Google AI Edge Foresight for Mac, integrates fully local note-taking and cross-modal retrieval capabilities.

Using EmbeddingGemma 2 and Gemma 4 models, Foresight indexes system audio, transcripts, and files, enabling offline, private search and retrieval of details within recorded meetings or documents.

According to Google, this approach protects user privacy, as data remains on the device.

The company has outlined broader support for developers. Google plans to offer EmbeddingGemma 2 as a service for Android via ML Kit, supporting hardware acceleration across various devices.

Integration with MediaPipe Tasks and LiteRT will allow cross-platform deployment for iOS, Windows, macOS, Linux and web applications, providing tools for embedding and semantic retrieval as well as optimised inference.

Google describes EmbeddingGemma 2 as a model for on-device multimodal embeddings, designed to map combinations of text, audio, images, and video into a single unified embedding space.

Technical notes highlight quantisation-aware training, modularity, and memory efficiency as key features enabling the model to run on memory-constrained hardware.

EmbeddingGemma 2 builds on the previous version’s text embedding functions, adding support for code, images, video and audio.

The model, based on the Gemma 4 architecture, offers an extended context window, up to 8,000 tokens, and provides dynamic vector size adjustment, reducing storage and memory requirements.

According to Google, EmbeddingGemma 2 achieved improved results in benchmarks such as the Massive Text Embedding Benchmark Code and Massive Audio Embedding Benchmark. The earlier version of the model has been downloaded more than 20 million times by developers.

Model weights and further deployment tools are available through platforms including Hugging Face and Kaggle.