Wednesday, October 7, 2026
Tech Beat
Oct 6, 2026, 7:57 PMArtificial Intelligence

Google EmbeddingGemma 2 Brings Multimodal AI Search On Device

Google’s 740M-parameter EmbeddingGemma 2 unifies text, code, images, audio and video for private, offline search and RAG on consumer devices with Apache 2.0.

Listen to this briefingAudio briefing

Summary

On October 6, 2026, Google DeepMind engineers Sahil Dua and Henrique Schechter Vera launched EmbeddingGemma 2, a commercially permissive Apache 2.0 model with 740 million parameters, built on Gemma 4 and Gemini Embedding technology. It maps text, code, images, video and audio into one space for on-device search, retrieval, routing and RAG. Its text-only predecessor exceeded 20 million downloads.

Google says it leads sub-1B multimodal embedders on MTEB Code and MAEB, matches EmbeddingGemma’s multilingual text quality, and beats some specialist models over twice its size. MTEB Code rose 9.92 points, from 68.76 to 78.68. The modular design uses 270M parameters for text, plus optional 170M vision and 300M audio encoders. MRL shortens 768-dimension vectors to 512, 256 or 128, cutting local vector database storage and memory by up to 6x. Quantized Pixel 11 Pro use is about 191MB active RAM for text weights and 567MB for the full model. Its 8K context, 4x EmbeddingGemma 1, handles 5.5 minutes of audio, 29 images, 58 video frames or mixed inputs.

Local generation keeps data private, lowers latency and works offline. Pairing it with Gemma 4 creates multimodal on-device RAG with lower combined memory because both share a tokenizer and audio encoder. Google AI Edge Gallery offers Instant Media Search and Video Moments Finder; AI Edge Foresight combines retrieval with Gemma 4 reasoning; MediaPipe Decision Task API supports real-time classification, routing and prediction. Weights are on Hugging Face and Kaggle, with LiteRT Community builds and Gemini Enterprise Agent Platform Model Garden access coming soon. Deployment supports Google AI Edge MediaPipe, LiteRT, transformers.js, WebGPU, transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio; Qdrant stores vectors, while Unsloth provides fine-tuning guidance.

Positives

  • MTEB Code performance increased 9.92 points, from 68.76 to 78.68, while multilingual text quality matched the original EmbeddingGemma.
  • MRL reduces 768-dimension vectors to 512, 256 or 128, lowering local database storage and memory by up to 6x.
  • 191MB of active RAM runs quantized text weights on a Pixel 11 Pro, supporting resource-constrained on-device applications.
  • An 8K context window processes 5.5 minutes of audio, 29 images, 58 video frames or interleaved inputs locally.
  • Apache 2.0 licensing and support across major deployment frameworks make commercial adoption and integration easier.
  • More than 20 million EmbeddingGemma downloads demonstrate substantial developer demand for lightweight, privacy-focused retrieval models.

Risks & concerns

  • 567MB of active RAM is still required for the quantized full multimodal model on a Pixel 11 Pro.
  • The 8K context limits each pass to 5.5 audio minutes, 29 images or 58 video frames.
  • Full multimodal support requires adding 170M vision and 300M audio encoders to the 270M text configuration.
  • Gemini Enterprise Agent Platform Model Garden availability is still pending, with no specific release date provided.
Primary sourceGoogle DeepMind Newshttps://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceOct 6

OpenAI Makes ChatGPT textGrain Watermarks Default in EU

Artificial IntelligenceOct 6

Underdog Launches Private On-Device AI With Stripe Fee Model

Artificial IntelligenceOct 6

Musubi Launches Open-Weight PolicyLM-1.7B for Real-Time AI Moderation