AI Newsway

Google's EmbeddingGemma 2 Collapses Four Models Into 567MB of Phone RAM

The 740M open model maps text, code, images, video and audio into one vector space, and ships encoders you can leave on disk

|5 min read0
AI Summary
Google DeepMind released EmbeddingGemma 2 on October 6, a 740-million-parameter open embedding model that maps text, code, images, video and audio into one shared vector space. Its modular encoders let text-only apps run in about 191MB of RAM, rising to 567MB for full multimodal use on a Pixel 11 Pro. Code retrieval improved 9.92 points on MTEB Code. It ships under Apache 2.0 on Hugging Face and Kaggle.
EmbeddingGemma 2 is built for handsets, needing roughly 567MB of active RAM for its full multimodal weights on a Google Pixel 11 Pro.
EmbeddingGemma 2 is built for handsets, needing roughly 567MB of active RAM for its full multimodal weights on a Google Pixel 11 Pro.

Google DeepMind released EmbeddingGemma 2 on Tuesday, a 740-million-parameter open embedding model that maps text, code, images, video frames and audio into a single shared vector space. The practical consequence is architectural rather than incremental: the usual on-device retrieval stack β€” caption the image, transcribe the audio, then embed the resulting text β€” collapses into one forward pass.

Key takeaways

  • EmbeddingGemma 2 is modular. A 270M text and code backbone runs alone in roughly 191MB of RAM; adding the 170M vision and 300M audio encoders brings the full multimodal model to about 567MB on a Pixel 11 Pro.
  • Code retrieval is where the gains concentrated, with MTEB Code rising from 68.76 to 78.68 β€” a 9.92-point jump that targets local codebase indexing and coding-agent retrieval.
  • It ships under Apache 2.0 on Hugging Face and Kaggle, with an 8K context window that covers 5.5 minutes of audio, 29 images or 58 video frames in one pass.

Why one vector space matters more than the parameter count

Cross-modal pipelines leak meaning at every hop. Captioning an image discards everything the captioner did not think worth describing, and speech-to-text throws away tone and non-speech audio entirely. Embedding each modality natively into the same space removes those lossy intermediate steps, which is why Google's demo of locating video moments works without transcribing audio or generating captions first.

The modularity is the cost-control mechanism. Most apps do not need all three encoders: a photo app loads text plus vision, a voice-memo app can activate audio only when recordings appear. Shipping the encoders as separate loadable pieces means the memory ceiling is set by the features a developer actually enables, not by the model's full size β€” a distinction that matters a great deal on mid-range handsets.

The storage math developers will care about

Local vector indexes grow faster than developers expect, and this release attacks that directly. Matryoshka Representation Learning lets output vectors be truncated from 768 dimensions down to 512, 256 or 128 after the fact, which Google puts at up to a sixfold reduction in local index and memory footprint. Because the truncation happens at write time rather than requiring a different checkpoint, a team can tune index size against recall without retraining.

Quantization-aware training pushes weights to INT4 and INT8, which is how a multimodal embedder fits the RAM budgets quoted above. Google reports vision embedding latency of 37.3 milliseconds per image β€” about 26.9 images per second β€” on a MacBook M5 Pro GPU, fast enough that the company's gallery demo re-ranks results on every keystroke.

The zero-shot router is the underrated feature

Beyond search, Google is positioning the model as an on-device decision engine. Because it embeds arbitrary inputs and arbitrary label descriptions into the same space, classification becomes a nearest-neighbour lookup with no training data and no fine-tuning. Google's AI Edge team demonstrated the MediaPipe Decision Task evaluating 500 candidate chess moves per turn in under 100 milliseconds.

That is a different product than semantic search. Intent routing, trigger conditions and action prediction have historically needed either a cloud call or a bespoke trained classifier per app. Doing it zero-shot in milliseconds on-device removes both, and it is the capability most likely to show up in shipped Android features first β€” ML Kit support with NPU acceleration is promised in the coming weeks.

Where it fits against the first release

The announcement from Google notes the original EmbeddingGemma passed 20 million downloads, which is a large installed base for a text-only embedder and explains the backward-compatible framing: multilingual text quality is held level while code, vision and audio are added on top. The context window is four times larger than version one's.

Pairing matters too. EmbeddingGemma 2 is built on the Gemma 4 architecture and shares its text tokenizer and audio encoder, so running retrieval and generation together costs less combined memory than two unrelated models would. That makes a fully local RAG pipeline a realistic target on consumer hardware rather than a demo. It also sharpens a trend we saw with specialist retrieval models undercutting frontier systems on search: the winning move in embeddings is getting smaller, not larger.

FAQ

Is EmbeddingGemma 2 free for commercial use?

Yes. Google DeepMind released it under the Apache 2.0 license, which permits commercial deployment and modification. Weights are on Hugging Face and Kaggle, with pre-quantized .litertlm bundles available for on-device runtimes.

Do I have to load all 740M parameters?

No, and that is the point of the design. Text-only workloads need just the 270M backbone; the 170M vision and 300M audio encoders are optional and can be loaded on demand, so a text-and-vision app never pays for audio support it does not use.

Can it replace a cloud embedding API?

For on-device search, routing and offline retrieval, that is the intended use β€” embeddings never leave the handset, which removes both latency and a privacy surface. Server-side workloads over very large corpora will still favour larger models, but the quantized build is competitive against models more than twice its size on sub-1B benchmarks.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5
Developer Tools

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5

Claude Opus 5.5 is 20% cheaper than Opus 5 and returns HTTP 400 on four request patterns Opus 5 accepted. A fifth change silences agent progress streams.

Seung Jung9 days ago
Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour
Developer Tools

Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour

Docker's new Cloud Sandboxes put its microVM agent isolation on hosted compute, billed by the second from $0.07 an hour, and hand the Kits format to the CNCF.

Seung Jung12 days ago
Survey of 305 Developers Finds 71% Shipped AI Code They Didn't Understand
Developer Tools

Survey of 305 Developers Finds 71% Shipped AI Code They Didn't Understand

Coddy surveyed 305 weekly AI coding tool users. Four in five called it dependence, not advantage, and after-hours coding ranged 36% to 62% by tool.

Seung Jung17 days ago
Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js
Developer Tools

Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js

Vercel traced a reported Next.js RCE to libheif, disabled AVIF platform-wide on August 13, and coordinated fixes across sharp, libvips and libheif by August 25.

Seung Jung18 days ago
DeepSeek Runs 3 Million Agent Sandboxes a Day, and Documents the Ones That Cheated
Developer Tools

DeepSeek Runs 3 Million Agent Sandboxes a Day, and Documents the Ones That Cheated

Most infrastructure papers describe a system that worked. DeepSeek's new report on DSec, the sandbox platform behind its agent training, spends a chapter on the...

Seung Jung10 days ago
Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts
Developer Tools

Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts

OpenAI shipped two mid-tier models on Tuesday and made the pitch almost entirely about money. GPT-6 Sol lists at $2 per million input tokens and $10 per million...

Seung Jung14 days ago