Google releases open model for on-device multimodal search
Google has released EmbeddingGemma 2, an open model designed to let software search text, code, images, audio and video together on consumer devices rather than relying on cloud processing. Google says the model has 740 million parameters and is available under the Apache 2.0 license, which permits developers to use and adapt it under the license’s terms. The model converts different kinds of content into a shared embedding space, allowing search systems to retrieve material based on meaning. Google says the design supports local multimodal search and retrieval, including searches across personal media and files. The model’s weights are available on Hugging Face and Kaggle. EmbeddingGemma 2 is modular, with a 270-million-parameter text component and optional vision and audio encoders of 170 million and 300 million parameters, respectively. Google reports that a quantized version uses about 191 megabytes of active RAM for text-only processing and about 567 megabytes for the full multimodal model on a Pixel 11 Pro. Those figures are Google’s measurements; the supplied reporting does not provide independent tests of the model’s performance or memory use.