Google DeepMind released a multimodal embedding model for devices
Google DeepMind introduced EmbeddingGemma 2 — the first natively multimodal open model for creating embeddings, designed to run directly on the device.
The model significantly expands the capabilities of vector search by combining not only text, but also code, images, audio, and video into a single mathematical space. This allows for the creation of much more complex and accurate search systems.
EmbeddingGemma 2 has only 740 million parameters, making it extremely compact and efficient for devices. At the same time, it demonstrates high competitiveness, outperforming some specialized models in benchmarks that are significantly larger.
Developers can use this model to add multimodal search to their applications, for example, for searching specific moments in video or audio materials.
Why it matters
- —Enables multimodal search that works with different data types (video, audio, code).
- —Optimized for on-device operation, which increases speed and privacy.
- —Combines high performance with a compact size (740 million parameters), representing a technical breakthrough.
Key facts
- It is the first natively multimodal open model for embeddings.
- Supports encoding, images, audio, and video in a single space.
- The model has 740 million parameters.
- Optimized for deployment on edge devices.
The full text is in the original source. Here we provide a brief summary and key facts.