Google Ships Gemini Embedding 2, a Natively Multimodal Embedding Model
Google's new embedding model handles text, images, and video natively — a meaningful step toward multimodal search and retrieval that doesn't require separate pipelines for each modality.
Google AI introduced Gemini Embedding 2, describing it as a "natively multimodal embedding model" capable of powering "video analysis tools, visual shopping assistants," and similar applications. The key word is "natively" — previous multimodal embedding approaches typically required separate encoders for each modality, stitched together at the vector level. A natively multimodal model produces embeddings in a shared space from the start.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.