Google Ships Gemini Embedding 2, a Natively Multimodal Embedding Model

Google's new embedding model handles text, images, and video natively — a meaningful step toward multimodal search and retrieval that doesn't require separate pipelines for each modality.

Google AI introduced Gemini Embedding 2, describing it as a "natively multimodal embedding model" capable of powering "video analysis tools, visual shopping assistants," and similar applications. The key word is "natively" — previous multimodal embedding approaches typically required separate encoders for each modality, stitched together at the vector level. A natively multimodal model produces embeddings in a shared space from the start.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.