Perplexity AI releases multimodal embedding models for search
Perplexity AI has introduced two new embedding models—pplx-embed-v2-late. These are advanced late-interaction models that allow for the extraction and indexing of not only text but also images and entire pages, using a single cross-modal space.
Traditional dense embeddings often lose details when working with long or visually rich documents. The new models solve this problem by preserving a vector representation for every token (128 dimensions) and enabling information retrieval from PDFs, slides, and scans without the need for OCR, thus preserving the structure, tables, and figures.
The models achieved high performance across various tasks: on the Q2D-Web benchmark, the 9B model showed a Recall@1000 of 74.8%, and in the MADQA task (PDF search), it achieved 92.4%, surpassing competitors. The availability in two sizes (9B and 0.6B) and high performance make these models a powerful tool for RAG systems and search agents.
Why it matters
- —Solves the problem of losing details when indexing complex documents (PDFs, images).
- —Enables cross-modal search by combining text and visual content in a single space.
- —High performance in RAG and agentic search tasks, surpassing competitors.
Key facts
- Models are capable of extracting text, images, and pages using a unified embedding.
- Instead of one vector per document, a vector is used for every token (128 dimensions).
- On the Q2D-Web benchmark, the 9B model achieved 74.8% Recall@1000.
- In the MADQA task (PDF search), the 9B model showed 92.4% accuracy.
The full text is in the original source. Here we provide a brief summary and key facts.