New-ZZZ
RU / EN
LLM 30 September 2026

Perplexity AI improves contextual embeddings for better information retrieval

N
New-ZZZ desk
X @perplexity_ai · 2 hours ago

Perplexity AI introduced a new approach to training contextual embeddings—models that encode not just individual text fragments, but the entire document. This solves a fundamental problem in retrieval systems, which traditionally truncate long documents into small chunks, thereby losing surrounding context.

Instead, the new model first processes the entire document and then aggregates the fragment vectors. For training, it uses an innovative method that extracts relevance by evaluating every document token relative to the search query, which allows it to bypass the limitations of older methods based on "golden" chunks.

The model pplx-embed-v2-context-9b-preview demonstrated high performance on the new context-bench benchmark, leading in answer and evidence extraction tasks. Furthermore, it is significantly more efficient than its competitors, using 8 times less memory to store vectors.

Why it matters

  • —This is critically important for RAG (Retrieval-Augmented Generation) systems as it improves the quality of retrieved context, making answers more accurate and well-founded.
  • —The new approach, which encodes the entire document, solves the problem of context loss, which was the main weakness of previous systems.
  • —High efficiency and low resource consumption (8 times less memory) make the model attractive for commercial implementation.

Key facts

  • The pplx-embed-v2-context-9b-preview model encodes the entire document, not just its fragments.
  • Training is based on assessing the relevance of each token relative to the query.
  • The model leads on the ConTEB benchmark for answer and evidence retrieval tasks.
  • The model uses 8 times less memory (1 KB) compared to competitors (8 KB).
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related