Perplexity open-sources context-aware embedding model for long-document RAG

Perplexity open-sources context-aware embedding model for long-document RAG

N
News Editor
2026-10-01 14:06:32
Perplexity has released a new open-source embedding model, pplx-embed-v2-context-9b-preview, aimed at retrieval-augmented generation workflows. The model changes how long documents are encoded: instead of splitting a document into chunks and embedding each one in isolation, it encodes chunks with access to the broader document context. That design is meant to help with references that are unclear on their own, such as a sentence that says only "it grew 20%" without naming what "it" refers to. The model is available on Hugging Face under the MIT license. Perplexity said the training setup also taught the model to retrieve not only answer-bearing passages but also supporting passages that can be used to verify those answers. In a blind evaluation on turbopuffer's context-bench, the model posted an Answer@10 score of 45.5%, 14.4 percentage points above Voyage Context 4, and an Evidence Recall@10 score of 40.6%, up 5 percentage points. The benchmark covered 38,894 long documents and 2,099 queries, with questions kept private to reduce the risk of test-set contamination. Perplexity said the release remains a preview, and warned that weights, APIs, and generated embeddings may still change, meaning vectors created now may not be directly compatible with future versions.

Perplexity has open-sourced a new context-aware embedding model, pplx-embed-v2-context-9b-preview, built for retrieval-augmented generation, or RAG.

Standard embedding systems usually handle long documents by splitting them into smaller chunks and encoding each chunk on its own. Perplexity's new model takes a different route: it encodes those chunks together, allowing each vector to be generated with access to the context of the full document.

The company used a simple example to explain the point. If a passage says only "it grew 20%," that sentence can be hard to interpret when viewed in isolation because "it" is undefined. With document-level context, the model can use other parts of the text to resolve that reference.

The model is now available on Hugging Face under the MIT license.

Perplexity also said the training process was designed so the model learns to retrieve both the passage containing an answer and other passages that support that answer. In practice, that means downstream AI systems can receive not just an answer, but also the evidence needed to check it.

In a blind test on context-bench, a benchmark designed by turbopuffer, the model recorded an Answer@10 score of 45.5%, which was 14.4 percentage points higher than Voyage Context 4. Its Evidence Recall@10 came in at 40.6%, 5 percentage points higher.

The evaluation covered 38,894 long documents and 2,099 queries. The questions were not made public, a setup Perplexity said helps reduce the risk that the test set is contaminated by training data.

The release is still labeled as a Preview version. Perplexity said the model weights, API, and generated embeddings may change later, so vectors created now may not be directly usable with future versions.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.