Perplexity open-sources retrieval models that let a 0.6B model query indexes built by a 9B model

Perplexity open-sources retrieval models that let a 0.6B model query indexes built by a 9B model

N
News Editor
2026-10-08 11:58:04
Perplexity has open-sourced two multimodal embedding models under its pplx-embed-v2-late line, with parameter sizes of 0.6B and 9B. The key design point is vector compatibility between the two: developers can build a database with the 9B model and then run day-to-day searches with the smaller 0.6B model, avoiding the need to use the larger model for every query. The models support retrieval across text, images, and PDF pages. Perplexity said the system keeps a 128-dimensional vector for each token instead of compressing an entire passage into a single vector, a setup aimed at preserving more detail during matching. For PDFs, slide decks, and scanned files, page images can be converted directly into vectors without first extracting text through OCR, which also keeps charts, tables, and layout information. According to the company, both models were trained from the same 18B teacher model and share one vector space. In ViDoRe v3 image retrieval testing, Perplexity reported a score of 62.3% when the 0.6B model handled both indexing and querying. Using the 9B model for indexing and the 0.6B model for querying raised the score to 63.5%, while using the 9B model for both reached 65.2%. The model weights are available on Hugging Face under the MIT license.

Perplexity has released two multimodal embedding models, pplx-embed-v2-late, in 0.6B and 9B parameter versions. Their vectors are compatible with each other, which lets developers build a database with the 9B model and use the 0.6B model for routine search queries instead of running the larger model every time.

Text, image, and PDF page retrieval

The new models support retrieval across text, images, and PDF pages. Standard embedding models often compress a passage into a single vector, which can lose detail. Perplexity said the new setup keeps a 128-dimensional vector for each token so search terms can match the relevant parts of a document more precisely.

For PDFs, PPT files, and scanned documents, page images can be turned directly into vectors without first extracting text through OCR. That process also preserves charts, tables, and layout information.

Shared vector space from one teacher model

According to Perplexity, both models were trained from the same 18B teacher model and share a single vector space.

Benchmark results disclosed by the company

In ViDoRe v3 image retrieval testing, the company said a setup using the 0.6B model for both indexing and querying scored 62.3%. Switching to 9B for indexing while keeping 0.6B for querying lifted the score to 63.5%. Using the 9B model for both indexing and querying scored 65.2%.

The weights for both models are now available on Hugging Face under the MIT license.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.