pplx-embed-v2-context-9b-preview

pplx-embed-v2-context-9b-preview is a contextual embedding model for document chunks in RAG systems. A document is passed as a list of chunks; the chunks are encoded together, so each chunk's embedding reflects its surrounding context, and one embedding is returned per chunk.

This is a preview release, not a final model. Weights, embeddings, and the interface may change in later versions without backward compatibility, so embeddings produced with this preview should not be mixed with embeddings from a future release.

Queries and documents are encoded with different methods: use encode_queries for queries and encode for document chunks. The model is trained with separate query and document prefixes, and encoding queries with encode silently degrades retrieval quality.

Like pplx-embed-context-v1, the model natively produces unnormalized int8-quantized embeddings. Compare embeddings with cosine similarity, or pass normalize_embeddings=True and use the dot product.

Model

Model Dimensions MRL Quantization Instruction Pooling
pplx-embed-v2-context-9b-preview 2048 1024, 2048 INT8 No (fixed query/document prefixes) Mean

Usage

Requires transformers>=5.4.0, torch, numpy, safetensors, and tqdm. The model uses custom code, so load it with trust_remote_code=True.

from transformers import AutoModel

model = AutoModel.from_pretrained(
    "perplexity-ai/pplx-embed-v2-context-9b-preview",
    trust_remote_code=True,
).to("cuda")

doc_chunks = [
    [
        "Curiosity begins in childhood with endless questions about the world.",
        "As we grow, curiosity drives us to explore new ideas.",
        "Scientific breakthroughs often start with a curious question.",
    ],
    [
        "The curiosity rover explores Mars searching for ancient life.",
        "Each discovery on Mars sparks new questions about the universe.",
    ],
]

# One (chunk_count, 2048) array per document:
# doc_embeddings[0].shape == (3, 2048), doc_embeddings[1].shape == (2, 2048)
doc_embeddings = model.encode(doc_chunks, normalize_embeddings=True)

# Each query is a single-chunk row.
queries = [["What drives scientific breakthroughs?"]]
query_embeddings = model.encode_queries(queries, normalize_embeddings=True)

scores = doc_embeddings[0] @ query_embeddings[0][0]

Options

encode(documents, ...) and encode_queries(queries, ...) accept:

Argument Default Description
batch_size 32 Documents (or queries) per forward pass
normalize_embeddings False L2-normalize outputs
convert_to_numpy True Return NumPy arrays; False returns CPU tensors
show_progress_bar False Show a progress bar
device None Move the model to this device before encoding

Matryoshka dimensions

The model was trained with Matryoshka losses at 1024 and 2048 dimensions. To use 1024-dimensional embeddings, take the first 1024 values of each unnormalized embedding and normalize afterwards:

import numpy as np

emb = model.encode(doc_chunks)  # unnormalized int8 values
emb_1024 = [e[:, :1024] / np.linalg.norm(e[:, :1024], axis=-1, keepdims=True) for e in emb]

Other truncation sizes were not trained.

Downloads last month
429
Safetensors
Model size
8B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using perplexity-ai/pplx-embed-v2-context-9b-preview 1