• bitcoinBitcoin(BTC)$84,224.001.02%
  • ethereumEthereum(ETH)$2,715.231.55%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$770.920.90%
  • rippleXRP(XRP)$1.510.36%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.400.08%
  • tronTRON(TRX)$0.3375080.06%
  • zcashZcash(ZEC)$1,437.692.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.29%
  • HyperliquidHyperliquid(HYPE)$89.133.85%
  • dogecoinDogecoin(DOGE)$0.0957752.00%
  • chainlinkChainlink(LINK)$14.470.42%
  • moneroMonero(XMR)$551.571.53%
  • whitebitWhiteBIT Coin(WBT)$84.231.04%
  • USDSUSDS(USDS)$1.000.02%
  • cardanoCardano(ADA)$0.2536003.08%
  • RainRain(RAIN)$0.012383-0.76%
  • leo-tokenLEO Token(LEO)$8.88-1.64%
  • stellarStellar(XLM)$0.2277482.15%
  • nearNEAR Protocol(NEAR)$5.479.96%
  • bitcoin-cashBitcoin Cash(BCH)$308.760.14%
  • uniswapUniswap(UNI)$8.940.92%
  • litecoinLitecoin(LTC)$67.310.14%
  • CantonCanton(CC)$0.1264740.92%
  • avalanche-2Avalanche(AVAX)$11.11-1.64%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.170.98%
  • hedera-hashgraphHedera(HBAR)$0.1063361.81%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.520.53%
  • quant-networkQuant(QNT)$290.500.86%
  • BitwayBitway(BTW)$1.3816.16%
  • BittensorBittensor(TAO)$308.501.14%
  • shiba-inuShiba Inu(SHIB)$0.0000060.37%
  • crypto-com-chainCronos(CRO)$0.0685941.94%
  • tether-goldTether Gold(XAUT)$4,184.410.25%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • EthenaEthena(ENA)$0.27289410.82%
  • Pump.funPump.fun(PUMP)$0.0058732.22%
  • aaveAave(AAVE)$167.072.88%
  • okbOKB(OKB)$121.570.37%
  • OndoOndo(ONDO)$0.510.37%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.051.03%
  • mantleMantle(MNT)$0.691.06%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

October 1, 2026
in AI & Technology
Reading Time: 7 mins read
A A
Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
ShareShareShareShareShare

Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ P

Is it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. It is not yet on the Perplexity API. The model card notes that weights and interface may change without backward compatibility.

YOU MAY ALSO LIKE

Apple TV Is Down

How To Watch NASA’s Crew-13 Launch

Why the gold passage falls short

RAG systems split long documents into chunks. A chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address this with late chunking. The document is encoded in one pass, then pooled per chunk.

Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists 3 more problems. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. Labels are also tied to one chunking strategy.

How the training works

The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.

  • Chunk relevance: the mean of the top n token scores inside each chunk.
  • Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
  • Distillation loss: forward KL divergence between teacher and student distributions.
  • Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.

Each batch samples a random chunking strategy. Chunks are separated by a learned token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.

The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.

Interactive explainer



pplx-embed-v2 Contextual Embedding Explainer

An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.




Three lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.

QUERYWhen does 5 Park Avenue’s lease end and what is the current rent?

Pick a mode and press Run retrieval.

Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.

A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.

1 Teacher scores tokens2 Top-n mean per chunk3 Softmax target4 Student matches via KL

Gold-chunk label (one-hot)

Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.

context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.

pplx-embed-v2-context-9b-previewvoyage-context-4 (derived from reported gap)

Voyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.

Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.

Chunks

0

vector storage (vectors only, not full index)

Chunk-size sensitivity (64 to 512 tokens)

81.0% to 79.9%

Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.

context-bench: a new benchmark

context-bench is built and privately held by turbopuffer to limit training contamination. It holds 2,099 queries over 38,894 documents in 21 domains. Sentence chunking yields 2,458,072 chunks. Median target document length is roughly 6,100 tokens. Queries test 12 contextual capabilities, from pronoun resolution to table structure. Perplexity says the model was submitted blind.

Metrics are Document@K, Answer@K, Evidence Recall@K, and All-Evidence@K. Every model is ranked exhaustively against all chunks, so index settings play no role.

Results

  • context-bench at K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence recall.
  • Document recall: 15.2% at K = 1 and 61.6% at K = 10.
  • vs voyage-context-4: ahead by 14.4 points on answer recall and 5.0 on evidence recall at K = 10.
  • ConTEB: highest average nDCG@10 among models shown. pplx-embed-context-v1-4B wins NarrativeQA, and Nemotron-3-Embed-8B wins COVID-QA.
  • General retrieval: best average on query-to-chunk tasks. Slightly behind voyage-context-4 on query-to-document.
  • Storage: 1024-dim int8 (1 KB per vector) slightly beats voyage-context-4 at 2048-dim float32 (8 KB).
  • Chunk size: average score moves from 81.0% to 79.9% between 64 and 512 tokens.

Comparison with the closest competitors

Feature pplx-embed-v2-context-9b-preview voyage-context-4 pplx-embed-context-v1-4B Nemotron-3-Embed-8B
Chunk embeddings Contextual Contextual Contextual Independent per chunk
Access Open weights, MIT Hosted API (Voyage, MongoDB Atlas) Open weights, MIT; Perplexity API Open weights, OpenMDW-1.1
Parameters 9B per blog (Hugging Face lists 8B) Not disclosed (MoE backbone) 4B About 8B
Dimensions 2048, 1024 2048, 1024, 512, 256 2560 (Matryoshka) 4096, sliceable
Quantized output Native int8 int8, uint8, binary, ubinary int8, binary Float
Context Evaluated up to 32,768 tokens 32K per request; 120K with auto-chunking 32K 32,768
Auto-chunking No Yes No No
Price Self-host $0.12 per 1M tokens; first 200M free Self-host or API Self-host

Sources: Voyage docs, model cards linked above. Checked September 30, 2026.

Key Takeaways

  • A token-level teacher replaces the single gold-chunk label.
  • The model retrieves answers plus supporting evidence in one chunk index.
  • context-bench Answer@10 is 45.5%, 14.4 points above voyage-context-4.
  • Open MIT weights ship today; Perplexity API access is still pending.

Check out the technical details and model weights. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple TV Is Down
AI & Technology

Apple TV Is Down

October 1, 2026
How To Watch NASA’s Crew-13 Launch
AI & Technology

How To Watch NASA’s Crew-13 Launch

October 1, 2026
Elon Musk And Palmer Luckey Will Advise The Government On The Future Of Warfare
AI & Technology

Elon Musk And Palmer Luckey Will Advise The Government On The Future Of Warfare

September 30, 2026
Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense
AI & Technology

Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense

September 30, 2026
Next Post
Pennsylvania reports 5th measles-associated death as outbreak continues – The Washington Post

Pennsylvania reports 5th measles-associated death as outbreak continues - The Washington Post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Credo: The Rise Of Muse Validates The AI Trade (NASDAQ:CRDO)

Credo: The Rise Of Muse Validates The AI Trade (NASDAQ:CRDO)

September 25, 2026
STIP: Simple TIPS ETF, For Risk-Averse Investors Concerned About Inflation (NYSEARCA:STIP)

STIP: Simple TIPS ETF, For Risk-Averse Investors Concerned About Inflation (NYSEARCA:STIP)

September 25, 2026
Sam Altman, Mark Zuckerberg, other billionaires at Trump-Xi dinner

Sam Altman, Mark Zuckerberg, other billionaires at Trump-Xi dinner

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!