• bitcoinBitcoin(BTC)$84,292.000.47%
  • ethereumEthereum(ETH)$2,692.820.23%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$772.40-0.46%
  • rippleXRP(XRP)$1.52-3.04%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$120.75-0.73%
  • tronTRON(TRX)$0.333869-1.12%
  • zcashZcash(ZEC)$1,646.406.91%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.41%
  • HyperliquidHyperliquid(HYPE)$92.801.10%
  • dogecoinDogecoin(DOGE)$0.096238-2.39%
  • chainlinkChainlink(LINK)$14.060.33%
  • moneroMonero(XMR)$556.13-0.09%
  • whitebitWhiteBIT Coin(WBT)$84.090.39%
  • USDSUSDS(USDS)$1.00-0.02%
  • cardanoCardano(ADA)$0.251560-2.21%
  • RainRain(RAIN)$0.01278215.39%
  • leo-tokenLEO Token(LEO)$9.002.13%
  • stellarStellar(XLM)$0.215178-1.78%
  • bitcoin-cashBitcoin Cash(BCH)$333.84-1.95%
  • nearNEAR Protocol(NEAR)$5.042.63%
  • uniswapUniswap(UNI)$9.782.39%
  • litecoinLitecoin(LTC)$71.88-0.83%
  • CantonCanton(CC)$0.1358943.65%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.852.43%
  • suiSui(SUI)$1.17-0.94%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.589.22%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.092703-2.29%
  • BittensorBittensor(TAO)$318.581.87%
  • shiba-inuShiba Inu(SHIB)$0.000006-1.15%
  • crypto-com-chainCronos(CRO)$0.0678342.92%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.232.52%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • EthenaEthena(ENA)$0.2690640.97%
  • BitwayBitway(BTW)$1.00-25.14%
  • tether-goldTether Gold(XAUT)$4,278.95-0.03%
  • OndoOndo(ONDO)$0.54-1.70%
  • okbOKB(OKB)$120.70-0.06%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.43-0.64%
  • quant-networkQuant(QNT)$159.0461.83%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.14%
  • mantleMantle(MNT)$0.692.52%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

September 6, 2026
in AI & Technology
Reading Time: 12 mins read
A A
Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
ShareShareShareShareShare

Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second — the serving infrastructure behind pplx-embed and the ranking models used across Perplexity Search, Computer and the API Platform.

Perplexity team states that embedding inference on the GPU side has largely converged across engines on mature Hopper and Blackwell hardware. The wins sit in the runtime and harness around the model: CUDA graph management, an async result-tracking abstraction, and a Rust request path.

YOU MAY ALSO LIKE

Your Old GPU Could Be Worth More Than You Think

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

Two traffic patterns, one engine

Perplexity frames embedding serving as two workloads. Batch embedding happens when building or re-indexing the vector database, where throughput minimizes cost. Online embedding happens at query time, where a short query must be embedded fast. Scoring sits in between: after vector search, large document batches are ranked, balancing both.

The key decision is that Perplexity did not build a separate embedding engine. Because embedding models are small Transformers, batch embedding resembles compute-bound prefill and online embedding, often a few tokens, resembles memory-bound decode. So the research team reuses the prefill and decode kernels from its LLM stack.

Ivy, Tulip and ROSE

Three services handle a request:

  • Ivy is a Rust HTTP gateway. It does the CPU-side work — JSON parsing, tokenization, input templating, batch splitting — and translates requests into a custom gRPC protocol. It also splits large-batch requests into chunks and load-balances them across replicas, which corrects the load imbalance that arises when production payloads vary in size.
  • Tulip is the inference server interface: a gRPC server built with Rust, tokio and tonic, handling scheduling and batching before dispatching to the engine.
  • ROSE (Runtime-Optimized Serving Engine) implements model inference. It is primarily Python, provides kernels, layers and model definitions, manages CUDA graphs, and exposes a step() function to Tulip.

Why the scheduler is deliberately simple

Tulip picks sequences first-come, first-served while requests accumulate. That simplicity is justified by a measurement: for small embedding models at the sequence lengths Perplexity serves, the linear cost of dense layers dominates the quadratic cost of attention. Latency is therefore roughly proportional to token count, not sequence count. Once a batch saturates the GPU, around 512 tokens on a sub-billion-parameter model, packing in more sequences does not improve efficiency.

CUDA graphs and LazyTensors

On small batches, CPU-side kernel launching can outweigh GPU execution. Perplexity builds whole-model CUDA graphs for all embedding models, capturing every launch into a single driver call. Because embedding models are small, the inflection point where GPU work exceeds launch cost arrives at batches of thousands of tokens and tens of sequences. Some attention implementations block full-model graphs by depending on dynamic host-side inputs; Perplexity upstreamed changes to FlashInfer to enable capture.

Graphs must be captured per configuration, so token counts are padded to buckets that are multiples of 64 or 256. That still yields thousands of graphs and multiple minutes of capture per model. The fix is lazy capture: each configuration gets an eager warmup run, then triggers capture and replay on its second hit. This costs p99 latency at startup but spreads minutes of eager work across hours.

The second piece is the LazyTensor, which tracks a page-locked host buffer plus a cudaMemcpyAsync and a CUDA event. Instead of step() blocking on the device, it returns a LazyTensor, letting a Rust async task wait on batch N while the CPU enqueues N+1.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
AI & Technology

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

September 26, 2026
Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
AI & Technology

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

September 26, 2026
This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig
AI & Technology

This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

September 26, 2026
Next Post
Lululemon billionaire founder Chip Wilson files for divorce after 20 years of marriage — with no prenup in place

Lululemon billionaire founder Chip Wilson files for divorce after 20 years of marriage — with no prenup in place

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
David Ellison eyes Elon Musk for Paramount equity investment: report

David Ellison eyes Elon Musk for Paramount equity investment: report

September 23, 2026
Stay Tuned NOW Streaming Behind The Scenes! – Aug 28

Stay Tuned NOW Streaming Behind The Scenes! – Aug 28

September 21, 2026
Susan Sarandon and Hannah Einbinder arrested at anti-Netanyahu protest in New York – The Guardian

Susan Sarandon and Hannah Einbinder arrested at anti-Netanyahu protest in New York – The Guardian

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!