• bitcoinBitcoin(BTC)$76,815.00-0.68%
  • ethereumEthereum(ETH)$2,493.84-1.57%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$717.66-2.25%
  • rippleXRP(XRP)$1.35-1.45%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.92-1.83%
  • tronTRON(TRX)$0.3395260.02%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,108.66-3.61%
  • HyperliquidHyperliquid(HYPE)$77.84-1.93%
  • dogecoinDogecoin(DOGE)$0.083411-1.58%
  • RainRain(RAIN)$0.0154161.77%
  • moneroMonero(XMR)$538.85-0.67%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.77-0.78%
  • chainlinkChainlink(LINK)$11.31-2.04%
  • leo-tokenLEO Token(LEO)$9.06-0.65%
  • cardanoCardano(ADA)$0.203862-2.39%
  • stellarStellar(XLM)$0.177785-1.79%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$222.35-4.00%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.69-0.77%
  • uniswapUniswap(UNI)$6.19-2.60%
  • CantonCanton(CC)$0.095246-3.64%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.99%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0746970.29%
  • avalanche-2Avalanche(AVAX)$7.30-2.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.20%
  • nearNEAR Protocol(NEAR)$2.30-3.14%
  • suiSui(SUI)$0.71-2.61%
  • crypto-com-chainCronos(CRO)$0.0581750.94%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,347.37-0.06%
  • MemeCoreMemeCore(M)$1.15-2.65%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.82-0.26%
  • BittensorBittensor(TAO)$232.60-1.33%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.18%
  • aaveAave(AAVE)$124.40-1.67%
  • pax-goldPAX Gold(PAXG)$4,354.57-0.02%
  • AsterAster(ASTER)$0.69-0.02%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574331.09%
  • mantleMantle(MNT)$0.55-4.09%
  • polkadotPolkadot(DOT)$1.00-4.32%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research Introduces Fast and Expressive LLM Inference with RadixAttention and SGLang

January 24, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Research Introduces Fast and Expressive LLM Inference with RadixAttention and SGLang
ShareShareShareShareShare

Advanced prompting mechanisms, control flow, contact with external environments, many chained generation calls, and complex activities are expanding the utilization of Large Language Models (LLMs). On the other hand, effective methods for developing and running such programs are severely lacking. LMSYS ORG presents SGLang, a Structured Generation Language for LLMs that collaborates on the architecture of both the backend runtime system and the frontend languages. SGLang improves interactions with LLMs, making them faster and more controllable.

Backend: Automatic KV Cache Reuse with RadixAttention

To take advantage of these reuse opportunities systematically, the team provides RadixAttention, a new automatic KV cache reuse method while running. The KV cache is not removed from the radix tree when a generation request is completed; it is kept for both the generation results and the prompts. This data structure makes efficient search, insertion, and eviction of prefixes possible. To improve the cache hit rate, the researchers employ a cache-aware scheduling policy in conjunction with a Least Recently Used (LRU) eviction policy. It can be eagerly executed using an interpreter or traced as a dataflow graph and run with a graph executor. In the second scenario, compiler optimizations like code relocation, instruction selection, and auto-tuning become possible. 

Frontend: Easy LLM Programming with SGLang

The team also presents SGLang, an embedded domain-specific language in Python, on the front end. Complex methods of prompting, control flow, multi-modality, decoding limitations, and external interaction can be simply articulated using it. Users can run an SGLang function through local models, OpenAI, Anthropic, and Gemini.

As mentioned by the team, much of SGLang’s syntax takes cues from Guidance. Users also deal with batching and intra-program parallelism in addition to introducing new primitives. With all these new features, SGLang is much more powerful than before. Improve the cache hit rate with an eviction policy and a scheduling approach that considers cache awareness.

The researchers recorded the throughput their system attained when testing it on the following typical LLM workloads:

  • MMLU: A multi-tasking, 5-shot, multiple-choice test.
  • HellaSwag: An assessment tool for 20-shot, multiple-choice phrase completion.
  • An agent job based on prompt traces taken from the original ReAct paper is ReAct Agent.
  • Tree-of-Thought: A GSM-8K problem-solving prompt based on bespoke tree searches.
  • A JSON decoder can parse a Wikipedia article and return its data in a JSON format.
  • The chat (short) benchmark is a synthetic chat in which each conversation consists of four turns with brief LLM outputs.
  • This synthetic chat benchmark uses long LLM outputs and four turns per conversation.
  • DSPy RAG: A pipeline in the DSPy tutorial that uses retrieval to augment generation.
  • The LLaVA-in-the-wild benchmark is used to run the vision language model LLaVA v1.5.

Using the Llama-7B and Mixtral-8x7B models on NVIDIA A10G GPUs, the team applied SGLang to typical LLM workloads such as agent, reasoning, extraction, chat, and few-shot learning tasks. The researchers used Hugging Face TGI v1.3.0, advice v0.1.8, and vllm v0.2.5 as a starting point. SGLang outperforms current systems, specifically Guid, by a factor of up to five in terms of throughput. It also performed quite well in latency tests, especially those involving the initial token, where a prefix cache hit is very useful. Current systems do a terrible job of handling sophisticated LLM programs, but while developing the SGLang runtime, it was observed that a critical optimization opportunity: KV cache reuse. By reusing the KV cache, many prompts that share the same prefix can use the intermediate KV cache, which saves both memory and computation. Many other KV cache reuse methods, including ance and vLLM, can be found in complicated programs that use many LLM calls. The automatic KV cache reuse with RadixAttention, the interpreter’s ability to provide intra-program parallelism, and the fact that the frontend and backend systems were co-designed all contribute to these benefits. 


Check out the Code and Blog. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.



Credit: Source link

ShareTweetSendSharePin

Related Posts

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Next Post
Elon Musk: “I’m Jewish by association”

Elon Musk: “I’m Jewish by association”

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Start Investing Early – The Gap Between 25 and 35 Is Enormous

Start Investing Early – The Gap Between 25 and 35 Is Enormous

September 7, 2026
Missouri Supreme Court Sets Contempt Hearing in Fight Over House Map – The New York Times

Missouri Supreme Court Sets Contempt Hearing in Fight Over House Map – The New York Times

September 9, 2026
The Commuter’s Paradox: Why We Misjudge Social Connection (With Nick Epley)

The Commuter’s Paradox: Why We Misjudge Social Connection (With Nick Epley)

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!