• bitcoinBitcoin(BTC)$77,357.001.20%
  • ethereumEthereum(ETH)$2,474.581.64%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$751.503.60%
  • rippleXRP(XRP)$1.321.37%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$104.074.50%
  • tronTRON(TRX)$0.3358240.07%
  • zcashZcash(ZEC)$1,490.347.61%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.14%
  • HyperliquidHyperliquid(HYPE)$86.539.26%
  • dogecoinDogecoin(DOGE)$0.0837093.24%
  • moneroMonero(XMR)$518.083.32%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$79.651.38%
  • RainRain(RAIN)$0.012668-2.00%
  • chainlinkChainlink(LINK)$11.695.00%
  • leo-tokenLEO Token(LEO)$8.89-0.65%
  • cardanoCardano(ADA)$0.2146179.79%
  • stellarStellar(XLM)$0.1874702.56%
  • uniswapUniswap(UNI)$8.4625.63%
  • bitcoin-cashBitcoin Cash(BCH)$245.0311.19%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.4027.83%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.10869911.65%
  • litecoinLitecoin(LTC)$54.704.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.352.93%
  • avalanche-2Avalanche(AVAX)$7.803.54%
  • hedera-hashgraphHedera(HBAR)$0.0766183.44%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • suiSui(SUI)$0.787.99%
  • shiba-inuShiba Inu(SHIB)$0.0000057.29%
  • crypto-com-chainCronos(CRO)$0.0584520.09%
  • MemeCoreMemeCore(M)$1.2613.14%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BittensorBittensor(TAO)$240.487.33%
  • tether-goldTether Gold(XAUT)$4,334.811.15%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$113.532.16%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.16%
  • aaveAave(AAVE)$133.219.78%
  • AsterAster(ASTER)$0.754.44%
  • mantleMantle(MNT)$0.595.22%
  • Pump.funPump.fun(PUMP)$0.0041217.85%
  • OndoOndo(ONDO)$0.39246012.08%
  • BitwayBitway(BTW)$0.70-3.87%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Parrot: Optimizing End-to-End Performance in LLM Applications Through Semantic Variables

June 3, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Parrot: Optimizing End-to-End Performance in LLM Applications Through Semantic Variables
ShareShareShareShareShare

Large language models (LLMs) possess advanced language understanding, enabling a shift in application development where AI agents communicate with LLMs via natural language prompts to complete tasks collaboratively. Applications like Microsoft Teams and Google Meet use LLMs to summarize meetings, while search engines like Google and Bing enhance their capabilities with chat features. These LLM-based applications often require multiple API calls, creating complex workflows. Current API designs for LLM services are request-centric and lack application-level information, which results in sub-optimal performance.

The field of model serving has seen significant advancements with systems like Clipper, TensorFlow Serving, and AlpaServe addressing deep learning deployment challenges. These systems focus on batching, caching, and scheduling but often overlook the unique needs of LLMs. Orca and vLLM improve batching and memory utilization for LLM requests. Parrot enhances LLM serving by analyzing application-level data flow, and optimizing end-to-end performance. LLM orchestrator frameworks like LangChain and Semantic Kernel simplify LLM application management. Parrot integrates with these frameworks, utilizing Semantic Variables for optimization. Parrot also uses DAG information to optimize LLM applications, emphasizing prompt structure and request dependencies.

Researchers from Shanghai Jiao Tong University and Microsoft Research proposed Parrot, an LLM service system designed to treat LLM applications as first-class citizens, retaining application-level information through the use of Semantic Variables. A Semantic Variable is a text region in a prompt with a specific semantic purpose, such as task instructions or inputs, and it connects multiple LLM requests. By exposing prompt structures and request correlations, Parrot enables data flow analysis, optimizing end-to-end performance. Parrot’s unified abstraction facilitates joint optimizations, improving scheduling, latency hiding, and de-duplication.

Parrot treats LLM requests as semantic functions implemented in natural language, executed by LLMs. Semantic Variables, defined as input or output placeholders in prompts, maintain the prompt structure for inter-request analysis. In multi-agent applications, such as MetaGPT, semantic functions like WritePythonCode and WriteTestCode use Semantic Variables to connect and sequence tasks. Parrot’s asynchronous design allows submitting and fetching requests separately, facilitating just-in-time relationship analysis. Performance criteria can be annotated for each variable, optimizing and scheduling based on end-to-end requirements like latency or throughput. 

Evaluating Parrot on both production and open-source LLM-based applications reveals significant improvements, achieving up to 11.7× speedup and 12× higher throughput compared to state-of-the-art solutions. These applications require numerous LLM calls, leading to high user-perceived latency. Treating requests individually can double end-to-end latency, but Parrot’s batching approach eliminates this overhead. By scheduling consecutive requests together, Parrot directly feeds outputs from one step to the next, bypassing network and queuing delays.

This study introduces Parrot, which optimizes the end-to-end performance of LLM applications by treating them as first-class citizens rather than focusing solely on individual requests. It introduces Semantic Variable, an abstraction that reveals dependencies and commonalities among LLM requests, creating new optimization opportunities. The evaluation demonstrates Parrot can enhance LLM-based applications by up to 11.7×. This approach opens new research directions for improving scheduling features, such as ensuring the fairness of end-to-end performance in LLM applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 43k+ ML SubReddit | Also, check out our AI Events Platform


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Google’s Revamped CC Is An AI Agent For Families And Groups

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Next Post
Sacramento mom recalls last conversation with son shot by other 10-year-old

Sacramento mom recalls last conversation with son shot by other 10-year-old

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

September 17, 2026
Hyper Light Drifter Dev At Risk Of Closing After Laying Off ‘Nearly Everyone’

Hyper Light Drifter Dev At Risk Of Closing After Laying Off ‘Nearly Everyone’

September 16, 2026
Meet Anthropic CEO Dario Amodei’s handpicked super-woke globalists he thinks will save us from an AI apocalypse

Meet Anthropic CEO Dario Amodei’s handpicked super-woke globalists he thinks will save us from an AI apocalypse

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!