• bitcoinBitcoin(BTC)$85,209.005.61%
  • ethereumEthereum(ETH)$2,722.345.26%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$794.405.25%
  • rippleXRP(XRP)$1.487.15%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.828.54%
  • tronTRON(TRX)$0.344722-0.49%
  • zcashZcash(ZEC)$1,519.865.12%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.020.00%
  • HyperliquidHyperliquid(HYPE)$94.493.50%
  • dogecoinDogecoin(DOGE)$0.0928128.72%
  • moneroMonero(XMR)$573.045.88%
  • whitebitWhiteBIT Coin(WBT)$85.724.50%
  • RainRain(RAIN)$0.0140756.75%
  • chainlinkChainlink(LINK)$12.956.67%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2417748.54%
  • leo-tokenLEO Token(LEO)$8.990.82%
  • stellarStellar(XLM)$0.2075727.79%
  • uniswapUniswap(UNI)$8.821.92%
  • bitcoin-cashBitcoin Cash(BCH)$267.468.28%
  • nearNEAR Protocol(NEAR)$4.1010.27%
  • avalanche-2Avalanche(AVAX)$11.183.63%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • litecoinLitecoin(LTC)$62.058.44%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.1146389.09%
  • USD1USD1(USD1)$1.000.01%
  • suiSui(SUI)$1.0323.69%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.433.67%
  • hedera-hashgraphHedera(HBAR)$0.0904362.93%
  • MemeCoreMemeCore(M)$1.48-3.30%
  • shiba-inuShiba Inu(SHIB)$0.0000066.09%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • BittensorBittensor(TAO)$283.6512.04%
  • crypto-com-chainCronos(CRO)$0.0634208.07%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,329.05-0.93%
  • okbOKB(OKB)$122.094.96%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BitwayBitway(BTW)$0.8620.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.28%
  • EthenaEthena(ENA)$0.2222719.11%
  • aaveAave(AAVE)$145.037.96%
  • OndoOndo(ONDO)$0.4477068.95%
  • mantleMantle(MNT)$0.635.82%
  • AsterAster(ASTER)$0.752.70%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers at Princeton University Reveal Hidden Costs of State-of-the-Art AI Agents

July 6, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Researchers at Princeton University Reveal Hidden Costs of State-of-the-Art AI Agents
ShareShareShareShareShare

There has been a lot of development in AI agents recently. However, one single goal—accuracy—has dominated evaluation and is vital to agent development. According to a recent study out of Princeton University, agents that are unnecessarily complicated and costly to run are the result of focusing only on accuracy. The team suggests a change to an evaluation paradigm that takes cost into account, where accuracy and cost are optimized together.

Standard metrics for gauging an agent’s efficacy on a given job have long been used in agent evaluation. Aiming for ever-increasing precision through ever-more-complicated models is a common trend that arises from these standards. The computing needs of these models may prevent them from being useful in the real world, even when they perform quite well on the benchmark.

YOU MAY ALSO LIKE

A Laptop That Works Better With Your Android Phone

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

The team points out where the existing evaluation system needs to improve in their study: 

  • Firstly, there is a risk that agents developed with an overemphasis on accuracy will not apply to real-world situations. Deploying highly accurate agents in contexts with limited resources is sometimes not viable due to their high computational cost.
  • Second, there is a chasm between model developers and downstream developers due to the present method. Upstream developers care more about how much it will cost to run the agent in production, while model developers focus on how accurate the model will be on the benchmark. The result of this discord may be agents with high levels of verifiable accuracy that are impractically expensive to implement in real-world scenarios.

The researchers propose an evaluation paradigm that considers the costs of solving these problems. By representing the cost and accuracy of agents as a Pareto frontier, a new avenue for agent design becomes apparent: maximizing both costs and accuracy simultaneously, which can result in agents with lower costs without sacrificing accuracy. This elaboration can be applied to many agent design criteria, including latency, without restrictions. 

The overall expense of managing an agent encompasses both fixed and variable expenses. When you optimize the agent’s hyperparameters (temperate, prompt, etc.) for a certain task, you’ll incur fixed costs. Running the agent incurs variable expenses proportional to the input and output token counts. Variable costs become increasingly important as the agent’s usage increases. The team could balance the agent’s fixed and variable expenses by utilizing joint optimization. They could lower the variable cost of running an agent (for example, by discovering shorter prompts and few-shot examples while preserving accuracy) by investing more upfront in the one-time optimization of agent design. They mentioned that if users want to operate agents for less money without compromising accuracy, it can be done by model trimming and hardware acceleration.

The modified version of the DSPy framework is tested on the HotPotQA benchmark to show how effective joint optimization can be. Since HotPotQA has been published in multiple official tutorials by the developers and was used as a benchmark to demonstrate DSPy’s efficiency in the original paper, the team decided to use it. To find few-shot instances that may be employed with an agent that decreases cost while preserving accuracy, the Optuna hyperparameter optimization framework was used. Please be aware that we anticipate significantly better performance from more intricate joint optimization methods. Joint optimization opens up a huge, uncharted design space in agent design, and the findings are just the tip of the iceberg.

The team tests the efficacy of DSPy-based multi-hop question-answering with several agent designs. They employ ColBERTv2 to conduct a HotPotQA-based query on Wikipedia as a retrieval strategy. To measure performance, they compare the agent’s retrieval success rate of all ground-truth documents included in the HotPotQA task. One hundred HotPotQA samples are utilized from the training set to fine-tune the DSPy pipelines, and 200 samples from the evaluation set to assess the outcomes. Five different agent architectures are evaluated as follows:

  1. Uncompiled: Neither the agent’s prompt optimization nor the formatting instructions for HotPotQA queries are provided in the uncompiled version. Without few-shot examples or formatting instructions, each prompt only includes the task instructions and the core content (i.e., question, context, rationale).
  2. Formatting instructions only: Like the uncompiled baseline, this one also includes formatting instructions for retrieval query outputs.
  3. Few Shot:  DSPy was used to find effective few-shot examples from all 100 samples in the training set. Few-shot examples are instances where the model is trained on a few examples, typically less than 100, to make predictions on new, unseen data. Attached are the instructions for formatting. Selecting few-shot examples is done by looking at the number of successful predictions on the training set. Random Search: they apply DSPy’s random search optimizer on half of the training data (out of 100 samples) to choose the best few-shot examples. The optimizer’s performance on the other half of the samples is then used to inform its decision-making. Attached are the instructions for formatting.
  4. Joint optimization: 50% of the training set is iterated to get a set of potential few-shot instances that enhance the model’s accuracy. For validation, the remaining fifty samples were used. Using parameter search, the team wanted to maximize accuracy while minimizing the number of tokens used in the few-shot samples provided by the prompt. Although DSPy provides significant accuracy increases over uncompiled baselines, it does so at a cost. Luckily, it is possible to reduce the price by utilizing joint optimization. Compared to the default DSPy implementations, it results in a variable cost of 53% lower while maintaining the same level of accuracy for GPT-3.5. For Llama-3-70B, it’s the same story: it reduces costs by 41% without sacrificing precision.  

It’s crucial that we rethink our approach to agent benchmarks. The current benchmarks often lead to agents that perform well in the benchmark but struggle in real-world scenarios. By considering factors such as distribution changes and downstream developer requirements, we can design more practical and effective benchmarks, addressing the urgency of this change.

As AI agents become more sophisticated, the importance of safety evaluations cannot be overstated. While this study doesn’t specifically address security concerns, it underscores the vital role of existing frameworks in regulating agentic AI. It’s crucial that developers prioritize and deploy these frameworks to ensure responsible development and deployment of AI agents.

The team states that their research empowers individuals to evaluate the cost-effectiveness of capabilities that could pose risks. This way, the community can spot and prevent possible safety issues before they increase. For this reason, makers of AI safety benchmarks should incorporate cost assessments. Ultimately, this work suggests a change in how agents are evaluated. To create useful and feasible agents for deployment in the real world, the team highlights that researchers need to shift their focus from accuracy alone to cost considerations. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

A Laptop That Works Better With Your Android Phone
AI & Technology

A Laptop That Works Better With Your Android Phone

September 21, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI
AI & Technology

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters
AI & Technology

Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters

September 21, 2026
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
AI & Technology

StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work

September 21, 2026
Next Post
Sanofi CEO on AI’s Impact, Lung Disease Drug Clearance

Sanofi CEO on AI's Impact, Lung Disease Drug Clearance

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Building a portfolio from scratch? David Wagner goes rapid-fire on the top names that make the cut

Building a portfolio from scratch? David Wagner goes rapid-fire on the top names that make the cut

September 17, 2026
Roblox Pushes Deeper Into AI-Powered Gaming

Roblox Pushes Deeper Into AI-Powered Gaming

September 16, 2026
Dolly Parton-themed corn maze honors singer after her death

Dolly Parton-themed corn maze honors singer after her death

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!