• bitcoinBitcoin(BTC)$85,543.005.12%
  • ethereumEthereum(ETH)$2,730.152.54%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$787.511.70%
  • rippleXRP(XRP)$1.526.75%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.634.63%
  • tronTRON(TRX)$0.3486401.68%
  • zcashZcash(ZEC)$1,459.74-4.22%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$93.07-0.73%
  • dogecoinDogecoin(DOGE)$0.10196315.52%
  • moneroMonero(XMR)$571.60-0.24%
  • whitebitWhiteBIT Coin(WBT)$86.033.47%
  • RainRain(RAIN)$0.013784-2.70%
  • chainlinkChainlink(LINK)$12.942.18%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2466896.44%
  • leo-tokenLEO Token(LEO)$8.960.50%
  • stellarStellar(XLM)$0.2131647.59%
  • nearNEAR Protocol(NEAR)$4.35-1.78%
  • uniswapUniswap(UNI)$8.921.63%
  • bitcoin-cashBitcoin Cash(BCH)$264.764.06%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.92-5.02%
  • litecoinLitecoin(LTC)$60.883.72%
  • CantonCanton(CC)$0.1176634.97%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.038.78%
  • hedera-hashgraphHedera(HBAR)$0.0935267.30%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.15%
  • shiba-inuShiba Inu(SHIB)$0.00000610.53%
  • BittensorBittensor(TAO)$312.4717.17%
  • crypto-com-chainCronos(CRO)$0.0676478.84%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.40-7.81%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,341.68-0.47%
  • okbOKB(OKB)$121.901.78%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.25%
  • BitwayBitway(BTW)$0.8310.33%
  • aaveAave(AAVE)$143.072.72%
  • pepePepe(PEPE)$0.00000530.19%
  • mantleMantle(MNT)$0.656.49%
  • EthenaEthena(ENA)$0.210705-1.66%
  • OndoOndo(ONDO)$0.4350301.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

AppWorld: An AI Framework for Consistent Execution Environment and Benchmark for Interactive Coding for API-Based Tasks

August 2, 2024
in AI & Technology
Reading Time: 5 mins read
A A
AppWorld: An AI Framework for Consistent Execution Environment and Benchmark for Interactive Coding for API-Based Tasks
ShareShareShareShareShare

The prospects and scope for automation in digital lives are expanding with the advances in instruction following, coding, and tool-use abilities of large language models (LLMs). Most day-to-day digital tasks involve complex activities across various applications, with reasoning and decision-making based on intermediate results. However, the responsive development of such autonomous agents needs rigorous, reproducible, and strong evaluation using realistic tasks that account for the complexities and dynamics of real digital environments. The current benchmarks for tool-based solutions cannot solve this challenge as they use a linear sequence of API calls without rich or interactive coding, and their evaluations through reference solutions are not suitable for complex tasks with varied solutions.

The current benchmarks discussed in this paper are Tool-Usage Benchmarks (TUB) and Interactive Code Generation Benchmarks (ICGB). TUB either does not provide agents with executable tools or uses existing public APIs, with some offering implementations of simple ones. Current evaluation methods depend on LLMs or human judgment, which are unsuitable for tasks with multiple valid solutions. ICGB evaluates the ability of agents to generate executable code, such as HumanEval targeting short code snippets and SWEBench focusing on patch file generation. Intercode proposes solving coding tasks interactively by observing code execution outputs, while MINT allows agents to use a Python interpreter for reasoning and decision-making.

YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

Researchers from Stony Brook University, Allen Institute for AI, and Saarland University have proposed the AppWorld Engine, a high-quality execution environment comprising 60K lines of code. This environment includes 9 day-to-day apps operable through 457 APIs and simulates realistic digital activities for approximately 100 fictitious users. An AppWorld Benchmark, a collection of 750 diverse and complex tasks for autonomous agents, is developed that requires rich and interactive code generation. It enables robust programmatic evaluation with state-based unit tests, allowing for different task completion methods and checking for unexpected changes. 

The AppWorld Engine implements 9 applications across various domains, including emails (Gmail), money transfer (Venmo), shopping (Amazon), and local file systems. It features 457 APIs that closely resemble real app functionalities, averaging 50 APIs per app, and contains 1470 arguments. These APIs perform actions through read/write operations on a database, e.g. a send email API creates new entries in the email and email thread tables for both sender and recipient(s). Moreover, two supporting apps, ApiDocs and Supervisor, are implemented. ApiDocs provides APIs for interactive documentation, while Supervisor APIs provide information about the task assigner, such as addresses, payment cards, and account passwords.

The results show that all methods produce low task (TGC) and scenario (SGC) completion scores in both Test-N and Test-C. The strongest model, ReAct + GPT4O, achieves a TGC of 48.8 on Test-N, which decreases to 30.2 on Test-C. The 30-50% reduction from task to scenario scores shows that models do not consistently complete all task variants within the same scenario. The second-best model, GPT4Trb, falls significantly behind GPT4O, with open models performing even worse. GPT4Trb achieves a TGC of 32.7 and 17.5, while the best open LLM, FullCodeRefl + LLaMA3, gets a TGC of 24.4 on Test-N and 7.0 on Test-C. CodeAct and ToolLLaMA failed on all tasks due to their specialized narrow-domain training.

In summary, researchers have introduced the AppWorld Engine, a robust execution environment consisting of 60K lines of code. The AppWorld framework provides a consistent execution environment and a benchmark for interactive API-based tasks. Its programmatic evaluation suite and realistic challenges ensure thorough assessment. Benchmarking state-of-the-art models highlights the difficulty of AppWorld and the challenges that LLMs encounter in automating tasks. The system’s modularity and extensibility create opportunities for user interface control, coordination among multiple agents, and the examination of privacy and safety issues in digital assistants.


Check out the Paper, GitHub, and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
Stay Tuned NOW with Gadi Schwartz – April 30 | NBC News NOW

Stay Tuned NOW with Gadi Schwartz - April 30 | NBC News NOW

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fake NFL athlete accused of .3M romance fraud scam

Fake NFL athlete accused of $1.3M romance fraud scam

September 20, 2026
Interparfums: The Light At The End Of The Tunnel Is Getting Brighter

Interparfums: The Light At The End Of The Tunnel Is Getting Brighter

September 17, 2026
NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!