• bitcoinBitcoin(BTC)$76,948.00-0.99%
  • ethereumEthereum(ETH)$2,452.94-0.16%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$712.30-0.41%
  • rippleXRP(XRP)$1.33-3.43%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$99.38-1.50%
  • tronTRON(TRX)$0.336102-1.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.79%
  • zcashZcash(ZEC)$1,094.75-9.86%
  • HyperliquidHyperliquid(HYPE)$79.06-4.18%
  • dogecoinDogecoin(DOGE)$0.083460-1.94%
  • RainRain(RAIN)$0.015578-3.20%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$511.510.43%
  • whitebitWhiteBIT Coin(WBT)$79.59-0.92%
  • chainlinkChainlink(LINK)$11.38-3.55%
  • leo-tokenLEO Token(LEO)$9.15-0.88%
  • cardanoCardano(ADA)$0.202113-4.80%
  • stellarStellar(XLM)$0.174194-2.86%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$224.59-8.47%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$52.390.41%
  • CantonCanton(CC)$0.096392-4.51%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-1.75%
  • uniswapUniswap(UNI)$5.96-0.76%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.35-4.84%
  • hedera-hashgraphHedera(HBAR)$0.073804-3.09%
  • nearNEAR Protocol(NEAR)$2.441.58%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.98%
  • suiSui(SUI)$0.72-5.57%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056119-0.62%
  • MemeCoreMemeCore(M)$1.18-1.13%
  • tether-goldTether Gold(XAUT)$4,328.43-1.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$111.770.00%
  • BittensorBittensor(TAO)$231.41-7.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.16%
  • mantleMantle(MNT)$0.58-2.60%
  • aaveAave(AAVE)$122.17-0.51%
  • pax-goldPAX Gold(PAXG)$4,332.31-0.96%
  • AsterAster(ASTER)$0.69-3.41%
  • polkadotPolkadot(DOT)$1.09-0.82%
  • OndoOndo(ONDO)$0.346528-1.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes ML-BENCH: A Novel Artificial Intelligence Approach Developed to Assess the Effectiveness of LLMs in Leveraging Existing Functions in Open-Source Libraries

November 24, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Proposes ML-BENCH: A Novel Artificial Intelligence Approach Developed to Assess the Effectiveness of LLMs in Leveraging Existing Functions in Open-Source Libraries
ShareShareShareShareShare

LLM models have been increasingly deployed as potent linguistic agents capable of performing various programming-related activities. Despite these impressive advances, a sizable chasm still separates the capabilities demonstrated by these models in static experimental settings from the ever-changing demands of actual programming scenarios.

Standard code generation benchmarks test how well LLM can generate new code from scratch. However, programming conventions rarely necessitate the genesis of all code components from scratch. 

When writing code for real-world applications, using existing, publicly available libraries is common practice. These developed libraries offer robust, battle-tested answers to various challenges. Therefore, the success of code LLMs should be evaluated in more ways than only function production, such as their skill in running code derived from open-source libraries with correct parameter usage.

A new study by Yale University, Nanjing University, and Peking University presents ML-BENCH, a realistic and comprehensive benchmark dataset for evaluating LLMs’ abilities to comprehend user instructions, navigate GitHub repositories, and produce executable code. High-quality, instructable ground truth code that satisfies the instructions’ requirements is made available by ML-BENCH. There are 9,444 examples, among 130 tasks and 14 popular machines learning GitHub repositories that make up ML-BENCH.

The researchers use Pass@k and Parameter Hit Precision as metrics in their investigations. Using these tools, they explore the possibilities of GPT-3.5-16k, GPT-4-32k, Claude 2, and CodeLlama in ML-BENCH environments. ML-BENCH suggests new tests for LLMs. The empirical results show that GPT models and Claude 2 outperformed CodeLlama by a wide margin. Although GPT-4 shows a significant performance increase over other LLMs, it still only completes 39.73% of the tasks in the experiments. Other well-known LLms experience hallucinations and underachieve. The findings suggest that LLMs must do more than just write code; they must also understand lengthy documentation. The key technological contribution is the proposal of the ML-AGENT, an autonomous language agent designed to address the deficiencies discovered through their error analysis. These agents can comprehend human language and instructions, generate efficient code, and do difficult tasks. 

ML-Bench and ML-Agent represent a significant advancement in the state of the art of automated machine learning processes. The researchers hope that this interests other researchers and practitioners alike. 


Check out the Paper and Project Page. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

How These XL Phones Compete

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


↗ Step by Step Tutorial on ‘How to Build LLM Apps that can See Hear Speak’

Credit: Source link

ShareTweetSendSharePin

Related Posts

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
AI & Technology

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

September 11, 2026
How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
Next Post
Trump ‘got frustrated’ and ‘stormed out’ of courtroom during Wednesday trial

Trump 'got frustrated' and 'stormed out' of courtroom during Wednesday trial

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Appeals court rejects Biden request to block recordings

Appeals court rejects Biden request to block recordings

September 7, 2026
How AI Turned Our Small Marketing Team into a Full-Service Agency – Unite.AI

How AI Turned Our Small Marketing Team into a Full-Service Agency – Unite.AI

September 4, 2026
Tesla’s Cybercab Is Under Federal Investigation Almost Immediately After Hitting The Streets

Tesla’s Cybercab Is Under Federal Investigation Almost Immediately After Hitting The Streets

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!