• bitcoinBitcoin(BTC)$86,213.00-0.45%
  • ethereumEthereum(ETH)$2,750.14-0.53%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$786.04-1.77%
  • rippleXRP(XRP)$1.585.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.06-0.21%
  • tronTRON(TRX)$0.341545-0.81%
  • zcashZcash(ZEC)$1,518.663.19%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.042.73%
  • HyperliquidHyperliquid(HYPE)$96.163.12%
  • dogecoinDogecoin(DOGE)$0.0998551.19%
  • moneroMonero(XMR)$565.80-2.57%
  • whitebitWhiteBIT Coin(WBT)$86.65-0.56%
  • chainlinkChainlink(LINK)$12.98-0.38%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2514502.74%
  • RainRain(RAIN)$0.013128-6.04%
  • leo-tokenLEO Token(LEO)$8.970.71%
  • stellarStellar(XLM)$0.2155181.46%
  • bitcoin-cashBitcoin Cash(BCH)$335.7925.79%
  • uniswapUniswap(UNI)$9.214.30%
  • nearNEAR Protocol(NEAR)$4.354.59%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$62.760.64%
  • avalanche-2Avalanche(AVAX)$10.99-0.70%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.113966-2.06%
  • USD1USD1(USD1)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0985117.74%
  • suiSui(SUI)$1.00-2.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.45-0.16%
  • shiba-inuShiba Inu(SHIB)$0.0000061.75%
  • BittensorBittensor(TAO)$313.371.83%
  • crypto-com-chainCronos(CRO)$0.0670754.93%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.30-13.40%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,363.910.42%
  • okbOKB(OKB)$122.41-0.55%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BitwayBitway(BTW)$0.87-10.26%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.19%
  • aaveAave(AAVE)$144.11-1.12%
  • mantleMantle(MNT)$0.662.11%
  • OndoOndo(ONDO)$0.436107-1.78%
  • EthenaEthena(ENA)$0.207427-2.75%
  • Pump.funPump.fun(PUMP)$0.0044403.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

AutoBencher: A Metrics-Driven AI Approach Towards Constructing New Datasets for Language Models

July 17, 2024
in AI & Technology
Reading Time: 4 mins read
A A
AutoBencher: A Metrics-Driven AI Approach Towards Constructing New Datasets for Language Models
ShareShareShareShareShare

This paper addresses the challenge of effectively evaluating language models (LMs). Evaluation is crucial for assessing model capabilities, tracking scientific progress, and informing model selection.  Traditional benchmarks often fail to highlight novel performance trends and are sometimes too easy for advanced models, providing little room for growth. The research identifies three key desiderata that existing benchmarks often lack: salience (testing practically important capabilities), novelty (revealing previously unknown performance trends), and difficulty (posing challenges for existing models).

Current methods for evaluating language models involve constructing benchmarks that test specific capabilities, such as mathematical reasoning or understanding academic subjects. Prior works have constructed high-quality benchmarks guided by salience and difficulty. While these benchmarks are valuable, they often yield similar performance trends across different models, limiting their ability to highlight unique strengths and weaknesses.

YOU MAY ALSO LIKE

The Latest PlayStation Update Made PSSR 2.0 The Default For PS5 Pro Owners

Do USB Extenders Really Work And Are They Safe To Use?

The researchers of this paper propose a new tool, AutoBencher, which automatically generates datasets that fulfill the three desiderata: salience, novelty, and difficulty. AutoBencher uses a language model to search for and construct datasets from privileged information sources. This approach allows creation of more challenging and insightful benchmarks compared to existing ones. For instance, AutoBencher can identify gaps in LM knowledge that are not captured by current benchmarks, such as performance discrepancies on less common topics like the Permian Extinction or Fordism.

AutoBencher operates by leveraging a language model to propose evaluation topics within a broad domain (e.g., history) and constructing small datasets for each topic using reliable sources like Wikipedia. The tool evaluates each dataset based on its salience, novelty, and difficulty, selecting the best ones for inclusion in the benchmark. This iterative and adaptive process allows the tool to refine its dataset generation to maximize the desired properties continuously.

Additionally, AutoBencher employs an adaptive search process, where the trajectory of past generated benchmarks is used to improve the difficulty of proposed topics. This allows AutoBencher to identify and select topics that jointly maximize novelty and difficulty, subject to a salience constraint specified by the user.

To ensure high-quality datasets, AutoBencher incorporates privileged information that the evaluated LMs cannot access, such as detailed documents or specific data relevant to the topic. This privileged information helps generate accurate and challenging questions. The results show that AutoBencher-created benchmarks are, on average, 27% more novel and 22% more difficult than existing human-constructed benchmarks. The tool has been used to create datasets across various domains, including math, history, science, economics, and multilingualism, revealing new trends and gaps in model performance.

The problem of effectively evaluating language models is critical for guiding their development and assessing their capabilities. AutoBencher offers a promising solution by automating the creation of salient, novel, and difficult benchmarks, thereby providing a more comprehensive and challenging evaluation framework for language models. The authors demonstrate the effectiveness of their approach by generating diverse benchmarks that uncover previously unknown performance trends across a range of language models, providing valuable insights to guide future model development and selection. This approach highlights existing gaps in model knowledge and paves the way for future improvements.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Shreya Maji is a consulting intern at MarktechPost. She is pursued her B.Tech at the Indian Institute of Technology (IIT), Bhubaneswar. An AI enthusiast, she enjoys staying updated on the latest advancements. Shreya is particularly interested in the real-life applications of cutting-edge technology, especially in the field of data science.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

The Latest PlayStation Update Made PSSR 2.0 The Default For PS5 Pro Owners
AI & Technology

The Latest PlayStation Update Made PSSR 2.0 The Default For PS5 Pro Owners

September 22, 2026
Do USB Extenders Really Work And Are They Safe To Use?
AI & Technology

Do USB Extenders Really Work And Are They Safe To Use?

September 22, 2026
How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
Next Post
Eruption at Hawaii’s Kīlauea volcano seen from helicopter

Eruption at Hawaii's Kīlauea volcano seen from helicopter

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Build columns, not Trump’s arch, D.C. official tells administration – The Washington Post

Build columns, not Trump’s arch, D.C. official tells administration – The Washington Post

September 18, 2026
Comedian debuts surprise documentary on Theranos founder

Comedian debuts surprise documentary on Theranos founder

September 16, 2026
The Case for AI Guardrails Without a Slowdown

The Case for AI Guardrails Without a Slowdown

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!