• bitcoinBitcoin(BTC)$81,321.004.33%
  • ethereumEthereum(ETH)$2,640.795.60%
  • tetherTether(USDT)$1.000.05%
  • binancecoinBNB(BNB)$767.702.86%
  • rippleXRP(XRP)$1.438.43%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$111.846.09%
  • tronTRON(TRX)$0.3380010.12%
  • zcashZcash(ZEC)$1,544.886.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.22%
  • HyperliquidHyperliquid(HYPE)$92.743.06%
  • dogecoinDogecoin(DOGE)$0.0883503.86%
  • moneroMonero(XMR)$584.629.01%
  • whitebitWhiteBIT Coin(WBT)$83.183.60%
  • RainRain(RAIN)$0.0139338.64%
  • USDSUSDS(USDS)$1.000.02%
  • chainlinkChainlink(LINK)$12.526.20%
  • cardanoCardano(ADA)$0.2266756.27%
  • leo-tokenLEO Token(LEO)$8.89-0.15%
  • stellarStellar(XLM)$0.1947375.08%
  • uniswapUniswap(UNI)$9.114.28%
  • bitcoin-cashBitcoin Cash(BCH)$251.121.41%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • nearNEAR Protocol(NEAR)$3.675.32%
  • daiDai(DAI)$1.00-0.02%
  • litecoinLitecoin(LTC)$58.165.79%
  • CantonCanton(CC)$0.1110703.85%
  • USD1USD1(USD1)$1.000.05%
  • avalanche-2Avalanche(AVAX)$9.2315.88%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.42%
  • hedera-hashgraphHedera(HBAR)$0.0807874.81%
  • suiSui(SUI)$0.868.69%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000052.00%
  • BittensorBittensor(TAO)$268.779.94%
  • crypto-com-chainCronos(CRO)$0.0599541.46%
  • MemeCoreMemeCore(M)$1.290.45%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,373.580.13%
  • okbOKB(OKB)$122.517.84%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • aaveAave(AAVE)$143.115.88%
  • AsterAster(ASTER)$0.773.03%
  • mantleMantle(MNT)$0.613.86%
  • OndoOndo(ONDO)$0.4129415.95%
  • EthenaEthena(ENA)$0.19668920.56%
  • Pump.funPump.fun(PUMP)$0.004154-0.14%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

How Scale Impacts Predicting Downstream Capabilities of Frontier AI Models: Understanding the Elusiveness

June 12, 2024
in AI & Technology
Reading Time: 4 mins read
A A
How Scale Impacts Predicting Downstream Capabilities of Frontier AI Models: Understanding the Elusiveness
ShareShareShareShareShare

Predicting the scaling behavior of frontier AI systems like GPT-4, Claude, and Gemini is essential for understanding their potential and making decisions about their development and use. However, it is difficult to predict how these systems will perform on specific tasks as they scale up, despite the well-established relation between parameters, data, compute, and pretraining loss defined by the scaling laws. For example, performance on standard NLP benchmarks can sometimes show unpredictable changes with scale. Some studies suggest these unpredictable changes might be due to choices of metrics and lack of resolution.

This paper contains two main directions. The first is “Beyond Multiple Choice Benchmarks”, where the study focuses on benchmarks evaluated using loglikelihood-based multiple-choice formats. While this focus is valuable due to the usefulness and prevalence of such tasks, it limits the broader application of the findings. The second direction is “Predicting Benchmark Performance A Priori”, which explains why multiple-choice benchmark performance is difficult to predict using metrics like Accuracy and Brier Score. However, the analyses assume access to the scores of entire model families across various orders of magnitude of pretraining FLOPs and do not utilize backtesting.

Researchers from the University of Cambridge, Stanford CS, EleutherAI, and MILA have shown that common multiple-choice metrics, such as Accuracy, Brier Score, and Probability Correct, can be evaluated from raw model outputs. This is achieved through a sequence of transformations that gradually degrades the statistical relationship between these metrics and the scaling parameters. The main reason is that these metrics depend on a direct comparison between the correct output and a limited set of specific incorrect outputs. Therefore, accurately predicting downstream performance needs modeling how the probability mass fluctuates among particular incorrect alternatives.

Researchers worked on how probability mass on incorrect choices fluctuates with increasing compute. This helps in understanding why individual downstream metrics can be unpredictable, while pretraining loss scaling laws are more consistent since they don’t depend on specific incorrect choices. To design evaluations that effectively track the progress of advanced AI capabilities, it’s important to understand what affects downstream performance. Moreover, to see how the downstream capabilities on specific tasks change with scale for different model families, per-sample scores are generated from various model families and multiple-choice NLP benchmarks.

To accurately predict performance on multiple-choice question-answering tests, it’s important to understand how the probability of choosing the correct answer changes with scale as well as how the probability of choosing the wrong answer changes with scale. For metrics such as Accuracy, these predictions need to be made for each question because knowing the average probability of choosing wrong answers across many questions doesn’t specify the probability of choosing a specific wrong answer for a particular question. It is especially important to look at how the probabilities of choosing the correct and incorrect answers change together as more computational power is used.

In conclusion, researchers have found a factor that causes unpredictability in multiple-choice tests for frontier AI models. This factor is the probability of choosing incorrect answers. The results can influence to design the of future evaluations for frontier AI models that are reliably predictable with scaling. Future work focuses on creating more predictable evaluations for AI systems, particularly for complex and important capabilities. The researchers gave several future directions for extending the work and adopting their framework to further improve scaling-predictable evaluations. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit


YOU MAY ALSO LIKE

Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model

GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model
AI & Technology

Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model

September 19, 2026
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
AI & Technology

GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)

September 19, 2026
Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI
AI & Technology

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI

September 19, 2026
How Focus Mode Has Changed In iOS 27
AI & Technology

How Focus Mode Has Changed In iOS 27

September 18, 2026
Next Post
Poll shows Biden losing support among young voters ahead of 2024 election

Poll shows Biden losing support among young voters ahead of 2024 election

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Pappas questions Trump’s ‘golden age’ after primary win

Pappas questions Trump’s ‘golden age’ after primary win

September 15, 2026
AIO: Don't Buy It To Win The AI Trade – Buy It Not To Lose

AIO: Don't Buy It To Win The AI Trade – Buy It Not To Lose

September 17, 2026
Morning News NOW Full Episode – Sept. 2

Morning News NOW Full Episode – Sept. 2

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!