• bitcoinBitcoin(BTC)$83,039.00-2.16%
  • ethereumEthereum(ETH)$2,667.03-1.56%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$763.15-2.12%
  • rippleXRP(XRP)$1.49-3.04%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.62-4.31%
  • tronTRON(TRX)$0.333958-0.07%
  • zcashZcash(ZEC)$1,566.91-5.66%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$89.91-3.49%
  • dogecoinDogecoin(DOGE)$0.093167-4.85%
  • chainlinkChainlink(LINK)$14.08-1.40%
  • moneroMonero(XMR)$530.42-4.17%
  • whitebitWhiteBIT Coin(WBT)$83.06-1.93%
  • USDSUSDS(USDS)$1.00-0.02%
  • cardanoCardano(ADA)$0.245678-4.38%
  • RainRain(RAIN)$0.012533-1.12%
  • leo-tokenLEO Token(LEO)$9.060.46%
  • stellarStellar(XLM)$0.212733-2.25%
  • nearNEAR Protocol(NEAR)$5.12-1.27%
  • bitcoin-cashBitcoin Cash(BCH)$310.33-8.45%
  • uniswapUniswap(UNI)$8.97-9.74%
  • litecoinLitecoin(LTC)$71.20-0.64%
  • CantonCanton(CC)$0.131061-5.18%
  • hedera-hashgraphHedera(HBAR)$0.11740923.96%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.50-4.49%
  • suiSui(SUI)$1.18-5.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.662.34%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.00%
  • BittensorBittensor(TAO)$306.15-7.28%
  • BitwayBitway(BTW)$1.2710.50%
  • quant-networkQuant(QNT)$234.9343.67%
  • shiba-inuShiba Inu(SHIB)$0.000006-4.88%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.064855-4.53%
  • tether-goldTether Gold(XAUT)$4,155.29-2.90%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.17-2.80%
  • EthenaEthena(ENA)$0.263660-4.17%
  • OndoOndo(ONDO)$0.52-4.11%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$117.57-3.70%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Pump.funPump.fun(PUMP)$0.00502111.73%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • aaveAave(AAVE)$147.88-5.41%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Ai2 Researchers are Changing the Benchmarking Game by Introducing Fluid Benchmarking that Enhances Evaluation along Several Dimensions

September 17, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Ai2 Researchers are Changing the Benchmarking Game by Introducing Fluid Benchmarking that Enhances Evaluation along Several Dimensions
ShareShareShareShareShare

A team of researchers from Allen Institute for Artificial Intelligence (Ai2), University of Washington and CMU introduce Fluid Benchmarking, an adaptive LLM evaluation method that replaces static accuracy with 2-parameter IRT ability estimation and Fisher-information–driven item selection. By asking only the most informative questions for a model’s current ability, it yields smoother training curves, delays benchmark saturation, improves external validity at small budgets, and filters mislabeled items.

Fluid Benchmarking replaces static accuracy with an adaptive, psychometrics-grounded procedure. A two-parameter logistic IRT model maps responses to a latent ability score and selects each next item by maximizing Fisher information at the model’s current ability estimate. Across six popular benchmarks and multiple model checkpoints, it improves validity (smaller rank distance), reduces variance (lower normalized total variation), delays saturation (more monotonic training curves), and avoids mislabeled items by ~100× compared to random sampling at equal budget.

YOU MAY ALSO LIKE

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

You Can Now Preorder The Tiny Boox Picco Ereader

What problem does Fluid Benchmarking solve?

Static subsets and plain accuracy conflate item quality and item difficulty, inflate step-to-step variance, and hit benchmark saturation early (training curves flatten while the model still improves). Fluid Benchmarking reframes both aggregation and selection: score in a latent ability space and adapt the item subset to the current ability, rather than treating all items equally or fixing them a priori.

How does it work?

1) Ability, not accuracy

Fit a 2-parameter logistic (2PL) IRT model on historical LM responses: for item j with discrimination aj​ and difficulty bj​, the probability a model with ability θi​ answers correctly is

p(uij​=1)=logistic(aj​(θi​−bj​))

At evaluation, estimate the MAP ability θ^i​ for the candidate LM by maximizing the 2PL likelihood over its observed right/wrong responses on the administered items. Items are weighted by their discrimination and difficulty, unlike accuracy which weights all equally

2) Dynamic item selection via Fisher information

At each step t, select the next item qj​ that maximizes Fisher information at the current ability estimate θ^(t):

I(θi​,aj​,bj​)=aj2​logistic(aj​(θi​−bj​))(1−logistic(aj​(θi​−bj​)))

High-information items minimize the variance of the ability estimate. As training progresses, the most informative items shift from easy to hard, so the administered subset evolves with model capability.

What does “better evaluation” mean here?

Fluid evaluates four dimensions with concrete metrics:

  • Validity: external agreement with “true” model ranking; measured by mean rank distance (lower is better).
  • Variance: normalized total variation of the training curve across checkpoints (lower is better).
  • Saturation: monotonicity (Spearman rank correlation between checkpoint index and predicted performance; higher is better).
  • Efficiency: quality at small item budgets.

How strong are the results?

Across six benchmarks (e.g., ARC-C, GSM8K, HellaSwag, MMLU, TruthfulQA, WinoGrande) and six LMs with 61–94 checkpoints each:

  • Validity: On the smallest subset (AP-10), mean rank distance drops from 20.0 → 10.1; on AP-50, 15.2 → 8.8.
  • Variance: Total variation shrinks markedly; e.g., 28.3 → 10.7 (AP-10) and 19.1 → 6.5 (AP-50).
  • Saturation: Monotonicity improves from 0.48 → 0.76 (AP-10) and 0.62 → 0.86 (AP-50).
  • Small-budget efficiency: With 10 items, Fluid improves mean rank distance by 9.9 vs. random; at 500 items, the improvement is 0.8—consistent with diminishing returns as budget grows.

In pretraining runs, accuracy space often looks flat late in training, but ability space continues to rise, delaying apparent saturation (e.g., HellaSwag monotonicity 0.91 → 0.99 for random vs. Fluid).

Fluid also avoids mislabeled items: on MMLU-Redux with 100-item budgets, mislabeled items per session drop from 0.75 (random) to 0.01 (Fluid)—about two orders of magnitude fewer.

Ablations isolate where the gains come from: IRT aggregation raises validity, but only dynamic selection lowers variance; “RANDOM-IRT” can even exceed random’s variance at large budgets, underscoring selection as the key lever.

Does it stop early when confident?

Yes. Fluid supports dynamic stopping using the standard error of the ability estimate; terminate when SE falls below the average ability gap between rank-adjacent LMs on the Open LLM Leaderboard. In practice, required items vary widely over training (≈20 early, >80 mid-run), showing why fixed budgets are suboptimal.

Where does it fit in the evaluation stack?

Fluid is benchmark-refinement: it does not invent new tasks; it re-weights and re-orders existing items to maximize information against a latent ability metric. It generalizes beyond pretraining to post-training and to other modalities, assuming enough responses to fit/update an IRT model. As models improve, IRT parameters must be refreshed to resolve difficulty among items that were previously “too hard,” otherwise the top of the scale compresses.

Summary

Fluid Benchmarking makes LLM evaluation budget-efficient and stable by scoring models in ability space and selecting items by Fisher information, yielding lower variance, better rank validity, and delayed saturation with far fewer questions. The trade-offs are operational: maintain fresh response matrices, periodically refit IRT parameters, and ensure reliable right/wrong binarization for open-ended tasks. As these practices standardize, Fluid becomes a practical default for in-loop pretraining and post-training evals across evolving benchmarks.


Check out the Paper, GitHub Page and Technical details. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

[Recommended Read] 🧵 NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

🔥[Recommended Read] NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI

Credit: Source link

ShareTweetSendSharePin

Related Posts

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
AI & Technology

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

September 28, 2026
Next Post
Channing Tatum’s extreme weight loss for role raises concerns about increasing eating disorders

Channing Tatum's extreme weight loss for role raises concerns about increasing eating disorders

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
A rare look inside a suspected cell of covert North Korean IT workers

A rare look inside a suspected cell of covert North Korean IT workers

September 21, 2026
Fighting Over The Cost Of Freezer Meals?

Fighting Over The Cost Of Freezer Meals?

September 22, 2026
NFL Week 3 grades for all 32 teams: Ravens get 'B+' for wild Rio win, Raiders earn 'A-' for beating Saints – CBS Sports

NFL Week 3 grades for all 32 teams: Ravens get 'B+' for wild Rio win, Raiders earn 'A-' for beating Saints – CBS Sports

September 28, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!