• bitcoinBitcoin(BTC)$76,621.001.06%
  • ethereumEthereum(ETH)$2,445.101.85%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$725.872.33%
  • rippleXRP(XRP)$1.311.40%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.623.61%
  • tronTRON(TRX)$0.3349610.06%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.64%
  • zcashZcash(ZEC)$1,372.5615.36%
  • HyperliquidHyperliquid(HYPE)$80.633.85%
  • dogecoinDogecoin(DOGE)$0.0812451.94%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$494.00-1.57%
  • whitebitWhiteBIT Coin(WBT)$78.771.06%
  • RainRain(RAIN)$0.012525-8.96%
  • chainlinkChainlink(LINK)$11.173.39%
  • leo-tokenLEO Token(LEO)$8.930.43%
  • cardanoCardano(ADA)$0.1996893.07%
  • stellarStellar(XLM)$0.1826424.37%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.00-0.03%
  • bitcoin-cashBitcoin Cash(BCH)$225.003.10%
  • USD1USD1(USD1)$1.00-0.01%
  • uniswapUniswap(UNI)$6.929.99%
  • litecoinLitecoin(LTC)$52.774.00%
  • CantonCanton(CC)$0.10106111.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.320.62%
  • nearNEAR Protocol(NEAR)$2.8217.10%
  • avalanche-2Avalanche(AVAX)$7.564.18%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.074032-0.39%
  • shiba-inuShiba Inu(SHIB)$0.0000054.00%
  • suiSui(SUI)$0.725.21%
  • crypto-com-chainCronos(CRO)$0.0579074.31%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.05%
  • tether-goldTether Gold(XAUT)$4,304.49-0.64%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$227.846.01%
  • MemeCoreMemeCore(M)$1.132.60%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.001.82%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.24%
  • AsterAster(ASTER)$0.739.26%
  • BitwayBitway(BTW)$0.71-8.27%
  • aaveAave(AAVE)$123.053.19%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0589203.38%
  • pax-goldPAX Gold(PAXG)$4,307.25-0.68%
  • mantleMantle(MNT)$0.562.34%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

LMSYS ORG Introduces Arena-Hard: A Data Pipeline to Build High-Quality Benchmarks from Live Data in Chatbot Arena, which is a Crowd-Sourced Platform for LLM Evals

April 28, 2024
in AI & Technology
Reading Time: 4 mins read
A A
LMSYS ORG Introduces Arena-Hard: A Data Pipeline to Build High-Quality Benchmarks from Live Data in Chatbot Arena, which is a Crowd-Sourced Platform for LLM Evals
ShareShareShareShareShare

In Large language models(LLM), developers and researchers face a significant challenge in accurately measuring and comparing the capabilities of different chatbot models. A good benchmark for evaluating these models should accurately reflect real-world usage, distinguish between different models’ abilities, and regularly update to incorporate new data and avoid biases.

Traditionally, benchmarks for large language models, such as multiple-choice question-answering systems, have been static. These benchmarks do not frequently update and fail to capture real-world application nuances. They also may not effectively demonstrate the differences between more closely performing models, which is crucial for developers aiming to improve their systems.

‘Arena-Hard‘ has been developed by LMSYS ORG to address these shortcomings. This system creates benchmarks from live data collected from a platform where users continuously evaluate large language models. This method ensures the benchmarks are up-to-date and rooted in fundamental user interactions, providing a more dynamic and relevant evaluation tool.

To adapt this for real-world benchmarking of LLMs:

  1. Continuously Update the Predictions and Reference Outcomes: As new data or models become available, the benchmark should update its predictions and recalibrate based on actual performance outcomes.
  2. Incorporate a Diversity of Model Comparisons: Ensure a wide range of model pairs is considered to capture various capabilities and weaknesses.
  3. Transparent Reporting: Regularly publish details on the benchmark’s performance, prediction accuracy, and areas for improvement.

The effectiveness of Arena-Hard is measured by two primary metrics: its ability to agree with human preferences and its capacity to separate different models based on their performance. Compared with existing benchmarks, Arena-Hard showed significantly better performance in both metrics. It demonstrated a high agreement rate with human preferences. It proved more capable of distinguishing between top-performing models, with a notable percentage of model comparisons having precise, non-overlapping confidence intervals.

In conclusion, Arena-Hard represents a significant advancement in benchmarking language model chatbots. By leveraging live user data and focusing on metrics that reflect both human preferences and clear separability of model capabilities, this new benchmark provides a more accurate, reliable, and relevant tool for developers. This can drive the development of more effective and nuanced language models, ultimately enhancing user experience across various applications.


Check out the GitHub page and Blog. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

An iOS 27 Bug Can Temporarily Freeze Your iPhone

Niharika is a Technical consulting intern at Marktechpost. She is a third year undergraduate, currently pursuing her B.Tech from Indian Institute of Technology(IIT), Kharagpur. She is a highly enthusiastic individual with a keen interest in Machine learning, Data science and AI and an avid reader of the latest developments in these fields.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
AI & Technology

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

September 17, 2026
An iOS 27 Bug Can Temporarily Freeze Your iPhone
AI & Technology

An iOS 27 Bug Can Temporarily Freeze Your iPhone

September 17, 2026
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Next Post
Valley National Bank: Q1 Results Should Calm Fears (NASDAQ:VLY)

Valley National Bank: Q1 Results Should Calm Fears (NASDAQ:VLY)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Man accused of stealing gold bars working towards plea deal

Man accused of stealing gold bars working towards plea deal

September 12, 2026
‘There was so much chaos’: 9/11 survivor reflects on the day of the attacks

‘There was so much chaos’: 9/11 survivor reflects on the day of the attacks

September 12, 2026
Canoodling lawyers’ firm rocked by another major defection as top attorney bolts

Canoodling lawyers’ firm rocked by another major defection as top attorney bolts

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!