• bitcoinBitcoin(BTC)$75,915.00-1.40%
  • ethereumEthereum(ETH)$2,403.98-2.96%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$710.26-1.07%
  • rippleXRP(XRP)$1.29-8.08%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$97.26-3.64%
  • tronTRON(TRX)$0.334845-1.23%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,212.196.08%
  • HyperliquidHyperliquid(HYPE)$77.75-1.94%
  • dogecoinDogecoin(DOGE)$0.079887-3.43%
  • RainRain(RAIN)$0.0137703.36%
  • USDSUSDS(USDS)$1.00-0.03%
  • moneroMonero(XMR)$502.28-3.15%
  • whitebitWhiteBIT Coin(WBT)$78.03-2.07%
  • leo-tokenLEO Token(LEO)$8.88-1.23%
  • chainlinkChainlink(LINK)$10.84-4.93%
  • cardanoCardano(ADA)$0.194496-5.11%
  • stellarStellar(XLM)$0.175710-9.50%
  • Ethena USDeEthena USDe(USDE)$1.00-0.06%
  • daiDai(DAI)$1.000.04%
  • bitcoin-cashBitcoin Cash(BCH)$218.74-1.64%
  • USD1USD1(USD1)$1.00-0.03%
  • uniswapUniswap(UNI)$6.36-4.41%
  • litecoinLitecoin(LTC)$50.91-3.20%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-2.11%
  • CantonCanton(CC)$0.090996-4.79%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074441-4.08%
  • avalanche-2Avalanche(AVAX)$7.28-3.26%
  • nearNEAR Protocol(NEAR)$2.431.09%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.71%
  • suiSui(SUI)$0.69-3.14%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.055649-3.57%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,331.841.42%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-1.01%
  • BittensorBittensor(TAO)$216.25-4.80%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$110.26-2.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.05%
  • BitwayBitway(BTW)$0.788.51%
  • pax-goldPAX Gold(PAXG)$4,336.901.48%
  • aaveAave(AAVE)$119.66-6.29%
  • AsterAster(ASTER)$0.67-2.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056907-0.52%
  • mantleMantle(MNT)$0.54-2.69%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Prometheus 2: An Open Source Language Model that Closely Mirrors Human and GPT-4 Judgements in Evaluating Other Language Models

May 5, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Prometheus 2: An Open Source Language Model that Closely Mirrors Human and GPT-4 Judgements in Evaluating Other Language Models
ShareShareShareShareShare

Natural Language Processing (NLP) seeks to enable computers to comprehend and interact using human language. A critical challenge in NLP is evaluating language models (LMs), which generate responses across various tasks. The diversity of these tasks makes it difficult to assess the quality of responses effectively. With the increasing sophistication of LMs, such as GPT-4, proprietary models often provide strong evaluation capabilities but suffer from transparency, control, and cost issues. This necessitates the development of reliable open-source alternatives that can effectively judge language outputs without compromising on these aspects.

The problem is multifaceted, involving the evaluation of responses and the scalability of evaluation mechanisms. Existing evaluation tools, particularly open-source models, have several limitations. Many models fail to provide direct assessment and pairwise ranking functionalities, the two most prevalent evaluation forms. This limits their adaptability to diverse real-life scenarios. They prioritize general attributes like helpfulness and harmlessness while issuing scores that significantly diverge from human evaluations. This inconsistency leads to unreliable assessments and requires improved evaluator models that closely mirror human judgments.

Research teams have attempted to address these gaps through various methods. However, most approaches lack comprehensive flexibility and fail to simulate human assessments accurately. Current proprietary models like GPT-4 remain expensive and non-transparent, which impedes widespread evaluation usage. The research team from KAIST AI, LG AI Research, Carnegie Mellon University, MIT, Allen Institute for AI, and the University of Illinois Chicago introduced Prometheus 2, a novel open-source evaluator designed to assess language models to resolve it. This model was developed to provide transparent, scalable, and controllable assessments while matching the evaluation quality of proprietary models.

Prometheus 2 was developed by merging two evaluator LMs: one trained exclusively for direct assessment and another for pairwise ranking. The merging of these models created a unified evaluator that excels in both evaluation formats. The researchers utilized the newly developed Preference Collection dataset, which features 1,000 evaluation criteria, to refine the model’s capabilities further. By effectively combining the two training formats, Prometheus 2 can evaluate LM responses using direct assessment and pairwise ranking methods. The merged model leverages a linear merging approach to blend the strengths of both evaluation formats, achieving high performance across evaluation tasks.

The model demonstrated the highest correlation with human and proprietary evaluators in benchmarking tests on four direct assessment benchmarks: Vicuna Bench, MT Bench, FLASK, and Feedback Bench. Pearson correlations exceeded 0.5 on all benchmarks, reaching 0.878 and 0.898 on the Feedback Bench for the 7B and 8x7B models, respectively. On four pairwise ranking benchmarks, including HHH Alignment, MT Bench Human Judgment, Auto-J Eval, and Preference Bench, Prometheus 2 outperformed existing open-source models, achieving accuracy scores surpassing 85%. The Preference Bench, an in-domain test set for Prometheus 2, indicated the model’s robustness and versatility.

Prometheus 2 narrowed the performance gap with proprietary evaluators, such as GPT-4, across various benchmarks. The model halved the correlation difference between humans and GPT-4 on the FLASK benchmark and achieved 84% accuracy in HHH Alignment evaluations. This highlights the significant potential of open-source evaluators to replace expensive proprietary solutions while ensuring comprehensive and accurate assessments.

In conclusion, the lack of transparent, scalable, and adaptable language model evaluators closely reflecting human judgment is a significant challenge in NLP. Researchers developed Prometheus 2, a novel open-source evaluator, to address it. They used a linear merging approach, combining two models trained separately on direct assessment and pairwise ranking. This unified model surpassed previous open-source models in benchmarking tests, showcasing high accuracy and correlation while substantially closing the performance gap with proprietary models. Prometheus 2 represents a significant advancement in open-source evaluation, offering a robust alternative to proprietary solutions.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 41k+ ML SubReddit


YOU MAY ALSO LIKE

Roblox Pushes Deeper Into AI-Powered Gaming

AI Leaders Debate Slowing the Frontier

Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.


✅ [FREE AI WEBINAR Alert] Using AWS Bedrock & LangChain for Private LLM App Dev: May 6, 2024 10:00am – 11:00am PDT


Credit: Source link

ShareTweetSendSharePin

Related Posts

Roblox Pushes Deeper Into AI-Powered Gaming
AI & Technology

Roblox Pushes Deeper Into AI-Powered Gaming

September 16, 2026
AI Leaders Debate Slowing the Frontier
AI & Technology

AI Leaders Debate Slowing the Frontier

September 16, 2026
Can Independent Testing Make AI Safer?
AI & Technology

Can Independent Testing Make AI Safer?

September 16, 2026
AI Needs a Full Stop, Not a Slowdown: Parmy Olson
AI & Technology

AI Needs a Full Stop, Not a Slowdown: Parmy Olson

September 16, 2026
Next Post
Amy Brown, wife of GOP Senate candidate Sam Brown, opens up about her abortion for the first time

Amy Brown, wife of GOP Senate candidate Sam Brown, opens up about her abortion for the first time

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI

Salesforce Debuts Koa Reasoning Model for Agentforce, Trained on Nemotron – Unite.AI

September 15, 2026
Considering A Level 2 EV Charger? How To Know If You Need One

Considering A Level 2 EV Charger? How To Know If You Need One

September 16, 2026
LA’s Boyle Heights vegan restaurant to close after 16 years

LA’s Boyle Heights vegan restaurant to close after 16 years

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!