• bitcoinBitcoin(BTC)$86,276.000.10%
  • ethereumEthereum(ETH)$2,745.68-0.06%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$787.32-1.99%
  • rippleXRP(XRP)$1.574.37%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.97-1.54%
  • tronTRON(TRX)$0.341591-1.05%
  • zcashZcash(ZEC)$1,548.641.64%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.020.00%
  • HyperliquidHyperliquid(HYPE)$94.780.44%
  • dogecoinDogecoin(DOGE)$0.1000092.55%
  • moneroMonero(XMR)$574.37-0.05%
  • whitebitWhiteBIT Coin(WBT)$86.68-0.02%
  • chainlinkChainlink(LINK)$13.01-0.56%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013437-5.02%
  • cardanoCardano(ADA)$0.2513802.32%
  • leo-tokenLEO Token(LEO)$8.972.18%
  • stellarStellar(XLM)$0.2137271.42%
  • bitcoin-cashBitcoin Cash(BCH)$322.9321.88%
  • nearNEAR Protocol(NEAR)$4.427.34%
  • uniswapUniswap(UNI)$9.081.96%
  • avalanche-2Avalanche(AVAX)$11.13-0.53%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • litecoinLitecoin(LTC)$61.57-2.09%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.115808-1.04%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0973235.99%
  • suiSui(SUI)$1.01-3.05%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.44-0.38%
  • BittensorBittensor(TAO)$317.4910.55%
  • shiba-inuShiba Inu(SHIB)$0.0000061.56%
  • crypto-com-chainCronos(CRO)$0.0668543.36%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.31-12.68%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,335.15-0.31%
  • okbOKB(OKB)$121.65-1.98%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BitwayBitway(BTW)$0.87-4.69%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • aaveAave(AAVE)$143.16-2.02%
  • mantleMantle(MNT)$0.662.46%
  • EthenaEthena(ENA)$0.209592-5.80%
  • OndoOndo(ONDO)$0.429511-4.16%
  • pepePepe(PEPE)$0.0000052.24%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta presents Self-Taught Evaluators: A New AI Approach that Aims to Improve Evaluators without Human Annotations and Outperforms Commonly Used LLM Judges Such as GPT-4

August 7, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Meta presents Self-Taught Evaluators: A New AI Approach that Aims to Improve Evaluators without Human Annotations and Outperforms Commonly Used LLM Judges Such as GPT-4
ShareShareShareShareShare

Advancements in NLP have led to the development of large language models (LLMs) capable of performing complex language-related tasks with high accuracy. These advancements have opened up new possibilities in technology and communication, allowing for more natural and effective human-computer interactions.

A significant problem in NLP is the reliance on human annotations for model evaluation. Human-generated data is essential for training and validating models, but collecting this data is both costly and time-consuming. Furthermore, as models improve, previously collected annotations may need to be updated, reducing their utility in evaluating newer models. This creates a continuous need for fresh data, which poses challenges for scaling and sustaining effective model evaluations. Addressing this problem is crucial for advancing NLP technologies and their applications.

YOU MAY ALSO LIKE

How To Enter VR Mode On Steam

Peloton Has Made A Foldable (Treadmill)

Current methods for model evaluation typically involve collecting large amounts of human preference judgments over model responses. These methods include using automated metrics for tasks with reference answers or employing classifiers that output scores directly. However, these methods face limitations, especially for complex tasks where multiple valid responses are possible, such as creative writing or coding. The high variance in human judgments and the associated costs highlight the need for more efficient and scalable evaluation techniques.

Researchers at Meta FAIR have introduced a novel approach called the “Self-Taught Evaluator.” This method eliminates the need for human annotations by using synthetically generated data for training. The process begins with a seed model, which produces contrasting synthetic preference pairs. The model then evaluates these pairs and improves iteratively, using its judgments to enhance its performance in subsequent iterations. This approach leverages the model’s capability to generate and evaluate data, significantly reducing dependency on human-generated annotations.

The proposed method involves several key steps. Initially, a baseline response is generated for a given instruction using a seed LLM. A modified version of the instruction is then created, prompting the LLM to generate a new response designed to be lower quality than the original. These paired responses form the basis for training data. The model, acting as an LLM-as-a-Judge, generates reasoning traces and judgments for these pairs. This process is repeated iteratively, with the model continually improving its judgment accuracy through self-generated and self-evaluated data, effectively creating a cycle of self-improvement.

The performance of the Self-Taught Evaluator was tested using the Llama-3-70B-Instruct model. The method improved the model’s accuracy on the RewardBench benchmark from 75.4 to 88.7, matching or surpassing the performance of models trained with human annotations. This significant improvement demonstrates the effectiveness of synthetic data in enhancing model evaluation. Furthermore, the researchers conducted multiple iterations, further refining the model’s capabilities. The final model achieved 88.3 accuracy with a single inference and 88.7 with majority voting, showcasing its robustness and reliability.

In conclusion, the Self-Taught Evaluator offers a scalable and efficient NLP model evaluation solution. By leveraging synthetic data and iterative self-improvement, it addresses the challenges of relying on human annotations and keeps pace with the rapid advancements in language model development. This approach enhances model performance and reduces the dependency on human-generated data, paving the way for more autonomous and efficient NLP systems. The research team’s work at Meta FAIR marks a significant step forward in the quest for more advanced and autonomous evaluation methods in the field of NLP.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here



Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
Next Post
Opening statements set to begin in Trump criminal trial

Opening statements set to begin in Trump criminal trial

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
This Morning’s Top Headlines – Aug. 31 | Morning News NOW

This Morning’s Top Headlines – Aug. 31 | Morning News NOW

September 20, 2026
Grand Canyon Rescue Operations; How to Quell Settler Violence in the West Bank | Sept. 1

Grand Canyon Rescue Operations; How to Quell Settler Violence in the West Bank | Sept. 1

September 19, 2026
Trump calls war with Iran ‘small potatoes’ for U.S.

Trump calls war with Iran ‘small potatoes’ for U.S.

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!