• bitcoinBitcoin(BTC)$84,071.00-0.45%
  • ethereumEthereum(ETH)$2,675.39-0.52%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$773.930.01%
  • rippleXRP(XRP)$1.532.02%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.491.01%
  • tronTRON(TRX)$0.338211-1.10%
  • zcashZcash(ZEC)$1,566.712.95%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.74%
  • HyperliquidHyperliquid(HYPE)$92.92-0.63%
  • dogecoinDogecoin(DOGE)$0.0953801.07%
  • moneroMonero(XMR)$570.131.80%
  • chainlinkChainlink(LINK)$13.578.84%
  • whitebitWhiteBIT Coin(WBT)$83.80-0.83%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2492933.07%
  • RainRain(RAIN)$0.011928-1.77%
  • leo-tokenLEO Token(LEO)$8.81-2.14%
  • stellarStellar(XLM)$0.2169936.48%
  • bitcoin-cashBitcoin Cash(BCH)$332.75-1.74%
  • nearNEAR Protocol(NEAR)$4.534.68%
  • uniswapUniswap(UNI)$9.13-0.93%
  • litecoinLitecoin(LTC)$71.102.71%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.1168886.30%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.19-0.65%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.024.92%
  • hedera-hashgraphHedera(HBAR)$0.0921730.97%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-1.14%
  • shiba-inuShiba Inu(SHIB)$0.0000060.78%
  • BittensorBittensor(TAO)$297.702.21%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0644893.15%
  • OndoOndo(ONDO)$0.5730.42%
  • BitwayBitway(BTW)$1.01-1.48%
  • MemeCoreMemeCore(M)$1.19-4.52%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,274.22-0.15%
  • okbOKB(OKB)$119.49-0.52%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.22%
  • EthenaEthena(ENA)$0.2214817.09%
  • aaveAave(AAVE)$144.082.90%
  • mantleMantle(MNT)$0.67-3.43%
  • MorphoMorpho(MORPHO)$2.864.45%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Google AI Introduces FLAMe: A Foundational Large Autorater Model for Reliable and Efficient LLM Evaluation

July 20, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper from Google AI Introduces FLAMe: A Foundational Large Autorater Model for Reliable and Efficient LLM Evaluation
ShareShareShareShareShare

Evaluating large language models (LLMs) has become increasingly challenging due to their complexity and versatility. Ensuring the reliability and quality of these models’ outputs is crucial for advancing AI technologies and applications. Researchers need help developing reliable evaluation methods to assess the accuracy and impartiality of LLMs’ outputs, given human evaluations’ subjective, inconsistent, and costly nature.

Current evaluation metrics like BLEU and ROUGE mainly focus on lexical overlaps and fail to capture the nuanced quality of LLM outputs. Although recent methods have utilized pretrained models to measure distributional similarity and token probabilities, these approaches still need to be revised in generalizability and consistency. The high cost and time required for human evaluations further complicate the process, making it impractical for large-scale assessments.

YOU MAY ALSO LIKE

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

A research team from Google DeepMind, Google, and UMass Amherst have introduced FLAMe, a family of Foundational Large Autorater Models designed to improve the evaluation of LLMs. FLAMe leverages a large and diverse collection of quality assessment tasks derived from human judgments to train and standardize autoraters. FLAMe is trained using supervised multitask fine-tuning on over 100 quality assessment tasks, encompassing more than 5 million human judgments. This training employs a text-to-text format, facilitating effective transfer learning across functions. The approach enables FLAMe to generalize to new tasks, outperforming existing models like GPT-4 and Claude-3.

The training of FLAMe involves a meticulous process of data collection and standardization. The research team curated human evaluations from previous studies, focusing on tasks such as machine translation quality and AI assistant instruction. This extensive dataset was then reformatted into a unified text-to-text format, where each quality assessment task was converted into input-target pairs. The inputs include task-specific contexts, while the targets contain expected human evaluations. By training on this large and diverse dataset, FLAMe learns robust patterns of human judgment, minimizing the impact of noisy or low-quality data. The FLAMe-RM variant, specifically fine-tuned for reward modeling evaluation, exemplifies this methodology’s effectiveness. Fine-tuned for only 50 steps on a mixture of four datasets covering chat, reasoning, and safety, FLAMe-RM demonstrates significant improvements in performance.

The performance of FLAMe is noteworthy across various benchmarks. The FLAMe-RM-24B model, a variant fine-tuned for reward modeling evaluation, achieved an accuracy of 87.8% on RewardBench, surpassing both GPT-4-0125 (85.9%) and GPT-4o (84.7%). On the CoBBLEr bias benchmark, FLAMe exhibits significantly lower bias compared to other autorater models. In addition to RewardBench, FLAMe’s performance is strong on other benchmarks. The FLAMe models outperform existing LLMs on 8 out of 12 automated evaluation benchmarks, covering 53 quality assessment tasks. This includes tasks such as summary comparisons, helpfulness evaluations, and factual accuracy assessments. The results demonstrate FLAMe’s broad applicability and robust performance across diverse evaluation scenarios.

FLAMe-Opt-RM, a computationally efficient variant, optimizes the multitask mixture for reward modeling evaluation using a novel tail-patch fine-tuning strategy. This method fine-tunes the initial instruction-tuned PaLM-2-24B checkpoint on an optimized mixture for 5000 steps, achieving competitive RewardBench performance with approximately 25 times fewer training data points. The research highlights that longer training and additional fine-tuning can improve performance, suggesting that FLAMe-Opt-RM is a versatile and efficient model.

To conclude, the research highlights the importance of reliable and efficient evaluation methods for LLMs. FLAMe offers a robust solution by leveraging standardized human evaluations, demonstrating significant improvements in performance and bias reduction. This advancement is poised to enhance the development and deployment of AI technologies. The FLAMe family of models, developed by a collaborative team from Google DeepMind, Google, and UMass Amherst, represents a significant step forward in evaluating large language models, ensuring their outputs are reliable, unbiased, and of high quality.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
AI & Technology

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

September 25, 2026
Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120
AI & Technology

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

September 25, 2026
Warzone Is Adding A Button To Hide All The Goofy Skins
AI & Technology

Warzone Is Adding A Button To Hide All The Goofy Skins

September 24, 2026
How These AI Glasses Compare
AI & Technology

How These AI Glasses Compare

September 24, 2026
Next Post
Andreessen, Horowitz Latest Tech VCs to Support Trump

Andreessen, Horowitz Latest Tech VCs to Support Trump

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The Street Fighter Movie Popcorn Bucket Is Gloriously Goofy

The Street Fighter Movie Popcorn Bucket Is Gloriously Goofy

September 18, 2026
Rescued mountain climber speaks out after daring 15-hour mission

Rescued mountain climber speaks out after daring 15-hour mission

September 21, 2026
Top Story with Tom Llamas – Aug. 24 | NBC News NOW

Top Story with Tom Llamas – Aug. 24 | NBC News NOW

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!