• bitcoinBitcoin(BTC)$77,053.00-0.42%
  • ethereumEthereum(ETH)$2,489.91-2.04%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$719.01-2.19%
  • rippleXRP(XRP)$1.35-1.56%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$100.41-1.53%
  • tronTRON(TRX)$0.3409140.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.51%
  • zcashZcash(ZEC)$1,098.10-4.46%
  • HyperliquidHyperliquid(HYPE)$78.36-2.69%
  • dogecoinDogecoin(DOGE)$0.083598-1.71%
  • RainRain(RAIN)$0.0153021.45%
  • moneroMonero(XMR)$531.89-0.07%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$79.89-0.74%
  • chainlinkChainlink(LINK)$11.33-2.13%
  • leo-tokenLEO Token(LEO)$9.06-0.66%
  • cardanoCardano(ADA)$0.207416-0.60%
  • stellarStellar(XLM)$0.179606-0.69%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$225.71-2.08%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.310.47%
  • uniswapUniswap(UNI)$6.34-0.66%
  • CantonCanton(CC)$0.095501-2.74%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.73%
  • hedera-hashgraphHedera(HBAR)$0.0764582.52%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.40-0.43%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.43%
  • nearNEAR Protocol(NEAR)$2.31-2.81%
  • suiSui(SUI)$0.72-1.49%
  • crypto-com-chainCronos(CRO)$0.058524-0.10%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,349.480.02%
  • MemeCoreMemeCore(M)$1.16-1.52%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.13-0.70%
  • BittensorBittensor(TAO)$236.460.77%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.02%
  • aaveAave(AAVE)$127.351.24%
  • AsterAster(ASTER)$0.702.45%
  • pax-goldPAX Gold(PAXG)$4,352.48-0.06%
  • BitwayBitway(BTW)$0.6926.94%
  • mantleMantle(MNT)$0.57-1.10%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575160.85%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Overcoming Hallucinations in AI: How Factually Augmented RLHF Optimizes Vision-Language Alignment in Large Multimodal Models

October 7, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Overcoming Hallucinations in AI: How Factually Augmented RLHF Optimizes Vision-Language Alignment in Large Multimodal Models
ShareShareShareShareShare

By additional pre-training using image-text pairings or fine-tuning them with specialized visual instruction tuning datasets, Large Language Models may dive into the multimodal domain, giving rise to potent Large Multimodal Models. However, there are obstacles to building LMMs, chief among them the disparity between the quantity and quality of multimodal data and text-only datasets. Take the LLaVA model, initialized from a pre-trained visual encoder and a language model tweaked for instructions. It is trained on far fewer instances than text-only models, which use over 100M examples over 1800 tasks. It is only trained on 150K artificial image-based conversations. Due to such data restrictions, the visual and language modalities may not be aligned. 

As a result, LMMs could generate hallucinatory outputs that are inaccurately tied to the context that pictures give. Researchers from UC Berkeley, CMU, UIUC, UW–Madison, UMass Amherst Microsoft Research, and MIT-IBM Watson AI Lab present LLaVA-RLHF, a vision-language model trained for enhanced multimodal alignment, to address the issues brought on by the absence of high-quality visual instruction tuning data for LMM training. One of their major contributions is adapting the multimodal alignment for LMMs to the universal and scalable alignment paradigm known as Reinforcement Learning from Human Feedback, which has demonstrated remarkable effectiveness for text-based AI agents. To fine-tune LMM, it collects human preferences focusing on recognizing hallucinations and uses those preferences in reinforcement learning. 

This strategy may improve the multimodal alignment at a relatively cheap annotation cost, such as $3000 for gathering 10K human preferences for image-based discussions. As far as they know, this strategy is the first effective use of RLHF for multimodal alignment. Gaining high ratings from the reward model only sometimes equates to improving human judgments, which is reward hacking. It is a possible problem with the present RLHF paradigm. Previous research suggested iteratively gathering “fresh” human feedback to stop incentive hacking, but this method is typically expensive and cannot properly use existing human preference data. This study suggests a more data-efficient option, attempting to make the reward model capable of using the knowledge and data already present in bigger language models that humans have annotated. 

Figure 1: A diagram illustrating the possibility of hallucinations during the Supervised Fine-Tuning (SFT) phase of LMM training and the way Factually Augmented RLHF addresses the problem of low capacity in the reward model, which is initialized from the SFT model.

First, they use a superior visual encoder with higher resolutions and a bigger language model to enhance the reward model’s overall functionality. Second, they present the Factually Augmented RLHF algorithm, which, as shown in Fig. 1, calibrates the reward signals by supplementing them with extra information like picture descriptions or a ground-truth multi-choice option. They further augment the synthetic vision instruction tuning data with existing high-quality human-annotated multimodal data in the conversation format to enhance the general capabilities of LMMs during the Supervised Fine-Tuning stage. They specifically transform Flickr30k into a Spotting Captioning assignment, VQA-v2, and A-OKVQA into a multi-round QA task, and both train the LLaVA-SFT+ models using the new data set. 

Finally, they consider how to evaluate the multimodal alignment of LMMs in situations of real-world creation, paying particular attention to penalizing any hallucinations. The benchmark questions they develop, MMHAL-BENCH, cover all 12 of COCO’s key object categories and comprise eight job kinds. According to their analysis, this benchmark dataset closely matches human assessments, especially if scores are considered for anti-hallucinations. As the first LMM trained with RLHF, LLaVA-RLHF performs admirably in their experimental assessment. They saw an improvement of 94% on the LLaVA-Bench, a 60% improvement on the MMHAL-BENCH, and they set new performance records for LLaVA with 52.4% on MMBench and 82.7% F1 on POPE. On GitHub, they have made their code, model, and data accessible to the public.


Check out the Paper and Project. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


▶️ Now Watch AI Research Updates On Our Youtube Channel [Watch Now]

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Get Your Cut Of PlayStation’s .85 Million Settlement
AI & Technology

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

September 13, 2026
AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
AI & Technology

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Next Post
Huawei’s Surprise Comeback Marks New Phase in the Tech Cold War

Huawei’s Surprise Comeback Marks New Phase in the Tech Cold War

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

September 10, 2026
Everybody’s Business: Unpacking Apple’s Upcoming Launches

Everybody’s Business: Unpacking Apple’s Upcoming Launches

September 12, 2026
When Are Portable Apple CarPlay Screens Actually Worth It?

When Are Portable Apple CarPlay Screens Actually Worth It?

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!