• bitcoinBitcoin(BTC)$76,987.00-1.28%
  • ethereumEthereum(ETH)$2,464.14-0.17%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$712.83-0.67%
  • rippleXRP(XRP)$1.34-2.98%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.32-1.80%
  • tronTRON(TRX)$0.338255-0.61%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.80%
  • zcashZcash(ZEC)$1,100.90-10.07%
  • HyperliquidHyperliquid(HYPE)$79.22-4.39%
  • dogecoinDogecoin(DOGE)$0.083504-2.05%
  • RainRain(RAIN)$0.015664-3.30%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$508.770.48%
  • whitebitWhiteBIT Coin(WBT)$79.73-1.05%
  • chainlinkChainlink(LINK)$11.39-3.40%
  • leo-tokenLEO Token(LEO)$9.04-1.70%
  • cardanoCardano(ADA)$0.203115-4.43%
  • stellarStellar(XLM)$0.174528-2.78%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$224.92-8.53%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$52.30-0.12%
  • CantonCanton(CC)$0.095033-6.18%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.50%
  • uniswapUniswap(UNI)$5.99-0.59%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.38-4.62%
  • hedera-hashgraphHedera(HBAR)$0.073653-3.27%
  • nearNEAR Protocol(NEAR)$2.472.08%
  • suiSui(SUI)$0.73-4.37%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.94%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056354-0.05%
  • MemeCoreMemeCore(M)$1.18-1.50%
  • tether-goldTether Gold(XAUT)$4,344.73-0.84%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$113.971.75%
  • BittensorBittensor(TAO)$232.48-7.72%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • mantleMantle(MNT)$0.58-2.08%
  • aaveAave(AAVE)$122.16-0.90%
  • pax-goldPAX Gold(PAXG)$4,349.12-0.78%
  • AsterAster(ASTER)$0.69-3.10%
  • polkadotPolkadot(DOT)$1.09-1.17%
  • OndoOndo(ONDO)$0.346661-1.97%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet ‘DRESS’: A Large Vision Language Model (LVLM) that Align and Interact with Humans via Natural Language Feedback

December 1, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Meet ‘DRESS’: A Large Vision Language Model (LVLM) that Align and Interact with Humans via Natural Language Feedback
ShareShareShareShareShare

Big vision-language models, or LVLMs, can interpret visual cues and provide easy replies for users to interact with. This is accomplished by skillfully fusing large language models (LLMs) with large-scale visual instruction finetuning. Nevertheless, LVLMs only need hand-crafted or LLM-generated datasets for alignment by supervised fine-tuning (SFT). Although it works well to change LVLMs from caption generators to models that obey instructions, LVLMs can still produce replies that are hurtful, ill-intentioned, or useless. This suggests that they still need to be more aligned with human preferences. Furthermore, while previous research encourages the organization of visual instruction tuning samples in multi-turn forms, the LVLMs’ capacity to interact is limited by the weak connections and interdependence between different turns. Here, the interaction ability assesses how well LVLMs can adjust their replies using the prior context in multi-turn interactions. These two drawbacks limit the practical use of LVLMs as visual helpers. 

The research team from  SRI International and the University of Illinois Urbana-Champaign presents DRESS, an LVLM that is uniquely taught using Natural Language Feedback (NLF) produced by LLMs in this work (refer to Figure 1). The research team instructs LLMs to provide fine-grained feedback on the LVLM’s replies by providing them with specific rules and extensive photo annotation. In keeping with the process of creating human-aligned LLMs, this feedback annotation considers the three H criteria: helpfulness, honesty, and harmlessness. The feedback measures the replies’ overall quality along the 3H criteria and provides a numerical score and NLF. The research team’s method divides NLF into critique and refining. This is a novel classification. While the refinement NLF offers precise recommendations to LVLMs on improving their replies to align with the ground truth reference, the critique NLF evaluates the responses’ strengths and faults. This classification provides a natural application of two kinds of NLF to make LVLMs more palatable to humans and enhance their interaction capabilities. 

Figure 1: Researchers direct DRESS to use natural language input, which is divided into two categories, critique and refinement, to enhance both alignment with human preferences and interaction capacity.

The research team generalizes the conditional reinforcement learning technique to meet the non-differentiable character of NLF and trains the LVLMs with such feedback. Specifically, the research team uses linguistic modeling (LM) loss on the replies to train DRESS to generate equivalent responses conditioned on the two NLFs. The research team refines DRESS by analyzing and interpreting the numerical results to match user preferences better. Through multi-turn interactions during inference, the research team trains DRESS to learn the meta-skill of refining its original replies by employing refinement NLF. 

The research team assesses DRESS on multi-turn interactions, adversarial prompting for harmlessness assessment, picture captioning for honesty assessment, and open-ended visual question responding for helpfulness evaluation. The experiments’ findings show that, compared to earlier LVLMs, DRESS can provide replies that align with human values and have superior interaction capabilities that allow it to learn from feedback and modify responses as needed efficiently. To their knowledge, the research team’s effort is the first to address the interaction ability and all three 3H criteria for LVLMs. 

The research team’s contributions are summed up as follows: 

• The research team suggests using natural language feedback (NLF), which may be divided into critique and refining NLF, to enhance LVLMs’ ability to interact and align with human preferences. 

• By training the model to provide matching responses conditioned on the NLF, the research team generalizes the conditional reinforcement learning method to accommodate the non-differentiable NLF successfully. Compared to the previous SOTA, the research team’s suggested model, DRESS, demonstrates relative improvements of 9.76%, 11.52%, and 21.03% based on a systematic evaluation of helpfulness, honesty, and harmlessness alignment. 

• The research group generates and makes 63K annotated language NLF examples available for public use, including 3H characteristics. Furthermore, the research team created a publicly available dataset of 4.7K samples for harmlessness alignment and LVLM assessment. 


Check out the Paper and Dataset. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

How These XL Phones Compete

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


[SPONSORED] Step by Step Tutorial on ‘How to Build LLM Apps that can See Hear Speak’

Credit: Source link

ShareTweetSendSharePin

Related Posts

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
AI & Technology

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

September 11, 2026
How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
Next Post
Full Lindsey Graham: ‘Every death, going forward, I blame on Hamas, not Israel’

Full Lindsey Graham: 'Every death, going forward, I blame on Hamas, not Israel'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How the Nepal-Tibet disaster became the latest victim of China’s ‘Clean Internet’ campaign – The Guardian

How the Nepal-Tibet disaster became the latest victim of China’s ‘Clean Internet’ campaign – The Guardian

September 8, 2026
Term Life Insurance Is Almost Always the Right Choice Over Whole Life

Term Life Insurance Is Almost Always the Right Choice Over Whole Life

September 7, 2026
Expedition to find Amelia Earhart’s plane set to begin

Expedition to find Amelia Earhart’s plane set to begin

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!