• bitcoinBitcoin(BTC)$86,839.001.63%
  • ethereumEthereum(ETH)$2,775.771.66%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$796.740.99%
  • rippleXRP(XRP)$1.626.83%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.302.38%
  • tronTRON(TRX)$0.344509-0.86%
  • zcashZcash(ZEC)$1,617.4610.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.77%
  • HyperliquidHyperliquid(HYPE)$97.074.65%
  • dogecoinDogecoin(DOGE)$0.1035943.37%
  • moneroMonero(XMR)$569.67-0.89%
  • whitebitWhiteBIT Coin(WBT)$87.351.59%
  • chainlinkChainlink(LINK)$13.191.93%
  • cardanoCardano(ADA)$0.2575924.51%
  • USDSUSDS(USDS)$1.00-0.02%
  • RainRain(RAIN)$0.013178-4.42%
  • leo-tokenLEO Token(LEO)$8.980.18%
  • stellarStellar(XLM)$0.2208703.75%
  • bitcoin-cashBitcoin Cash(BCH)$340.4128.19%
  • uniswapUniswap(UNI)$10.7218.38%
  • nearNEAR Protocol(NEAR)$4.420.37%
  • avalanche-2Avalanche(AVAX)$11.323.05%
  • litecoinLitecoin(LTC)$63.715.07%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.115446-1.40%
  • hedera-hashgraphHedera(HBAR)$0.1010998.41%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.03-0.50%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.472.08%
  • shiba-inuShiba Inu(SHIB)$0.0000063.19%
  • BittensorBittensor(TAO)$316.860.82%
  • crypto-com-chainCronos(CRO)$0.0688852.41%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.30-9.09%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,341.34-0.07%
  • okbOKB(OKB)$124.371.96%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.909.11%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • aaveAave(AAVE)$152.556.97%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • mantleMantle(MNT)$0.696.65%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.49%
  • EthenaEthena(ENA)$0.2191933.60%
  • OndoOndo(ONDO)$0.4467532.10%
  • pepePepe(PEPE)$0.000005-3.86%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

UCLA Researchers Released OpenVLThinker-7B: A Reinforcement Learning Driven Model for Enhancing Complex Visual Reasoning and Step-by-Step Problem Solving in Multimodal Systems

March 29, 2025
in AI & Technology
Reading Time: 5 mins read
A A
UCLA Researchers Released OpenVLThinker-7B: A Reinforcement Learning Driven Model for Enhancing Complex Visual Reasoning and Step-by-Step Problem Solving in Multimodal Systems
ShareShareShareShareShare

Large vision-language models (LVLMs) integrate large language models with image processing capabilities, enabling them to interpret images and generate coherent textual responses. While they excel at recognizing visual objects and responding to prompts, they often falter when presented with problems requiring multi-step reasoning. Vision-language tasks like understanding charts, solving visual math questions, or interpreting diagrams demand more than recognition; they need the ability to follow logical steps based on visual cues. Despite advancements in model architecture, current systems consistently struggle to produce accurate and interpretable answers in such complex scenarios.

A major limitation in current vision-language models is their inability to perform complex reasoning that involves multiple steps of logical deduction, especially when interpreting images in conjunction with textual queries. These models cannot often internally verify or correct their reasoning, leading to incorrect or shallow outputs. Also, the reasoning chains these models follow are typically not transparent or verifiable, making it difficult to ensure the robustness of their conclusions. The challenge lies in bridging this reasoning gap, which text-only models have begun to address effectively through reinforcement learning techniques but vision-language models have yet to embrace fully.

YOU MAY ALSO LIKE

The Pros And Cons Of Using A Password Manager Over An Authenticator App

How To Hide Or Replace The Audio Button In iMessages

Before this study, efforts to enhance reasoning in such systems mostly relied on standard fine-tuning or prompting techniques. Though helpful in basic tasks, these approaches often resulted in verbose or repetitive outputs with limited depth. Vision-language models like Qwen2.5-VL-7B showed promise due to their visual instruction-following abilities but lacked the multi-step reasoning comparable to their text-only counterparts, such as DeepSeek-R1. Even when prompted with structured queries, these models struggled to reflect upon their outputs or validate intermediate reasoning steps. This was a significant bottleneck, particularly for use cases requiring structured decision-making, such as visual problem-solving or educational support tools.

Researchers from the University of California, Los Angeles, introduced a model named OpenVLThinker-7B. This model was developed through a novel training method that combines supervised fine-tuning (SFT) and reinforcement learning (RL) in an iterative loop. The process started by generating image captions using Qwen2.5-VL-3B and feeding these into a distilled version of DeepSeek-R1 to produce structured reasoning chains. These outputs formed the training data for the first round of SFT, guiding the model in learning basic reasoning structures. Following this, a reinforcement learning stage using Group Relative Policy Optimization (GRPO) was applied to refine the model’s reasoning based on reward feedback. This combination enabled the model to progressively self-improve, using each iteration’s refined outputs as new training data for the next cycle.

The method involved careful data curation and multiple training phases. In the first iteration, 25,000 examples were used for SFT, sourced from datasets like FigureQA, Geometry3K, TabMWP, and VizWiz. These examples were filtered to remove overly verbose or redundant reflections, improving training quality. GRPO was then applied to a smaller, more difficult dataset of 5,000 samples. This led to a performance increase from 62.5% to 65.6% accuracy on the MathVista benchmark. In the second iteration, another 5,000 high-quality examples were used for SFT, raising accuracy to 66.1%. A second round of GRPO pushed performance to 69.4%. Across these phases, the model was evaluated on multiple benchmarks, MathVista, MathVerse, and MathVision, showing consistent performance gains with each iteration.

Quantitatively, OpenVLThinker-7B outperformed its base model, Qwen2.5-VL-7B, significantly. On MathVista, it reached 70.2% accuracy compared to the base model’s 50.2%. On MathVerse, the improvement was from 46.8% to 68.5%. MathVision full test accuracy rose from 24.0% to 29.6%, and MathVision testmini improved from 25.3% to 30.4%. These improvements indicate that the model learned to follow reasoning patterns and generalized better to unseen multimodal tasks. Each iteration of training contributed measurable gains, showcasing the strength of combining fine-tuning with reward-based learning in a looped structure.

The core of this model’s strength lies in its iterative structure. Rather than relying solely on vast datasets, it focuses on quality and structure. Each cycle of SFT and RL improves the model’s capacity to understand the relationship between images, questions, and answers. Self-verification and correction behaviors, initially lacking in standard LVLMs, emerged as a byproduct of reinforcement learning with verifiable reward signals. This allowed OpenVLThinker-7B to produce reasoning traces that were logically consistent and interpretable. Even subtle improvements, such as reduced redundant self-reflections or increased accuracy with shorter reasoning chains, contributed to its overall performance gains.

Some Key Takeaways from the Research: 

  • UCLA researchers developed OpenVLThinker-7B using a combined SFT and RL approach, starting from the Qwen2.5-VL-7B base model.
  • Used iterative training cycles involving caption generation, reasoning distillation, and alternating SFT and GRPO reinforcement learning.
  • The initial SFT used 25,000 filtered examples, while the RL phases used smaller sets of 5,000 harder samples from datasets like Geometry3K and SuperCLEVR.
  • On MathVista, accuracy improved from 50.2% (base model) to 70.2%. MathVerse accuracy jumped from 46.8% to 68.5%, and other datasets also saw notable gains.
  • GRPO effectively refined reasoning behaviors by rewarding correct answers, reducing verbosity, and improving logical consistency.
  • Each training iteration led to incremental performance gains, confirming the effectiveness of the self-improvement strategy.
  • Establishes a viable route to bring R1-style multi-step reasoning into multimodal models, useful for educational, visual analytics, and assistive tech applications.

Check out the Paper, Model on Hugging Face and GitHub Page. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

Credit: Source link

ShareTweetSendSharePin

Related Posts

The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
How To Hide Or Replace The Audio Button In iMessages
AI & Technology

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026
Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)
AI & Technology

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)

September 22, 2026
Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor
AI & Technology

Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor

September 22, 2026
Next Post
Rapper Young Scooter dies from injuries after fleeing Atlanta officers, authorities say –  The Atlanta Journal Constitution

Rapper Young Scooter dies from injuries after fleeing Atlanta officers, authorities say - The Atlanta Journal Constitution

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Israel launches deadly strike on southern Lebanon

Israel launches deadly strike on southern Lebanon

September 16, 2026
Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
College football winners, losers: Alabama QB Keelon Russell has arrived, Texas A&M faceplants – cbssports.com

College football winners, losers: Alabama QB Keelon Russell has arrived, Texas A&M faceplants – cbssports.com

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!