• bitcoinBitcoin(BTC)$84,013.00-0.01%
  • ethereumEthereum(ETH)$2,685.01-0.36%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$770.75-0.57%
  • rippleXRP(XRP)$1.52-2.93%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$121.21-0.71%
  • tronTRON(TRX)$0.336006-0.49%
  • zcashZcash(ZEC)$1,557.17-0.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.19%
  • HyperliquidHyperliquid(HYPE)$91.68-0.12%
  • dogecoinDogecoin(DOGE)$0.097247-0.52%
  • chainlinkChainlink(LINK)$14.172.17%
  • moneroMonero(XMR)$552.81-0.94%
  • whitebitWhiteBIT Coin(WBT)$83.83-0.12%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2551030.06%
  • RainRain(RAIN)$0.0130839.61%
  • leo-tokenLEO Token(LEO)$8.961.52%
  • stellarStellar(XLM)$0.218252-0.53%
  • bitcoin-cashBitcoin Cash(BCH)$337.68-1.24%
  • nearNEAR Protocol(NEAR)$4.82-5.60%
  • uniswapUniswap(UNI)$9.56-0.47%
  • litecoinLitecoin(LTC)$72.011.55%
  • CantonCanton(CC)$0.1340472.48%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.803.12%
  • suiSui(SUI)$1.162.72%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.558.56%
  • hedera-hashgraphHedera(HBAR)$0.093565-0.60%
  • BittensorBittensor(TAO)$322.904.89%
  • shiba-inuShiba Inu(SHIB)$0.0000061.52%
  • crypto-com-chainCronos(CRO)$0.0658710.07%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • BitwayBitway(BTW)$1.06-16.79%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • EthenaEthena(ENA)$0.2727786.15%
  • MemeCoreMemeCore(M)$1.211.07%
  • tether-goldTether Gold(XAUT)$4,278.98-0.26%
  • OndoOndo(ONDO)$0.540.46%
  • okbOKB(OKB)$120.870.13%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.720.04%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.04%
  • mantleMantle(MNT)$0.693.28%
  • polkadotPolkadot(DOT)$1.245.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Prefix-RFT: A Unified Machine Learning Framework to blend Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT)

August 24, 2025
in AI & Technology
Reading Time: 7 mins read
A A
Prefix-RFT: A Unified Machine Learning Framework to blend Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT)
ShareShareShareShareShare

Large language models are typically refined after pretraining using either supervised fine-tuning (SFT) or reinforcement fine-tuning (RFT), each with distinct strengths and limitations. SFT is effective in teaching instruction-following through example-based learning, but it can lead to rigid behavior and poor generalization. RFT, on the other hand, optimizes models for task success using reward signals, which can improve performance but also introduce instability and reliance on a strong starting policy. While these methods are often used sequentially, their interaction remains poorly understood. This raises an important question: how can we design a unified framework that combines SFT’s structure with RFT’s goal-driven learning? 

Research at the intersection of RL and LLM post-training has gained momentum, particularly for training reasoning-capable models. Offline RL, which learns from fixed datasets, often yields suboptimal policies due to the limited diversity of the data. This has sparked interest in combining offline and online RL approaches to improve performance. In LLMs, the dominant strategy is to first apply SFT to teach desirable behaviors, then use RFT to optimize outcomes. However, the dynamics between SFT and RFT are still not well understood, and finding effective ways to integrate them remains an open research challenge. 

YOU MAY ALSO LIKE

You Can Use Your Old Laptop To Make A Smart Home Hub

TikTok Will Pay Alabama $100 Million To Settle Social Media Addiction Lawsuit

Researchers from the University of Edinburgh, Fudan University, Alibaba Group, Stepfun, and the University of Amsterdam propose a unified framework that combines supervised and reinforcement fine-tuning in a way called Prefix-RFT. This method guides exploration using partial demonstrations, allowing the model to continue generating solutions with flexibility and adaptability. Tested on math reasoning tasks, Prefix-RFT consistently outperforms standalone SFT, RFT, and mixed-policy methods. It integrates easily into existing frameworks and proves robust to changes in demonstration quality and quantity. Blending demonstration-based learning with exploration can lead to more effective and adaptive training of large language models. 

https://arxiv.org/abs/2507.01679

The study presents Prefix Reinforcement Fine-Tuning (Prefix-RFT) as a way to blend the strengths of SFT and RFT. While SFT offers stability by mimicking expert demonstrations, RFT encourages exploration through the use of reward signals. Prefix-RFT bridges the two by using a partial demonstration (a prefix) and letting the model generate the rest. This approach guides learning without relying too heavily on full supervision. It incorporates techniques like entropy-based clipping and a cosine decay scheduler to ensure stable training and efficient learning. Compared to prior methods, Prefix-RFT offers a more balanced and adaptive fine-tuning strategy. 

Prefix-RFT is a reward fine-tuning method that improves performance using high-quality offline math datasets, such as OpenR1-Math-220K (46k filtered problems). Tested on Qwen2.5-Math-7B, 1.5B, and LLaMA-3.1-8B, it was evaluated on benchmarks including AIME 2024/25, AMC, MATH500, Minerva, and OlympiadBench. Prefix-RFT achieved the highest avg@32 and pass@1 scores across tasks, outperforming RFT, SFT, ReLIFT, and LUFFY. Using Dr. GRPO, it updated only the top 20% high-entropy prefix tokens, with the prefix length decaying from 95% to 5%. It maintained intermediate SFT loss, indicating a strong balance between imitation and exploration, especially on difficult problems (Trainhard). 

https://arxiv.org/abs/2507.01679

In conclusion, Prefix-RFT combines the strengths of SFT and RFT by utilizing sampled demonstration prefixes to guide learning. Despite its simplicity, it consistently outperforms SFT, RFT, and hybrid baselines across various models and datasets. Even with only 1% of the training data (450 prompts), it maintains strong performance (avg@32 drops only from 40.8 to 37.6), showing efficiency and robustness. Its top-20% entropy-based token update strategy proves most effective, achieving the highest benchmark scores with shorter outputs. Moreover, using a cosine decay scheduler for prefix length enhances stability and learning dynamics compared to a uniform strategy, particularly on complex tasks such as AIME. 


Check out the Paper here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

Credit: Source link

ShareTweetSendSharePin

Related Posts

You Can Use Your Old Laptop To Make A Smart Home Hub
AI & Technology

You Can Use Your Old Laptop To Make A Smart Home Hub

September 26, 2026
TikTok Will Pay Alabama 0 Million To Settle Social Media Addiction Lawsuit
AI & Technology

TikTok Will Pay Alabama $100 Million To Settle Social Media Addiction Lawsuit

September 26, 2026
This App Lets You Use An Apple Watch With An Android Phone
AI & Technology

This App Lets You Use An Apple Watch With An Android Phone

September 26, 2026
These Xbox Players Got GTA 6 For Free The Hard Way
AI & Technology

These Xbox Players Got GTA 6 For Free The Hard Way

September 26, 2026
Next Post
NBC Nightly News Full Episode – Feb. 24

NBC Nightly News Full Episode - Feb. 24

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Freddie Mac: Multifamily Delinquency Rate Rises To Multi-Decade High

Freddie Mac: Multifamily Delinquency Rate Rises To Multi-Decade High

September 26, 2026
AI ‘Existential Risk’ Is Close to Zero: Databricks CEO

AI ‘Existential Risk’ Is Close to Zero: Databricks CEO

September 20, 2026
ICE officer in Texas shooting was recruit not using a body camera – The Washington Post

ICE officer in Texas shooting was recruit not using a body camera – The Washington Post

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!