• bitcoinBitcoin(BTC)$65,223.00-0.60%
  • ethereumEthereum(ETH)$1,874.37-2.40%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$566.60-0.60%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.11-2.40%
  • solanaSolana(SOL)$75.64-2.50%
  • tronTRON(TRX)$0.3293120.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.10%
  • whitebitWhiteBIT Coin(WBT)$56.71-1.20%
  • HyperliquidHyperliquid(HYPE)$58.49-1.20%
  • dogecoinDogecoin(DOGE)$0.069021-4.60%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013907-3.00%
  • leo-tokenLEO Token(LEO)$9.59-1.70%
  • zcashZcash(ZEC)$508.09-0.90%
  • moneroMonero(XMR)$354.890.70%
  • chainlinkChainlink(LINK)$8.43-1.90%
  • stellarStellar(XLM)$0.183834-0.20%
  • cardanoCardano(ADA)$0.167226-3.90%
  • CantonCanton(CC)$0.1221780.30%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$211.31-3.30%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.46-3.90%
  • litecoinLitecoin(LTC)$46.64-1.20%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.070995-2.30%
  • suiSui(SUI)$0.74-2.70%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • crypto-com-chainCronos(CRO)$0.057243-0.40%
  • avalanche-2Avalanche(AVAX)$6.26-4.10%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,025.87-2.20%
  • nearNEAR Protocol(NEAR)$1.891.50%
  • shiba-inuShiba Inu(SHIB)$0.000004-2.20%
  • uniswapUniswap(UNI)$3.78-0.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • OndoOndo(ONDO)$0.403474-1.30%
  • BittensorBittensor(TAO)$192.76-1.20%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0565672.80%
  • pax-goldPAX Gold(PAXG)$4,023.08-2.30%
  • okbOKB(OKB)$81.88-0.20%
  • AsterAster(ASTER)$0.620.30%
  • HTX DAOHTX DAO(HTX)$0.000002-0.60%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • usddUSDD(USDD)$1.000.00%
  • MemeCoreMemeCore(M)$1.161.50%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers find that retraining only small parts of AI models can cut costs and prevent forgetting

October 13, 2025
in AI & Technology
Reading Time: 3 mins read
A A
Researchers find that retraining only small parts of AI models can cut costs and prevent forgetting
ShareShareShareShareShare

Enterprises often find that when they fine-tune models, one effective approach to making a large language model (LLM) fit for purpose and grounded in data is to have the model lose some of its abilities. After fine-tuning, some models “forget” how to perform certain tasks or other tasks they already learned. 

YOU MAY ALSO LIKE

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat

Research from the University of Illinois Urbana-Champaign proposes a new method for retraining models that avoids “catastrophic forgetting,” in which the model loses some of its prior knowledge. The paper focuses on two specific LLMs that generate responses from images: LLaVA and Qwen 2.5-VL.

The approach encourages enterprises to retrain only narrow parts of an LLM to avoid retraining the entire model and incurring a significant increase in compute costs. The team claims that catastrophic forgetting isn’t true memory loss, but rather a side effect of bias drift. 

“Training a new LMM can cost millions of dollars, weeks of time, and emit hundreds of tons of CO2, so finding ways to more efficiently and effectively update existing models is a pressing concern,” the team wrote in the paper. “Guided by this result, we explore tuning recipes that preserve learning while limiting output shift.”

The researchers focused on a multi-layer perceptron (MLP), the model's internal decision-making component. 

Catastrophic forgetting 

The researchers wanted first to verify the existence and the cause of catastrophic forgetting in models. 

To do this, they created a set of target tasks for the models to complete. The models were then fine-tuned and evaluated to determine whether they led to substantial forgetting. But as the process went on, the researchers found that the models were recovering some of their abilities. 

“We also noticed a surprising result, that the model performance would drop significantly in held out benchmarks after training on the counting task, it would mostly recover on PathVQA, another specialized task that is not well represented in the benchmarks,” they said. “Meanwhile, while performing the forgetting mitigation experiments, we also tried separately tuning only the self-attention projection (SA Proj) or MLP layers, motivated by the finding that tuning only the LLM was generally better than tuning the full model. This led to another very surprising result – that tuning only self-attention projection layers led to very good learning of the target tasks with no drop in performance in held out tasks, even after training all five target tasks in a sequence.”

The researchers said they believe that “what looks like forgetting or interference after fine-tuning on a narrow target task is actually bias in the output distribution due to the task distribution shift.”

Narrow retraining

That finding turned out to be the key to the experiment. The researchers noted that tuning the MLP increases the likelihood of “outputting numeric tokens and a highly correlated drop in held out task accuracy.” What it showed is that a model forgetting some of its knowledge is only temporary and not a long-term matter. 

“To avoid biasing the output distribution, we tune the MLP up/gating projections while keeping the down projection frozen, and find that it achieves similar learning to full MLP tuning with little forgetting,” the researchers said. 

This allows for a more straightforward and more reproducible method for fine-tuning a model. 

By focusing on a narrow segment of the model, rather than a wholesale retraining, enterprises can cut compute costs. It also allows better control of output drift. 

However, the research focuses only on two models, specifically those dealing with vision and language. The researchers noted that due to limited resources, they are unable to try the experiment with other models.

Their findings, however, can be extended to other LLMs, especially for different modalities. 

Credit: Source link

ShareTweetSendSharePin

Related Posts

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI
AI & Technology

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

July 23, 2026
Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat
AI & Technology

Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat

July 23, 2026
Agentic coding goes hands-free as OpenAI brings GPT-Live’s full duplex voice control to Codex and ChatGPT on the desktop
AI & Technology

Agentic coding goes hands-free as OpenAI brings GPT-Live’s full duplex voice control to Codex and ChatGPT on the desktop

July 23, 2026
Meta’s Pro-AI Ad Campaign Is Conspicuously Light On AI
AI & Technology

Meta’s Pro-AI Ad Campaign Is Conspicuously Light On AI

July 23, 2026
Next Post
Microsoft debuts its first in-house AI image generator

Microsoft debuts its first in-house AI image generator

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Canada issues air quality warnings over US wildfire smoke after Trump tariff threat – The Guardian

Canada issues air quality warnings over US wildfire smoke after Trump tariff threat – The Guardian

July 21, 2026
VTEB: A Low-Cost Tax-Exempt Fund Retaining Appeal With Intermediate Duration

VTEB: A Low-Cost Tax-Exempt Fund Retaining Appeal With Intermediate Duration

July 21, 2026
Taylor Farms and Taco Bell remove iceberg lettuce amid parasite outbreak

Taylor Farms and Taco Bell remove iceberg lettuce amid parasite outbreak

July 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!