• bitcoinBitcoin(BTC)$84,104.00-0.11%
  • ethereumEthereum(ETH)$2,677.22-0.43%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$774.340.31%
  • rippleXRP(XRP)$1.531.41%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.421.13%
  • tronTRON(TRX)$0.338569-1.34%
  • zcashZcash(ZEC)$1,541.851.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.75%
  • HyperliquidHyperliquid(HYPE)$91.80-1.27%
  • dogecoinDogecoin(DOGE)$0.0950200.47%
  • moneroMonero(XMR)$568.071.37%
  • chainlinkChainlink(LINK)$13.397.77%
  • whitebitWhiteBIT Coin(WBT)$83.85-0.84%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2478072.66%
  • RainRain(RAIN)$0.011963-1.85%
  • leo-tokenLEO Token(LEO)$8.81-1.55%
  • stellarStellar(XLM)$0.2187127.55%
  • bitcoin-cashBitcoin Cash(BCH)$333.56-1.82%
  • nearNEAR Protocol(NEAR)$4.515.30%
  • uniswapUniswap(UNI)$9.12-2.09%
  • litecoinLitecoin(LTC)$71.374.37%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.1170437.30%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.17-1.10%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.025.43%
  • hedera-hashgraphHedera(HBAR)$0.0922320.29%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-1.49%
  • shiba-inuShiba Inu(SHIB)$0.0000060.73%
  • BittensorBittensor(TAO)$298.982.75%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0645283.47%
  • MemeCoreMemeCore(M)$1.22-3.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,273.23-0.18%
  • OndoOndo(ONDO)$0.5425.20%
  • okbOKB(OKB)$119.48-0.51%
  • BitwayBitway(BTW)$0.90-16.16%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • aaveAave(AAVE)$144.983.74%
  • mantleMantle(MNT)$0.681.87%
  • EthenaEthena(ENA)$0.2209864.65%
  • polkadotPolkadot(DOT)$1.162.71%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Branch-and-Merge Method: Enhancing Language Adaptation in AI Models by Mitigating Catastrophic Forgetting and Ensuring Retention of Base Language Capabilities while Learning New Languages

July 14, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Branch-and-Merge Method: Enhancing Language Adaptation in AI Models by Mitigating Catastrophic Forgetting and Ensuring Retention of Base Language Capabilities while Learning New Languages
ShareShareShareShareShare

Language model adaptation is a crucial area in artificial intelligence, focusing on enhancing large pre-trained language models to work effectively across various languages. This research is vital for enabling these models to understand and generate text in multiple languages, which is essential for global AI applications. Despite the impressive performance of LLMs in English, their capabilities significantly drop when adapted to less prevalent languages, making additional adaptation techniques necessary.

One of the significant challenges in adapting language models to new languages is catastrophic forgetting. This occurs when a model loses its proficiency in the original language while learning a new one, severely limiting its usefulness. Retaining the base model’s capabilities is essential for solving tasks in the new language, as skills such as math and coding learned in English are invaluable for problem-solving and reasoning in other languages.

YOU MAY ALSO LIKE

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

Warzone Is Adding A Button To Hide All The Goofy Skins

Current methods to address catastrophic forgetting include continued pretraining and instruction tuning with experience replay. Experience replay involves mixing data from the original language during training in the new language. However, this approach needs to be revised to fully mitigate forgetting, especially when the exact source data is unknown. The approximation of experience replay reduces its effectiveness, necessitating further regularization to maintain the model’s performance in the base language.

Researchers from INSAIT, LogicStar.ai, ETH Zurich, the University of Chicago, and Together AI  introduced a novel approach called Branch-and-Merge (BAM). This method iteratively merges multiple models, each fine-tuned on different subsets of training data, to achieve lower magnitude but higher quality weight changes. By combining these models, BAM reduces forgetting while maintaining learning efficiency. The BAM method splits the training data into several slices and fine-tunes the base model on these slices in parallel. The resulting models are merged to form a new base model for the next iteration. This iterative process minimizes the total weight change, reducing the risk of catastrophic forgetting. Additionally, by leveraging multiple training slices, BAM ensures the retention of essential skills from the base language.

In detail, BAM splits the training data into N slices and fine-tunes the base model on K (typically two) of these slices in parallel before merging the resulting models. This significantly reduces the total weight change, preserving most of the learning from the parallel training steps. The research team applied BAM to adapt models like MISTRAL-7B and LLAMA-3-8B from predominantly English to Bulgarian and German. They found that BAM consistently improved benchmark performance in target and source languages compared to standard training methods. For instance, the BAM-trained LLAMA-3-8B improved Bulgarian task performance by 10.9% and English task performance by 1.3%, demonstrating the method’s efficacy.

To further understand the performance of BAM, the researchers conducted an extensive empirical study. They applied BAM to adapt MISTRAL-7B and LLAMA-3-8B models, predominantly using English data, to Bulgarian and German languages. The results showed that BAM significantly reduced forgetting while matching or improving target domain performance compared to standard continued pretraining and fine-tuning instruction. Specifically, BAM allowed the LLAMA-3-8B model to outperform its standard counterpart by 10.9% in Bulgarian tasks and 1.3% in English tasks. This improvement is attributed to the smaller magnitude but more efficient weight changes induced by BAM.

BAM was evaluated using both approximate and minimal experience replay. The approximate experience replay involved a mix of 15.1 billion unique tokens from sources like OpenWebText, English Wikipedia, and GitHub repositories. In contrast, minimal experience replay used only 5 billion tokens from OpenWebText for German and 10 billion tokens for Bulgarian. The study found that approximate experience replay led to a stronger increase in target domain performance and reduced forgetting of the source domain compared to minimal experience replay.

The effectiveness of BAM was also demonstrated in instruction fine-tuning. Using 928,000 samples of English finetuning data mixed with German or Bulgarian data, BAM slightly improved learning in both target languages while significantly reducing forgetting. For instance, BAM-trained models outperformed the standard instruction fine-tuning models in the Bulgarian instruction tuning, achieving 10.8% better performance in Bulgarian tasks and 1.3% better in English tasks.

In conclusion, the Branch-and-Merge (BAM) method offers a robust solution for catastrophic forgetting in language model adaptation. Ensuring minimal yet effective weight changes preserves the model’s capabilities in the original language while enhancing its performance in the target language. This approach can significantly benefit practitioners working on multilingual AI applications, providing a more efficient way to adapt large language models to diverse linguistic environments. The research demonstrated that BAM could effectively balance learning and forgetting, making it a valuable method for continuous pretraining and instruction tuning in alphabet- and non-alphabet-sharing languages.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
AI & Technology

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

September 25, 2026
Warzone Is Adding A Button To Hide All The Goofy Skins
AI & Technology

Warzone Is Adding A Button To Hide All The Goofy Skins

September 24, 2026
How These AI Glasses Compare
AI & Technology

How These AI Glasses Compare

September 24, 2026
Nintendo Wins .5 Million From Lawsuit Over Pirated Switch Games
AI & Technology

Nintendo Wins $4.5 Million From Lawsuit Over Pirated Switch Games

September 24, 2026
Next Post
Amazon Prime Day 2024 is almost here and these are the best tech deals we could find from Apple, Bose, Samsung and more

Amazon Prime Day 2024 is almost here and these are the best tech deals we could find from Apple, Bose, Samsung and more

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Residents lose homes as firefighters battle Hawk Fire

Residents lose homes as firefighters battle Hawk Fire

September 24, 2026
How To Block And Unblock A Number On Your Android Phone

How To Block And Unblock A Number On Your Android Phone

September 19, 2026
'Fire everyone': How NC State lost to Vanderbilt on disastrous game-ending fumble – USA Today

'Fire everyone': How NC State lost to Vanderbilt on disastrous game-ending fumble – USA Today

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!