• bitcoinBitcoin(BTC)$84,803.000.94%
  • ethereumEthereum(ETH)$2,695.260.43%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$778.911.08%
  • rippleXRP(XRP)$1.530.65%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$123.291.71%
  • tronTRON(TRX)$0.334530-0.45%
  • zcashZcash(ZEC)$1,608.023.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.84%
  • HyperliquidHyperliquid(HYPE)$91.920.27%
  • dogecoinDogecoin(DOGE)$0.0974840.24%
  • chainlinkChainlink(LINK)$14.11-0.38%
  • moneroMonero(XMR)$547.47-0.98%
  • whitebitWhiteBIT Coin(WBT)$84.530.85%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2567830.71%
  • RainRain(RAIN)$0.012580-3.81%
  • leo-tokenLEO Token(LEO)$9.010.52%
  • stellarStellar(XLM)$0.217059-0.51%
  • nearNEAR Protocol(NEAR)$5.4613.33%
  • bitcoin-cashBitcoin Cash(BCH)$335.90-0.56%
  • uniswapUniswap(UNI)$9.782.35%
  • litecoinLitecoin(LTC)$71.23-1.08%
  • CantonCanton(CC)$0.1379542.82%
  • suiSui(SUI)$1.279.46%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • avalanche-2Avalanche(AVAX)$11.022.18%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.645.26%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0951651.74%
  • BittensorBittensor(TAO)$328.701.81%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.13%
  • crypto-com-chainCronos(CRO)$0.0673182.30%
  • BitwayBitway(BTW)$1.2013.37%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • EthenaEthena(ENA)$0.2911926.68%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.20-0.26%
  • quant-networkQuant(QNT)$188.1853.91%
  • OndoOndo(ONDO)$0.563.05%
  • tether-goldTether Gold(XAUT)$4,280.090.02%
  • okbOKB(OKB)$121.560.57%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.490.44%
  • Pump.funPump.fun(PUMP)$0.00496412.49%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.04%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research from Cohere for AI Compares Merging vs Data Mixing as a Recipe for Building High-Performant Aligned LLMs

October 21, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Research from Cohere for AI Compares Merging vs Data Mixing as a Recipe for Building High-Performant Aligned LLMs
ShareShareShareShareShare

Large language models (LLMs) have revolutionized the field of artificial intelligence by performing a wide range of tasks across different domains. These models are expected to work seamlessly in multiple languages, solving complex problems while ensuring safety. However, the challenge lies in maintaining safety without compromising performance, especially in multilingual settings. As AI technologies become globally pervasive, addressing the safety concerns that arise when models trained predominantly in English are deployed across various languages and cultural contexts is essential.

The core issue revolves around balancing performance and safety in LLMs. Safety concerns arise when models produce biased or harmful outputs, particularly in languages with limited training data. Typically, the methods to address this involve fine-tuning models on mixed datasets, combining general-purpose and safety tasks. However, these approaches can lead to undesirable trade-offs. In many cases, increasing the safety measures in LLMs can negatively impact their ability to perform well on general tasks. The challenge, therefore, is to develop an approach that improves both safety and performance in multilingual LLMs without requiring massive amounts of task-specific data.

YOU MAY ALSO LIKE

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

How To Improve Your Router’s Security In 10 Minutes

Current methods used to balance these objectives often rely on data-mixing techniques. This involves creating a single model by training it on various datasets from various tasks and languages. While these methods help achieve some level of multitasking ability, they can result in under-addressed safety concerns in languages other than English. In addition, the complexity of managing numerous tasks simultaneously often reduces the model’s ability to perform well in any of them. The lack of specialized attention to each task and language limits the model’s capacity to address safety and general performance effectively.

To overcome these limitations, researchers from Cohere AI have introduced an innovative approach based on model merging. Instead of relying on the traditional method of data mixing, where a single model is trained across multiple tasks and languages, the researchers propose merging separate models that have been independently fine-tuned for specific tasks and languages. This method allows for better specialization within each model before merging them into a unified system. By doing so, the models retain their unique capabilities, offering improvements in safety and general performance across diverse languages.

The merging process is conducted through multiple techniques. The primary method introduced by the researchers is Spherical Linear Interpolation (SLERP), which allows for smooth transitions between different models by blending their weights along a spherical path. This technique ensures that the unique properties of each model are preserved, allowing the merged model to handle various tasks without compromising safety or performance. Another method, TIES (Task Interference Elimination Strategy), focuses on resolving conflicts between task-specific fine-tuned models by adjusting model parameters to align better. The merging techniques also include linear merging and DARE-TIES, which further enhance the robustness of the final model by addressing interference issues and ensuring that the model parameters contribute positively to performance.

The results of this research show clear improvements in both general performance and safety. For instance, SLERP merging achieved an impressive 7% improvement in general performance and a 3.1% reduction in harmful outputs compared to traditional data mixing methods. On the other hand, TIES merging delivered a remarkable 10.4% reduction in harmful outputs, although it slightly reduced general performance by 7.4%. These numbers indicate that model merging significantly outperforms data mixing when balancing safety and performance. Moreover, when models were fine-tuned for individual languages and merged, the researchers observed up to a 6.6% reduction in harmful outputs and a 3.8% improvement in general benchmarks, further proving the effectiveness of language-specific model merging over multilingual model training.

The performance improvements were particularly noteworthy in some languages, with Russian showing the highest reduction in harmful generations (up to 15%) using TIES merging. Spanish, meanwhile, exhibited a 10% improvement in general performance with both SLERP and TIES methods. However, not all languages benefit equally. English models, for example, showed a decline in safety performance when merged, highlighting the variability in outcomes based on the underlying training data and merging strategy.

The research provides a comprehensive framework for building safer and more effective multilingual LLMs. By merging models fine-tuned for safety and performance on specific tasks and languages, the researchers from Cohere AI demonstrated a more efficient and scalable method for improving LLMs. The approach reduces the need for massive amounts of training data and allows for better alignment of safety protocols across languages, which is critically needed in today’s AI landscape.

In conclusion, model merging represents a promising step forward in addressing balancing performance and safety challenges in LLMs, particularly in multilingual settings. This method significantly improves LLMs’ ability to deliver safe and high-quality outputs, especially when applied to low-resource languages. As AI evolves, techniques like model merging could become essential tools for ensuring that AI systems are robust and safe across diverse linguistic and cultural contexts.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 50k+ ML SubReddit.

[Upcoming Live Webinar- Oct 29, 2024] The Best Platform for Serving Fine-Tuned Models: Predibase Inference Engine (Promoted)


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Listen to our latest AI podcasts and AI research videos here ➡️


Credit: Source link

ShareTweetSendSharePin

Related Posts

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8
AI & Technology

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

September 27, 2026
How To Improve Your Router’s Security In 10 Minutes
AI & Technology

How To Improve Your Router’s Security In 10 Minutes

September 27, 2026
Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)
AI & Technology

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

September 27, 2026
AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
Next Post
Lindsey Graham to Republicans supporting Kamala Harris: ‘What the hell are you doing?’

Lindsey Graham to Republicans supporting Kamala Harris: ‘What the hell are you doing?’

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
White House withholding 9 million of small business loans

White House withholding $289 million of small business loans

September 23, 2026
This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

September 26, 2026
Oil-Rates Correlation Jumps To A 35-Year High

Oil-Rates Correlation Jumps To A 35-Year High

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!