• bitcoinBitcoin(BTC)$77,176.000.27%
  • ethereumEthereum(ETH)$2,523.88-0.48%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$730.591.18%
  • rippleXRP(XRP)$1.371.03%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.881.23%
  • tronTRON(TRX)$0.3400371.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.27%
  • zcashZcash(ZEC)$1,136.46-2.20%
  • HyperliquidHyperliquid(HYPE)$80.490.08%
  • dogecoinDogecoin(DOGE)$0.0849221.25%
  • RainRain(RAIN)$0.015300-2.51%
  • moneroMonero(XMR)$533.534.17%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.240.11%
  • chainlinkChainlink(LINK)$11.52-0.77%
  • leo-tokenLEO Token(LEO)$9.11-0.48%
  • cardanoCardano(ADA)$0.2079251.77%
  • stellarStellar(XLM)$0.1809811.76%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$227.42-0.19%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.851.00%
  • uniswapUniswap(UNI)$6.314.65%
  • CantonCanton(CC)$0.0974900.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.382.37%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0746220.79%
  • avalanche-2Avalanche(AVAX)$7.39-1.09%
  • shiba-inuShiba Inu(SHIB)$0.0000052.51%
  • nearNEAR Protocol(NEAR)$2.37-6.90%
  • suiSui(SUI)$0.720.25%
  • crypto-com-chainCronos(CRO)$0.0585133.65%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.181.15%
  • tether-goldTether Gold(XAUT)$4,349.75-0.17%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.290.19%
  • BittensorBittensor(TAO)$233.66-0.12%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • aaveAave(AAVE)$126.261.81%
  • mantleMantle(MNT)$0.57-3.69%
  • pax-goldPAX Gold(PAXG)$4,355.59-0.13%
  • AsterAster(ASTER)$0.690.77%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0567367.21%
  • polkadotPolkadot(DOT)$1.03-1.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A New AI Research Introduces LoRAMoE: A Plugin Version of Mixture of Experts (Moe) for Maintaining World Knowledge in Language Model Alignment

January 3, 2024
in AI & Technology
Reading Time: 4 mins read
A A
A New AI Research Introduces LoRAMoE: A Plugin Version of Mixture of Experts (Moe) for Maintaining World Knowledge in Language Model Alignment
ShareShareShareShareShare

Large Language Models (LLMs) have proven remarkably effective in numerous jobs. To fully realize the potential of the models, supervised fine-tuning (SFT) is necessary to match them with human instructions. A simple option when the variety of tasks increases or when improved performance on a particular activity is needed is to increase the amount of data, even if some works have shown that models can follow human instruction successfully with a little fine-tuning of data. 

Several studies show that the significant growth of fine-tuning data presents new difficulties. In particular, researchers have found that performance significantly decreases with significant increases in fine-tuning data on the Natural Questions dataset from the Closed-Book Question Answering (CBQA) dataset. The collapse of previously learned and stored world knowledge in the pre-trained models could be linked to this notable performance loss. There are two phases involved in proving this proposition. Firstly, the CBQA dataset draws conclusions from the world information in the models. Second, large-scale fine-tuning can significantly alter the model’s parameters, erasing world information (i.e., knowledge forgetting), which is responsible for the notable decline in performance on the CBQA dataset. There is a conflict in vanilla-supervised fine-tuning between preserving LLM world information and enhancing performance on downstream tasks concurrently.

The best course of action is to designate a certain area of the model for storing global information, much like the human brain’s hippocampus, which is specialized for remembering. Nonetheless, direct and fine-tuning with a single plugin are comparable in form. An architecture known as a “Mixture of Experts” (MoE) includes several experts, and data with varying properties is sent to the appropriate experts for personalized processing. Using this concept, a group of researchers from Fudan University and Hikvision Inc. intend to offer numerous plugins as experts, enabling one portion to access the backup and another to carry out downstream operations.

Their new study presents LoRAMoE, which can improve the LLMs’ downstream task-solving capacities and mitigate the forgetting of world knowledge. A plugin version of MoE is called LoRAMoE. Introducing numerous parallel plugins that are specialists in every feed-forward layer and coupling them to routers modifies the model’s architecture. Next, they suggest creating separate groups of experts for each LoRAMoE layer using localized balancing constraints. To be more precise, one group works on downstream tasks, and the other is tasked with reducing knowledge forgetting by aligning human instructions with the world information included in the backbone model. Furthermore, the localized balancing constraint prohibits the routers from placing too much weight on only a few experts within the same expert group by balancing the relevance of all experts within the same expert group. It allows multiple professionals to work together, enhancing the capacity to complete jobs later. 

The experiment results demonstrate that LoRAMoE can successfully prevent large-scale fine-tuning from upsetting the world information included in language models. Furthermore, by visualizing the expert weight for tasks, the team validated LoRAMoE’s efficacy on capacity localization at an interpretable level. The findings indicate that the router prioritizes the output of experts who specialize in completing world knowledge benchmarks. On the other hand, the router concentrates on specialists from a different group for other downstream duties. LoRAMoE successfully settles the dispute by encouraging expert cooperation. Furthermore, the experiment results indicate that the proposed strategy improves learning on various downstream tasks, suggesting the method’s potential for multi-task learning.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, LinkedIn Group, Twitter, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

Anthropic’s CEO Proposes A Three-Step Plan To Curb AI Development

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🎯 Meet Meetgeek: your personal AI Meeting Assistant…. Try it now!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI
AI & Technology

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

September 12, 2026
Anthropic’s CEO Proposes A Three-Step Plan To Curb AI Development
AI & Technology

Anthropic’s CEO Proposes A Three-Step Plan To Curb AI Development

September 12, 2026
Amodei Calls for Slowing the Pace of AI Capability Improvement – Unite.AI
AI & Technology

Amodei Calls for Slowing the Pace of AI Capability Improvement – Unite.AI

September 12, 2026
Are You Using The Right Ethernet Port On Your Router? Here’s How To Know
AI & Technology

Are You Using The Right Ethernet Port On Your Router? Here’s How To Know

September 12, 2026
Next Post
Enhancing Accountability and Trust: Meet the ‘AI Foundation Model Transparency Act’

Enhancing Accountability and Trust: Meet the 'AI Foundation Model Transparency Act'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Wholesale inflation rises by strongest increase in 3 months

Wholesale inflation rises by strongest increase in 3 months

September 10, 2026
L.A. ‘graffiti towers’ to be cleaned in 90 days

L.A. ‘graffiti towers’ to be cleaned in 90 days

September 6, 2026
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!