• bitcoinBitcoin(BTC)$75,799.00-2.11%
  • ethereumEthereum(ETH)$2,401.75-3.47%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$713.62-0.80%
  • rippleXRP(XRP)$1.30-7.67%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$97.06-3.76%
  • tronTRON(TRX)$0.334531-0.94%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.43%
  • zcashZcash(ZEC)$1,151.980.89%
  • HyperliquidHyperliquid(HYPE)$77.21-2.24%
  • dogecoinDogecoin(DOGE)$0.080121-3.25%
  • RainRain(RAIN)$0.013998-0.29%
  • USDSUSDS(USDS)$1.00-0.03%
  • moneroMonero(XMR)$504.82-1.28%
  • whitebitWhiteBIT Coin(WBT)$77.89-2.70%
  • leo-tokenLEO Token(LEO)$8.87-1.01%
  • chainlinkChainlink(LINK)$10.87-4.80%
  • cardanoCardano(ADA)$0.195436-4.17%
  • stellarStellar(XLM)$0.176330-7.98%
  • Ethena USDeEthena USDe(USDE)$1.00-0.07%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$219.84-1.04%
  • USD1USD1(USD1)$1.00-0.03%
  • litecoinLitecoin(LTC)$51.09-2.95%
  • uniswapUniswap(UNI)$6.34-4.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-1.90%
  • CantonCanton(CC)$0.091584-4.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.074473-2.87%
  • avalanche-2Avalanche(AVAX)$7.32-2.18%
  • nearNEAR Protocol(NEAR)$2.36-2.79%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.05%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • suiSui(SUI)$0.69-2.45%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,322.940.70%
  • crypto-com-chainCronos(CRO)$0.055315-3.20%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.112.22%
  • BittensorBittensor(TAO)$217.33-6.29%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$110.98-1.58%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.14%
  • BitwayBitway(BTW)$0.7711.45%
  • pax-goldPAX Gold(PAXG)$4,328.010.74%
  • aaveAave(AAVE)$121.11-4.73%
  • AsterAster(ASTER)$0.68-1.85%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0569360.08%
  • mantleMantle(MNT)$0.54-5.55%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing AI Model’s Scalability and Performance: A Study on Multi-Head Mixture-of-Experts

April 26, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Enhancing AI Model’s Scalability and Performance: A Study on Multi-Head Mixture-of-Experts
ShareShareShareShareShare

Large capacity models, such as Large Language Models (LLMs) and Large Multi-modal Models (LMMs), have demonstrated effectiveness across various domains and tasks. Scaling up these models by increasing parameter count enhances performance but significantly reduces inference speed, limiting practicality. Sparse Mixtures of Experts (SMoE) offer a promising alternative, enabling model scalability while mitigating computational costs. However, SMoE faces two key challenges: i) low expert activation and ii) limited analytical capabilities, which hinder its effectiveness and scalability.

SMoE enhances model capacity while maintaining constant computational demand, yielding superior performance compared to densely-activated models. Unlike dense models, SMoE employs N-independent Feed-Forward Networks (FFN) as experts within each Mixture-of-Experts (MoE) layer and a gating function to distribute weights over these experts’ outputs. The routing mechanism selects the top-k experts from N experts, where k << N facilitates data and expert parallelism. Larger k values often improve model performance but can reduce training efficiency.

Researchers from Tsinghua University and Microsoft Research introduce Multi-Head Mixture-of-Experts (MH-MoE). MH-MoE utilises a multi-head mechanism to divide each input token into multiple sub-tokens and distribute them across different experts, achieving denser expert activation without increasing computational or parameter complexity. In contrast to SMoE, MH-MoE activates four experts for a single input token by splitting it into four sub-tokens. This allocation enables the model to focus on various representation spaces within experts, facilitating a more nuanced understanding of vision and language patterns. 

The architecture of MH-MoE addresses issues of low expert activation and token ambiguity by employing a multi-head mechanism to split tokens into sub-tokens and route them to various experts. In MH-MoE, each parallel layer contains a set of N experts, with a multi-head layer projecting inputs followed by token splitting and gating functions to route sub-tokens to experts. The top-k routing mechanism activates experts with the highest scores, and the resulting sub-tokens are processed by these activated experts and rearranged before token merging to maintain input-output shape consistency. The Token-Splitting-Merging (TSM) operation increases the data volume routed to specific experts, resulting in denser expert activation and improved understanding. This process ensures no additional computational cost in subsequent blocks, with a hyperparameter β used to balance parameters and computational complexity with the original SMoE.

The validation perplexity curves for all pretrained models and pre-training tasks are examined under two expert settings (8 experts and 32 experts). MH-MoE consistently maintains lower perplexity than the baselines across various experimental setups, indicating more effective learning. Also, increasing the number of experts correlates with a decrease in perplexity for MH-MoE, suggesting enhanced representation learning capabilities. Downstream evaluation across different pre-training tasks further validates the efficacy of MH-MoE. In English-focused language modeling, MH-MoE achieves the best performance across multiple benchmarks, demonstrating its effectiveness in improving language representation. Similarly, MH-MoE outperforms X-MoE consistently in multi-lingual language modeling, showcasing its superiority in modeling cross-lingual natural language. In masked multi-modal modeling tasks such as visual question answering, visual reasoning, and image captioning, MH-MoE consistently outperforms Dense and X-MoE baselines, underscoring its ability to capture diverse semantic and detailed information within visual data.

In conclusion, This paper investigates methods for achieving denser expert activation without introducing additional cost while enhancing fine-grained understanding ability. The proposed MH-MoE offers a straightforward implementation of these functionalities. Also, MH-MoE’s simplicity facilitates seamless integration with other SMoE frameworks, improving performance easily. Extensive empirical results across three tasks validate the effectiveness of MH-MoE in achieving these objectives.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera

Considering A Level 2 EV Charger? How To Know If You Need One

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera
AI & Technology

Canon’s R8 II Camera Borrowed Its Styling From A Classic SLR Film Camera

September 16, 2026
Considering A Level 2 EV Charger? How To Know If You Need One
AI & Technology

Considering A Level 2 EV Charger? How To Know If You Need One

September 16, 2026
How To Get Spotify’s Best Audio Quality
AI & Technology

How To Get Spotify’s Best Audio Quality

September 15, 2026
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
AI & Technology

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend

September 15, 2026
Next Post
‘Innocent lives are on the line’: Biden pledges food aid into Gaza

'Innocent lives are on the line': Biden pledges food aid into Gaza

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Aryna Sabalenka v Elena Rybakina: US Open 2026 women’s singles final – live – The Guardian

Aryna Sabalenka v Elena Rybakina: US Open 2026 women’s singles final – live – The Guardian

September 12, 2026
Waterspout swirls off northern Italy’s coast

Waterspout swirls off northern Italy’s coast

September 13, 2026
Crypto billionaire paid .5M to marry glamorous actress — now he wants it back after romance fizzled

Crypto billionaire paid $4.5M to marry glamorous actress — now he wants it back after romance fizzled

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!