• bitcoinBitcoin(BTC)$77,592.001.70%
  • ethereumEthereum(ETH)$2,484.942.06%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$752.824.05%
  • rippleXRP(XRP)$1.322.14%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$105.666.27%
  • tronTRON(TRX)$0.3358170.12%
  • zcashZcash(ZEC)$1,499.1311.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.16%
  • HyperliquidHyperliquid(HYPE)$87.6710.99%
  • dogecoinDogecoin(DOGE)$0.0842984.37%
  • moneroMonero(XMR)$529.416.48%
  • USDSUSDS(USDS)$1.000.05%
  • whitebitWhiteBIT Coin(WBT)$79.941.97%
  • RainRain(RAIN)$0.012796-0.61%
  • chainlinkChainlink(LINK)$11.805.97%
  • leo-tokenLEO Token(LEO)$8.91-0.36%
  • cardanoCardano(ADA)$0.2134338.66%
  • stellarStellar(XLM)$0.1870292.76%
  • uniswapUniswap(UNI)$8.5027.23%
  • bitcoin-cashBitcoin Cash(BCH)$247.0311.80%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.00%
  • nearNEAR Protocol(NEAR)$3.4829.98%
  • USD1USD1(USD1)$1.000.02%
  • CantonCanton(CC)$0.1084159.94%
  • litecoinLitecoin(LTC)$54.885.58%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.363.64%
  • avalanche-2Avalanche(AVAX)$7.895.06%
  • hedera-hashgraphHedera(HBAR)$0.0761463.43%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.788.07%
  • shiba-inuShiba Inu(SHIB)$0.0000056.48%
  • MemeCoreMemeCore(M)$1.2814.36%
  • crypto-com-chainCronos(CRO)$0.0584500.23%
  • paypal-usdPayPal USD(PYUSD)$1.000.03%
  • tether-goldTether Gold(XAUT)$4,385.461.87%
  • BittensorBittensor(TAO)$240.837.27%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$114.162.34%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.05%
  • aaveAave(AAVE)$134.5910.42%
  • AsterAster(ASTER)$0.753.10%
  • Pump.funPump.fun(PUMP)$0.00428310.18%
  • mantleMantle(MNT)$0.584.80%
  • polkadotPolkadot(DOT)$1.1312.15%
  • pax-goldPAX Gold(PAXG)$4,384.701.81%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Skywork Team Introduces Skywork-MoE: A High-Performance Mixture-of-Experts (MoE) Model with 146B Parameters, 16 Experts, and 22B Activated Parameters

June 5, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Skywork Team Introduces Skywork-MoE: A High-Performance Mixture-of-Experts (MoE) Model with 146B Parameters, 16 Experts, and 22B Activated Parameters
ShareShareShareShareShare

The development of large language models (LLMs) has been a focal point in advancing NLP capabilities. However, training these models poses substantial challenges due to the immense computational resources and costs involved. Researchers continuously explore more efficient methods to manage these demands while maintaining high performance.

A critical issue in LLM development is the extensive resources needed for training dense models. Dense models activate all parameters for each input token, leading to significant inefficiencies. This approach makes it difficult to scale up without incurring prohibitive costs. Consequently, there is a pressing need for more resource-efficient training methods that can still deliver competitive performance. The primary goal is to balance computational feasibility and the ability to handle complex NLP tasks effectively.

Traditionally, LLM training has relied on dense, resource-intensive models despite their high performance. These models require the activation of every parameter for each token, leading to a substantial computational load. Sparse models, such as Mixture-of-Experts (MoE), have emerged as a promising alternative. MoE models distribute computational tasks across several specialized sub-models or “experts.” This approach can match or surpass dense models’ performance using a fraction of the resources. The efficiency of MoE models lies in their ability to selectively activate only a subset of the experts for each token, thus optimizing resource usage.

The Skywork Team, Kunlun Inc. research team introduced Skywork-MoE, a high-performance MoE large language model with 146 billion parameters and 16 experts. This model builds on the foundational architecture of their previously developed Skywork-13B model, utilizing its dense checkpoints as the initial setup. The Skywork-MoE incorporates two novel training techniques: gating logit normalization and adaptive auxiliary loss coefficients. These innovations are designed to enhance the model’s efficiency and performance. By leveraging dense checkpoints, the model benefits from pre-existing data, which aids in the initial setup and subsequent training phases.

Skywork-MoE was trained using dense checkpoints from the Skywork-13B model, initialized from dense models pre-trained for 3.2 trillion tokens, and further trained on an additional 2 trillion tokens. The gating logit normalization technique ensures a distinct gate output distribution, which enhances export diversification. This method involves normalizing the gating layer outputs before applying the softmax function, which helps achieve a sharper and more focused distribution. The adaptive auxiliary loss coefficients allow for layer-specific adjustment, maintaining a balanced load across experts and preventing any single expert from becoming overloaded. These adjustments are based on monitoring the token drop rate and adapting the coefficients accordingly.

The performance of Skywork-MoE was evaluated across a variety of benchmarks. The model scored 82.2 on the CEVAL benchmark and 79.5 on the CMMLU benchmark, surpassing the Deepseek-67B model. The MMLU benchmark scored 77.4, which is competitive compared to higher-capacity models like Qwen1.5-72B. For mathematical reasoning tasks, Skywork-MoE scored 76.1 on GSM8K and 31.9 on MATH, comfortably outperforming models like Llama2-70B and Mixtral 8*7B. Skywork-MoE demonstrated robust performance in code synthesis tasks with a score of 43.9 on the HumanEval benchmark, exceeding all dense models in the comparison and slightly trailing behind the Deepseek-V2 model. These results highlight the model’s ability to effectively handle complex quantitative and logical reasoning tasks.

In conclusion, the research team from the Skywork team successfully addressed the issue of resource-intensive LLM training by developing Skywork-MoE, which leverages innovative techniques to enhance performance while reducing computational demands. Skywork-MoE, with its 146 billion parameters and advanced training methodologies, stands as a significant advancement in the field of NLP. The model’s strong performance across various benchmarks underscores the effectiveness of the gating logit normalization and adaptive auxiliary loss coefficients techniques. This research competes well with existing models and sets a new benchmark for the efficiency and efficacy of MoE models in large-scale language processing tasks.


YOU MAY ALSO LIKE

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

eGPUs Do Work, But They Come With Some Notable Limitations

Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…

Credit: Source link

ShareTweetSendSharePin

Related Posts

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI
AI & Technology

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

September 18, 2026
eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
Next Post
Meet the Texas six-year-old savant who joined a high-IQ society

Meet the Texas six-year-old savant who joined a high-IQ society

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

September 14, 2026
Smithsonian Secretary Lonnie Bunch to retire amid clashes with Trump administration

Smithsonian Secretary Lonnie Bunch to retire amid clashes with Trump administration

September 15, 2026
Apple unveils foldable iPhone

Apple unveils foldable iPhone

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!