• bitcoinBitcoin(BTC)$79,646.00-1.91%
  • ethereumEthereum(ETH)$2,450.97-1.89%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$718.87-0.74%
  • rippleXRP(XRP)$1.40-4.33%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$101.76-2.29%
  • tronTRON(TRX)$0.3317060.33%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.041.29%
  • HyperliquidHyperliquid(HYPE)$84.11-2.21%
  • zcashZcash(ZEC)$1,025.738.03%
  • dogecoinDogecoin(DOGE)$0.084594-3.53%
  • RainRain(RAIN)$0.016561-3.35%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$517.24-0.56%
  • chainlinkChainlink(LINK)$11.60-1.88%
  • whitebitWhiteBIT Coin(WBT)$73.19-1.11%
  • leo-tokenLEO Token(LEO)$9.21-1.25%
  • cardanoCardano(ADA)$0.210399-4.96%
  • stellarStellar(XLM)$0.179038-2.76%
  • bitcoin-cashBitcoin Cash(BCH)$249.41-3.22%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.106777-5.46%
  • litecoinLitecoin(LTC)$50.72-1.13%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.01%
  • uniswapUniswap(UNI)$6.12-4.59%
  • hedera-hashgraphHedera(HBAR)$0.077703-1.78%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.35-2.17%
  • suiSui(SUI)$0.75-3.75%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.32%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.138.94%
  • tether-goldTether Gold(XAUT)$4,427.14-0.87%
  • crypto-com-chainCronos(CRO)$0.055713-3.63%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • MemeCoreMemeCore(M)$1.126.76%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$108.17-0.95%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.20%
  • BittensorBittensor(TAO)$222.69-1.36%
  • aaveAave(AAVE)$129.85-2.45%
  • AsterAster(ASTER)$0.742.71%
  • pax-goldPAX Gold(PAXG)$4,432.86-0.94%
  • mantleMantle(MNT)$0.570.55%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056132-2.26%
  • MorphoMorpho(MORPHO)$2.563.06%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Huawei Researchers Develop Pangu-Σ: A Large Language Model With Sparse Architecture And 1.085 Trillion Parameters

July 11, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Huawei Researchers Develop Pangu-Σ: A Large Language Model With Sparse Architecture And 1.085 Trillion Parameters
ShareShareShareShareShare

Large Language Models (LLMs) have exhibited exceptional skills and potential in natural language processing, creation, and reasoning. By employing a large quantity of textual data, the performance of language models scales up with compute budget and model parameters, displaying significant zero/few-shot learning skills or even emerging abilities. Since GPT-3, several big language models have been developed and published, including the Megatron-Turing NLG, PanGu, ERNIE 3.0 Titan, Gopher, PaLM, OPT, Bloom, and GLM-130B. With more than one trillion parameters, researchers have begun constructing ever bigger language models. Generally, sparsely-activated models like Mixture-of-Experts (MoE) are used to achieve this.

Several notable works among the trillion-parameter models are available, including Switch-C, GLaM, MoE-1.1T, Wu Dao 2.0, and M6-10T. Unfortunately, only a chosen number have achieved the expected performance while publishing thorough assessment findings across various jobs. According to their observations, scaling efficiency is the main challenge. Current research on the scaling laws of language models shows that for LLMs to function at their best, there must be an adequate amount of training data and a reasonable computing budget. Designing a scalable model architecture and an effective distributed training system that can ingest the data with high training throughput is, therefore, one of the key motivations for this effort.

• Scaling the model: LLM model performance is anticipated to increase as the model size grows. Sparse architectures like a Mixture of Experts (MoE) are an intriguing option to scale the model size up without incurring a linear rise in computational cost compared to the high computational price for training dense Transformer models. Yet, issues such as an imbalanced workload and global communication delay plague MoE models. Also, there are still unresolved issues with adding MoE to an existing dense model and how many experts to place in each layer. Thus, developing a trillion-parameter sparse model with good performance and training efficiency is a critical but difficult challenge.

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!

• Scaling the system: It has been suggested to use frameworks like DeepSpeed 4 to enable training models with a trillion parameters. The primary constraint is frequently a constrained compute budget, or more precisely, the number of accelerating devices (such as GPU, NPU, and TPU) that may be employed. Practitioners may train trillion-parameter models with workable batch sizes using tensor parallelism, pipeline parallelism, zero redundancy optimizer, and rematerialization over thousands of accelerating devices. By using heterogeneous computing strategies, such as shifting a portion of the processing to host machines, practitioners can minimize the number of computing resources.

However, the poor bandwidth between the host and device and the CPUs’ limited computational power compared to accelerating devices make it impossible to feed big language models with a sufficient quantity of data and achieve optimal performance using the present methodologies. Consequently, the effectiveness of big language models depends on how to scale the system performance with a restricted computing budget. In this paper, researchers from Huawei introduce Pangu-Σ a large language model with sparse architecture and 1.085 trillion parameters. They create the Pangu-Σmodel within the MindSpore 5 framework and train it over 100 days on a cluster using 512 Ascend 910 AI Accelerators and 329 billion tokens.

PanGu’s built-in parameters are expanded using Random Routed Experts’ Transformer decoder architecture (RRE). RRE uses two levels of routing as opposed to traditional MoE. Experts are organized by task or domain at the first level, and tokens are evenly and randomly assigned to each group at the second level without using any learnable gating functions as in MoE. Using the RRE architecture, it is simple to extract sub-models from the Pangu-Σ for various downstream applications, including conversation, translation, code production, and interpreting natural language in general.

They suggest the Expert Computation and Storage Separation (ECSS) mechanism to make training systems efficient and scalable. This mechanism achieves 69905 tokens/s observed throughput in training 1.085 trillion Pangu-Σ on a cluster of 512 Ascend 910 accelerators and significantly reduces host-to-device and device-to-host communication as optimizer update computation. Overall, the training throughput is 6.3 times faster than it was for the model with the MoE architecture but with the same hyperparameters.

The sub-modal of Pangu-Σ in the Chinese domain significantly outperforms the previous SOTA models, including Pangu-Σ with 13B parameters and ERNIE 3.0 Titan with 260B parameters over 16 downstream tasks in six categories in the zero-shot setting without any multitask finetuning or instruction tuning. The Pangu-Σ model performs better in the relevant regions than the SOTA models. It uses 329B tokens in more than 40 natural and programming languages. Moreover, they evaluate how well Pangu-Σ has been tweaked in several application domains, including conversation, machine translation, and code production.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 16k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

Flock Cameras Are Officially Banned On State Roads In Florida

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Commits B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI
AI & Technology

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

September 4, 2026
Flock Cameras Are Officially Banned On State Roads In Florida
AI & Technology

Flock Cameras Are Officially Banned On State Roads In Florida

September 4, 2026
Researchers Document OpenAI Agent Swarm That Repurposed German Wiki – Unite.AI
AI & Technology

Researchers Document OpenAI Agent Swarm That Repurposed German Wiki – Unite.AI

September 4, 2026
Microsoft Brings OpenAI’s GPT-6 Astra to Foundry With Limited Access – Unite.AI
AI & Technology

Microsoft Brings OpenAI’s GPT-6 Astra to Foundry With Limited Access – Unite.AI

September 4, 2026
Next Post
Charter Communications to Acquire Time Warner Cable, Sending Shares of Both Cable Giants Higher

Charter Communications to Acquire Time Warner Cable, Sending Shares of Both Cable Giants Higher

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Bryan Kohberger files petition to change guilty plea

Bryan Kohberger files petition to change guilty plea

September 4, 2026
Revolution Medicines: Strong Buy RASONQUE FDA Approval And Zoldonrasib PDAC Expansions

Revolution Medicines: Strong Buy RASONQUE FDA Approval And Zoldonrasib PDAC Expansions

August 31, 2026
What You Should Look For In A USB To Bluetooth Adapter For PC

What You Should Look For In A USB To Bluetooth Adapter For PC

August 29, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!