• bitcoinBitcoin(BTC)$75,627.00-0.44%
  • ethereumEthereum(ETH)$2,393.650.07%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$714.99-0.12%
  • rippleXRP(XRP)$1.28-4.53%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$97.52-1.00%
  • tronTRON(TRX)$0.3351300.83%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.52%
  • zcashZcash(ZEC)$1,338.5420.51%
  • HyperliquidHyperliquid(HYPE)$79.383.81%
  • dogecoinDogecoin(DOGE)$0.079599-1.34%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$491.93-2.03%
  • whitebitWhiteBIT Coin(WBT)$77.66-0.81%
  • RainRain(RAIN)$0.012704-11.27%
  • leo-tokenLEO Token(LEO)$8.85-0.46%
  • chainlinkChainlink(LINK)$10.80-2.55%
  • cardanoCardano(ADA)$0.192358-2.99%
  • stellarStellar(XLM)$0.177881-3.27%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$216.45-1.15%
  • uniswapUniswap(UNI)$6.350.55%
  • litecoinLitecoin(LTC)$50.56-1.54%
  • CantonCanton(CC)$0.092314-0.44%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.29-1.78%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.517.49%
  • avalanche-2Avalanche(AVAX)$7.27-1.78%
  • hedera-hashgraphHedera(HBAR)$0.072781-4.33%
  • suiSui(SUI)$0.69-0.30%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.80%
  • crypto-com-chainCronos(CRO)$0.055596-0.81%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,283.38-0.34%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.120.31%
  • BittensorBittensor(TAO)$216.55-1.86%
  • Ripple USDRipple USD(RLUSD)$1.00-0.03%
  • okbOKB(OKB)$109.43-0.77%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.06%
  • BitwayBitway(BTW)$0.757.26%
  • pax-goldPAX Gold(PAXG)$4,286.01-0.33%
  • AsterAster(ASTER)$0.690.74%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0570140.59%
  • mantleMantle(MNT)$0.54-0.68%
  • aaveAave(AAVE)$116.05-6.24%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Linear Attention Sequence Parallel (LASP): An Efficient Machine Learning Method Tailored to Linear Attention-Based Language Models

April 7, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Linear Attention Sequence Parallel (LASP): An Efficient Machine Learning Method Tailored to Linear Attention-Based Language Models
ShareShareShareShareShare

Linear attention-based models are gaining attention for their faster processing speed and comparable performance to Softmax transformers. However, large language models (LLMs), due to their large size and longer sequence lengths, exert significant strain on contemporary GPU hardware because a single GPU’s memory confines a language model’s maximum sequence length.

Sequence Parallelism (SP) techniques are often utilized to divide a long sequence into several sub-sequences and train them on multiple GPUs separately. However, current SP methods underutilize linear attention features, resulting in inefficient parallelism and usability issues. 

Researchers from Shanghai AI Laboratory and TapTap present the linear attention sequence parallel (LASP) technique, which optimizes sequence parallelism on linear transformers. It employs point-to-point (P2P) communication for efficient state exchange among GPUs within or across nodes. LASP maximizes the use of right-product kernel tricks in linear attention. Importantly, it doesn’t rely on attention head partitioning, making it adaptable to multi-head, multi-query, and grouped-query attentions.

LASP employs a tiling approach to partition input sequences into sub-sequence chunks distributed across GPUs. It distinguishes attention computation into intra-chunks and inter-chunks for utilizing linear attention’s right-product advantage. Intra-chunks use conventional attention computation, while inter-chunks exploit kernel tricks. The method also includes data distribution, forward pass, and backward pass mechanisms to enhance parallel processing efficiency.

LASP achieves significant throughput enhancement for linear attention through efficient communication design, surpassing DeepSpeed-Ulysses by 38% and Megatron by 136% in throughput at 256K sequence length on 1B model. Moreover, LASP, with system optimizations like kernel fusion and KV State caching, supports longer sequence lengths within the same cluster, reaching 2048K for the 1B model and 512K for the 7B model.

Key contributions of this research are as follows: 

  • A new SP strategy tailored to linear attention: Enabling linear attention-based models to scale for long sequences without being limited by a single GPU. 
  • Sequence length-independent communication over-head: Their elegant communication mechanism harnesses the right-product kernel trick of linear attention to ensure that the exchanging of linear attention intermediate states is sequence length-independent.
  • GPU-friendly implementation: Optimized LASP’s execution on GPUs through meticulous system engineering, including kernel fusion and KV State caching.
  • Data-parallel compatibility: LASP is compatible with all batch-level DDP methods, such as PyTorch/Legacy DDP, FSDP, and ZeRO-series optimizers.

In conclusion,  LASP is introduced to overcome the limitations of existing SP methods on linear transformers by leveraging linear attention features to enhance parallelism efficiency and usability. Implementing P2P communication, kernel fusion, and KV state caching reduces communication traffic and improves GPU cluster utilization. Compatibility with batch-level DDP methods ensures practicality for large-scale distributed training. Experiments highlight LASP’s advantages in scalability, speed, memory usage, and convergence performance compared to existing SP methods.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI

MindsEye Developer Build A Rocket Boy Is Reportedly Shutting Down

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI
AI & Technology

Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI

September 16, 2026
MindsEye Developer Build A Rocket Boy Is Reportedly Shutting Down
AI & Technology

MindsEye Developer Build A Rocket Boy Is Reportedly Shutting Down

September 16, 2026
NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results – Unite.AI
AI & Technology

NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results – Unite.AI

September 16, 2026
Samsung Brings One UI 9 To The Rest Of The Galaxy S26 Series
AI & Technology

Samsung Brings One UI 9 To The Rest Of The Galaxy S26 Series

September 16, 2026
Next Post
Jobs Data and Apple Layoffs | Bloomberg Technology

Jobs Data and Apple Layoffs | Bloomberg Technology

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Rescuers airlift missing man from a North Carolina river

Rescuers airlift missing man from a North Carolina river

September 13, 2026
Focus on Process, Not Outcome

Focus on Process, Not Outcome

September 15, 2026
Husband Only Contributes 17% To Our Household

Husband Only Contributes 17% To Our Household

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!