• bitcoinBitcoin(BTC)$79,461.00-2.11%
  • ethereumEthereum(ETH)$2,454.17-2.14%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$716.67-0.94%
  • rippleXRP(XRP)$1.40-4.21%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.52-3.31%
  • tronTRON(TRX)$0.330203-0.36%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.09%
  • HyperliquidHyperliquid(HYPE)$84.600.04%
  • zcashZcash(ZEC)$983.974.00%
  • dogecoinDogecoin(DOGE)$0.084443-4.82%
  • RainRain(RAIN)$0.016588-2.51%
  • moneroMonero(XMR)$522.45-0.80%
  • USDSUSDS(USDS)$1.00-0.03%
  • chainlinkChainlink(LINK)$11.66-0.87%
  • whitebitWhiteBIT Coin(WBT)$73.08-1.52%
  • leo-tokenLEO Token(LEO)$9.29-0.69%
  • cardanoCardano(ADA)$0.213387-2.58%
  • stellarStellar(XLM)$0.179150-3.58%
  • bitcoin-cashBitcoin Cash(BCH)$251.37-2.32%
  • daiDai(DAI)$1.00-0.03%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • CantonCanton(CC)$0.108016-4.59%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.291.02%
  • litecoinLitecoin(LTC)$50.26-1.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.07%
  • hedera-hashgraphHedera(HBAR)$0.077368-1.92%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$7.37-1.57%
  • suiSui(SUI)$0.75-4.65%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.19%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056181-0.76%
  • tether-goldTether Gold(XAUT)$4,438.28-0.99%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • nearNEAR Protocol(NEAR)$1.94-2.28%
  • MemeCoreMemeCore(M)$1.113.80%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$107.92-2.38%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.12%
  • BittensorBittensor(TAO)$223.78-2.61%
  • aaveAave(AAVE)$131.43-1.78%
  • AsterAster(ASTER)$0.730.18%
  • pax-goldPAX Gold(PAXG)$4,444.57-1.06%
  • mantleMantle(MNT)$0.571.04%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057120-1.60%
  • MorphoMorpho(MORPHO)$2.521.82%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Stanford Research Introduces FlashAttention-2: A Leap in Speed and Efficiency for Long-Context Language Models

July 20, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Stanford Research Introduces FlashAttention-2: A Leap in Speed and Efficiency for Long-Context Language Models
ShareShareShareShareShare

In the past year, natural language processing has seen remarkable advancements with the emergence of language models equipped with significantly longer contexts. Among these models are GPT-4 with a context length of 32k, MosaicML’s MPT with 65k context, and Anthropic’s Claude, boasting an impressive 100k context length. As applications such as long document querying and story writing continue to grow, the need for language models with extended context becomes evident. However, the challenge lies in scaling up the context length of Transformers, as their attention layer has computational and memory requirements that grow quadratically with the input sequence length.

Addressing this challenge, FlashAttention, an innovative algorithm released just a year ago, gained rapid adoption across various organizations and research labs. This algorithm successfully accelerated attention computation while reducing its memory footprint without sacrificing accuracy or approximating the results. With 2-4 times faster performance than optimized baselines at its initial release, FlashAttention proved to be a groundbreaking advancement. Yet, it still had untapped potential, as it fell short of the blazing-fast optimized matrix-multiply (GEMM) operations that achieved up to 124 TFLOPs/s on A100 GPUs.

Taking the next leap forward, the developers of FlashAttention have now introduced FlashAttention-2, a reinvented version that significantly surpasses its predecessor. Leveraging Nvidia’s CUTLASS 3.x and CuTe core library, FlashAttention-2 achieves a remarkable 2x speedup, reaching up to 230 TFLOPs/s on A100 GPUs. Moreover, in end-to-end training of GPT-style language models, FlashAttention-2 attains a training speed of up to 225 TFLOPs/s, with an impressive 72% model FLOP utilization.

🚀 Build high-quality training datasets with Kili Technology and solve NLP machine learning challenges to develop powerful ML applications

The key enhancements of FlashAttention-2 lie in its better parallelism and work partitioning strategies. Initially, FlashAttention parallelized over batch size and number of heads, effectively utilizing the compute resources on the GPU. However, for long sequences with smaller batch sizes or fewer heads, FlashAttention-2 now parallelizes over the sequence length dimension, resulting in significant speedup in these scenarios.

Another improvement involves efficiently partitioning work between different warps within each thread block. In FlashAttention, splitting K and V across four warps while keeping Q accessible by all warps, referred to as the “sliced-K” scheme, led to unnecessary shared memory reads and writes, slowing down the computation. FlashAttention-2 takes a different approach, now splitting Q across four warps while keeping K and V accessible to all warps. This eliminates the need for communication between warps and significantly reduces shared memory reads/writes, further boosting performance.

FlashAttention-2 introduces several new features to broaden its applicability and enhance its capabilities. It now supports head dimensions up to 256, accommodating models like GPT-J, CodeGen, CodeGen2, and StableDiffusion 1.x, opening up more speedup and memory-saving opportunities. Additionally, FlashAttention-2 embraces multi-query attention (MQA) and grouped-query attention (GQA) variants, where multiple heads of the query can attend to the same head of key and value, leading to higher inference throughput and better performance.

The performance of FlashAttention-2 is truly impressive. Benchmarked on an A100 80GB SXM4 GPU, it achieves around 2x speedup compared to its predecessor and up to 9x speedup compared to a standard attention implementation in PyTorch. Moreover, when used for end-to-end training of GPT-style models, FlashAttention-2 unlocks up to 225 TFLOPs/s on A100 GPUs, representing a 1.3x end-to-end speedup over already highly optimized models with FlashAttention.

Looking ahead, the potential applications of FlashAttention-2 are promising. With the ability to train models with 16k longer context for the same price as previous 8k context models, this technology can help analyze long books, reports, high-resolution images, audio, and video. Plans for broader applicability on devices like H100 GPUs and AMD GPUs and optimizing for new data types like fp8 are underway. Furthermore, combining the low-level optimizations of FlashAttention-2 with high-level algorithmic changes could pave the way for training AI models with unprecedentedly longer context. Collaboration with compiler researchers to enhance programmability is also on the horizon, promising a bright future for the next generation of language models.


Check out the Paper and Github. Don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 900+ AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking

How AI Turned Our Small Marketing Team into a Full-Service Agency – Unite.AI

Niharika is a Technical consulting intern at Marktechpost. She is a third year undergraduate, currently pursuing her B.Tech from Indian Institute of Technology(IIT), Kharagpur. She is a highly enthusiastic individual with a keen interest in Machine learning, Data science and AI and an avid reader of the latest developments in these fields.


🔥 Gain a competitive
edge with data: Actionable market intelligence for global brands, retailers, analysts, and investors. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking
AI & Technology

Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking

September 4, 2026
How AI Turned Our Small Marketing Team into a Full-Service Agency – Unite.AI
AI & Technology

How AI Turned Our Small Marketing Team into a Full-Service Agency – Unite.AI

September 4, 2026
A Worthy Android Ereader, With Some Tradeoffs
AI & Technology

A Worthy Android Ereader, With Some Tradeoffs

September 4, 2026
Insurance Spent Years Talking About AI. This Year It Actually Used It – Unite.AI
AI & Technology

Insurance Spent Years Talking About AI. This Year It Actually Used It – Unite.AI

September 4, 2026
Next Post
‘Bloomberg Technology’ Full Show (01/13/2020)

'Bloomberg Technology' Full Show (01/13/2020)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
FBI investigating possible ties between Iran and cyberattacks on Minnesota water systems

FBI investigating possible ties between Iran and cyberattacks on Minnesota water systems

September 1, 2026
Trump says he’ll revive ‘anti-weaponization’ fund if Blanche’s nomination for AG is blocked

Trump says he’ll revive ‘anti-weaponization’ fund if Blanche’s nomination for AG is blocked

August 31, 2026
OpenAI’s ‘rogue’ agents hacked into more systems than initially reported

OpenAI’s ‘rogue’ agents hacked into more systems than initially reported

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!