• bitcoinBitcoin(BTC)$77,226.00-2.14%
  • ethereumEthereum(ETH)$2,414.05-2.68%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$679.25-1.79%
  • rippleXRP(XRP)$1.35-2.32%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$99.88-3.67%
  • tronTRON(TRX)$0.323748-2.96%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-3.93%
  • HyperliquidHyperliquid(HYPE)$82.35-1.77%
  • zcashZcash(ZEC)$830.61-3.47%
  • dogecoinDogecoin(DOGE)$0.081915-1.66%
  • RainRain(RAIN)$0.016401-2.28%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$503.20-4.02%
  • leo-tokenLEO Token(LEO)$9.37-2.47%
  • whitebitWhiteBIT Coin(WBT)$71.07-2.25%
  • chainlinkChainlink(LINK)$11.20-1.87%
  • cardanoCardano(ADA)$0.196236-1.40%
  • stellarStellar(XLM)$0.176299-1.20%
  • bitcoin-cashBitcoin Cash(BCH)$246.59-0.26%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.114100-7.77%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.431.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-4.16%
  • uniswapUniswap(UNI)$5.658.35%
  • Global DollarGlobal Dollar(USDG)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0743431.30%
  • avalanche-2Avalanche(AVAX)$7.230.26%
  • shiba-inuShiba Inu(SHIB)$0.0000052.15%
  • suiSui(SUI)$0.72-0.35%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054945-2.06%
  • tether-goldTether Gold(XAUT)$4,330.37-2.49%
  • nearNEAR Protocol(NEAR)$1.901.64%
  • MemeCoreMemeCore(M)$1.06-2.82%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$110.17-1.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.10%
  • BittensorBittensor(TAO)$222.62-2.76%
  • aaveAave(AAVE)$125.791.27%
  • pax-goldPAX Gold(PAXG)$4,338.27-2.51%
  • AsterAster(ASTER)$0.69-1.21%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056680-2.39%
  • mantleMantle(MNT)$0.53-5.54%
  • MorphoMorpho(MORPHO)$2.53-4.65%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention

July 12, 2026
in AI & Technology
Reading Time: 4 mins read
A A
A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention
ShareShareShareShareShare

YOU MAY ALSO LIKE

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

if BACKEND == "triton":
   @triton.jit
   def _vadd_kernel(a_ptr, b_ptr, c_ptr, n, BLOCK: tl.constexpr):
       pid  = tl.program_id(0)
       offs = pid * BLOCK + tl.arange(0, BLOCK)
       mask = offs < n
       a = tl.load(a_ptr + offs, mask=mask)
       b = tl.load(b_ptr + offs, mask=mask)
       tl.store(c_ptr + offs, a + b, mask=mask)
   @triton.jit
   def _fused_gelu_kernel(x_ptr, w_ptr, b_ptr, o_ptr, n, BLOCK: tl.constexpr):
       pid  = tl.program_id(0)
       offs = pid * BLOCK + tl.arange(0, BLOCK)
       mask = offs < n
       x = tl.load(x_ptr + offs, mask=mask)
       w = tl.load(w_ptr + offs, mask=mask)
       b = tl.load(b_ptr + offs, mask=mask)
       h = x * w + b
       c = 0.7978845608028654
       z = c * (h + 0.044715 * h * h * h)
       e = tl.exp(-2.0 * z)
       tanh = (1.0 - e) / (1.0 + e)
       g = 0.5 * h * (1.0 + tanh)
       tl.store(o_ptr + offs, g, mask=mask)
   @triton.jit
   def _softmax_kernel(x_ptr, o_ptr, stride, n_cols, BLOCK: tl.constexpr):
       row  = tl.program_id(0)
       cols = tl.arange(0, BLOCK)
       mask = cols < n_cols
       ptr  = x_ptr + row * stride + cols
       x    = tl.load(ptr, mask=mask, other=-float("inf"))
       x    = x - tl.max(x, axis=0)
       num  = tl.exp(x)
       den  = tl.sum(num, axis=0)
       tl.store(o_ptr + row * stride + cols, num / den, mask=mask)
   @triton.jit
   def _matmul_kernel(A, B, C, M, N, K,
                      sam, sak, sbk, sbn, scm, scn,
                      BM: tl.constexpr, BN: tl.constexpr, BK: tl.constexpr):
       pid_m = tl.program_id(0)
       pid_n = tl.program_id(1)
       offs_m = pid_m * BM + tl.arange(0, BM)
       offs_n = pid_n * BN + tl.arange(0, BN)
       offs_k = tl.arange(0, BK)
       a_ptr = A + offs_m[:, None] * sam + offs_k[None, :] * sak
       b_ptr = B + offs_k[:, None] * sbk + offs_n[None, :] * sbn
       acc = tl.zeros((BM, BN), dtype=tl.float32)
       for k in range(0, K, BK):
           a = tl.load(a_ptr, mask=offs_k[None, :] < K - k, other=0.0)
           b = tl.load(b_ptr, mask=offs_k[:, None] < K - k, other=0.0)
           acc += tl.dot(a, b)
           a_ptr += BK * sak
           b_ptr += BK * sbk
       c_ptr = C + offs_m[:, None] * scm + offs_n[None, :] * scn
       cmask = (offs_m[:, None] < M) & (offs_n[None, :] < N)
       tl.store(c_ptr, acc.to(C.dtype.element_ty), mask=cmask)
   @triton.jit
   def _flash_kernel(Q, K, V, O, sqz, skz, svz, soz,
                     L, D, scale,
                     BL: tl.constexpr, BD: tl.constexpr):
       pid_l = tl.program_id(0)
       z     = tl.program_id(1)
       offs_l = pid_l * BL + tl.arange(0, BL)
       offs_d = tl.arange(0, BD)
       q_ptr = Q + z * sqz + offs_l[:, None] * D + offs_d[None, :]
       q = tl.load(q_ptr, mask=offs_l[:, None] < L, other=0.0)
       m_i = tl.full((BL,), -float("inf"), dtype=tl.float32)
       l_i = tl.zeros((BL,), dtype=tl.float32)
       acc = tl.zeros((BL, BD), dtype=tl.float32)
       for start in range(0, L, BL):
           offs_k = start + tl.arange(0, BL)
           k_ptr = K + z * skz + offs_k[:, None] * D + offs_d[None, :]
           v_ptr = V + z * svz + offs_k[:, None] * D + offs_d[None, :]
           k = tl.load(k_ptr, mask=offs_k[:, None] < L, other=0.0)
           v = tl.load(v_ptr, mask=offs_k[:, None] < L, other=0.0)
           s = tl.dot(q, tl.trans(k)) * scale
           s = tl.where(offs_k[None, :] < L, s, -float("inf"))
           m_ij = tl.maximum(m_i, tl.max(s, axis=1))
           p    = tl.exp(s - m_ij[:, None])
           alpha = tl.exp(m_i - m_ij)
           l_i  = l_i * alpha + tl.sum(p, axis=1)
           acc  = acc * alpha[:, None] + tl.dot(p.to(v.dtype), v)
           m_i  = m_ij
       acc = acc / l_i[:, None]
       o_ptr = O + z * soz + offs_l[:, None] * D + offs_d[None, :]
       tl.store(o_ptr, acc.to(O.dtype.element_ty), mask=offs_l[:, None] < L)
   def run_vadd(a, b):
       c = torch.empty_like(a); n = a.numel()
       grid = (triton.cdiv(n, 1024),)
       _vadd_kernel[grid](a, b, c, n, BLOCK=1024)
       return c
   def run_fused_gelu(x, w, b):
       o = torch.empty_like(x); n = x.numel()
       grid = (triton.cdiv(n, 1024),)
       _fused_gelu_kernel[grid](x, w, b, o, n, BLOCK=1024)
       return o
   def run_softmax(x):
       m, ncols = x.shape
       o = torch.empty_like(x)
       BLOCK = triton.next_power_of_2(ncols)
       _softmax_kernel[(m,)](x, o, x.stride(0), ncols, BLOCK=BLOCK)
       return o
   def run_matmul(a, b):
       M, K = a.shape; K2, N = b.shape
       c = torch.empty((M, N), device=a.device, dtype=a.dtype)
       BM = BN = 64; BK = 32
       grid = (triton.cdiv(M, BM), triton.cdiv(N, BN))
       _matmul_kernel[grid](a, b, c, M, N, K,
                            a.stride(0), a.stride(1), b.stride(0), b.stride(1),
                            c.stride(0), c.stride(1), BM=BM, BN=BN, BK=BK)
       return c
   def run_flash(q, k, v):
       Z, L, D = q.shape
       o = torch.empty_like(q)
       scale = 1.0 / math.sqrt(D)
       BL = 64
       grid = (triton.cdiv(L, BL), Z)
       _flash_kernel[grid](q, k, v, o,
                           q.stride(0), k.stride(0), v.stride(0), o.stride(0),
                           L, D, scale, BL=BL, BD=D)
       return o
else:
   def run_vadd(a, b):            return a + b
   def run_fused_gelu(x, w, b):   return torch.nn.functional.gelu(x * w + b, approximate="tanh")
   def run_softmax(x):            return torch.softmax(x, dim=-1)
   def run_matmul(a, b):          return a @ b
   def run_flash(q, k, v):        return torch.nn.functional.scaled_dot_product_attention(q, k, v)

Credit: Source link

ShareTweetSendSharePin

Related Posts

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways
AI & Technology

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

September 1, 2026
Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI
AI & Technology

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

September 1, 2026
Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance
AI & Technology

Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

September 1, 2026
AI is redefining the workforce — and most planning models aren’t ready
AI & Technology

AI is redefining the workforce — and most planning models aren’t ready

September 1, 2026
Next Post
Illinois representative talks bill that would regulate AI companies

Illinois representative talks bill that would regulate AI companies

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
When Content Is Infinite, Point of View Becomes the Scarce Asset – Unite.AI

When Content Is Infinite, Point of View Becomes the Scarce Asset – Unite.AI

August 28, 2026
The labor market was weaker early in Trump’s term, data shows – The Washington Post

The labor market was weaker early in Trump’s term, data shows – The Washington Post

August 29, 2026
Nvidia’s Sales Forecast Fails to Meet Loftiest Estimates

Nvidia’s Sales Forecast Fails to Meet Loftiest Estimates

August 29, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!