• bitcoinBitcoin(BTC)$76,285.000.48%
  • ethereumEthereum(ETH)$2,433.321.16%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$721.701.37%
  • rippleXRP(XRP)$1.300.51%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.942.82%
  • tronTRON(TRX)$0.334273-0.19%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.64%
  • zcashZcash(ZEC)$1,352.4111.54%
  • HyperliquidHyperliquid(HYPE)$79.610.72%
  • dogecoinDogecoin(DOGE)$0.0807700.96%
  • USDSUSDS(USDS)$1.000.04%
  • RainRain(RAIN)$0.013204-3.48%
  • moneroMonero(XMR)$498.97-0.73%
  • whitebitWhiteBIT Coin(WBT)$78.510.64%
  • chainlinkChainlink(LINK)$11.163.16%
  • leo-tokenLEO Token(LEO)$8.910.37%
  • cardanoCardano(ADA)$0.1978861.78%
  • stellarStellar(XLM)$0.1812493.31%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.432.48%
  • USD1USD1(USD1)$1.00-0.02%
  • uniswapUniswap(UNI)$6.868.52%
  • litecoinLitecoin(LTC)$52.744.12%
  • CantonCanton(CC)$0.10041410.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.30%
  • nearNEAR Protocol(NEAR)$2.8215.47%
  • avalanche-2Avalanche(AVAX)$7.543.70%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0743200.02%
  • shiba-inuShiba Inu(SHIB)$0.0000053.09%
  • suiSui(SUI)$0.724.25%
  • crypto-com-chainCronos(CRO)$0.0578744.29%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • tether-goldTether Gold(XAUT)$4,323.39-0.49%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$226.454.43%
  • MemeCoreMemeCore(M)$1.132.70%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$111.771.38%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.34%
  • AsterAster(ASTER)$0.748.97%
  • aaveAave(AAVE)$122.682.57%
  • BitwayBitway(BTW)$0.70-10.90%
  • pax-goldPAX Gold(PAXG)$4,326.87-0.54%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0591713.99%
  • mantleMantle(MNT)$0.563.28%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

September 17, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI
ShareShareShareShareShare

Z.ai on September 17, 2026 published a technical account describing how it built a complete production-grade inference service for its GLM-5.3-Flash model from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. According to the company, much of the work was carried out by an Infra Agent powered by GLM-5.3 rather than by infrastructure engineers alone, and all production inference for GLM-5.3-Flash runs on the system.

YOU MAY ALSO LIKE

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

An iOS 27 Bug Can Temporarily Freeze Your iPhone

Z.ai said no one had previously operated a cluster of Chinese-made accelerators at this scale. The company cited relatively limited on-chip memory capacity and bandwidth, a new model architecture, a 1M-token context window, and multimodal requests, alongside an immature ecosystem in which kernel support was incomplete and engineers had to guess at behavior that should have been documented.

GLM-5.3-Flash launched on August 26, 2026 as the first natively multimodal model in the GLM-5 series, with 320 billion total parameters and 18 billion active parameters under a hybrid architecture combining sparse and linear attention. Before release, Z.ai tested the model anonymously as ox-alpha on OpenCode and OpenRouter, and the company said it became the most-used model on both platforms within a week of launch, processing more than 62 trillion tokens in six days.

The Dense Feedback Method

At the center of the account is a systems problem: end-to-end metrics can tell an agent that results got worse, but not why. A failed numerical accuracy test, a 30% increase in time-to-first-token, or a 20% drop in output throughput does not show which layer is responsible or what to test next. Z.ai’s answer, which it calls dense feedback, folds correctness tests, runtime logs, execution traces, runtime events, microbenchmarks, and end-to-end metrics into repeatable workflows that let the agent validate each hypothesis locally rather than wait for a full deployment and load test after every change.

The company defines three required properties for such feedback. It must be local, tied wherever possible to specific launch parameters, code changes, kernels, input conditions, threads, execution intervals, or code paths. It must be inexpensive and timely to obtain. And it must support objective verification through reference implementations and controlled experiments, because observed correlations by themselves do not establish a root cause.

In the launch loop the account describes, engineers defined objectives and system boundaries and reviewed critical changes involving numerical semantics, concurrency behavior, and production risk, while the agent handled analysis, hypotheses, and code changes. The stack they jointly optimized combined intra-node tensor parallelism for linear attention and the LM Head, ReplaySSM, W8A8 quantization, mixed-precision INT8/FP8/BF16 cache quantization, and Layer Split, under an Encode-Prefill-Decode disaggregated architecture.

Three Engineering Cases

The first case concerns numerical correctness. Validation comparing partitioned and unpartitioned kernel execution paths exposed an accuracy problem in the KDA kernel’s Context Parallelism path: the tl.dot operation defaulted to TF32 computation even when its inputs were FP32, so errors accumulated during state merging and compounded as context length increased. The fix explicitly set the input precision to tf32x3, which uses three TF32 Tensor Core operations to yield a higher-precision result.

According to the account, the fixes were merged upstream into Flash Linear Attention. The pull request, opened and merged on August 27, 2026, applies the tf32x3 affine chain in the state-update and transformation-merging kernels as an opt-in accuracy path, adds Context Parallelism tests, and falls back explicitly to ieee precision on platforms without tf32 support (AMD, NPUs, and NVIDIA GPUs below compute capability 8.0).

The second case concerns a KV Transfer concurrency bottleneck. According to the account, engineers set an acceptance criterion that, under the same workload, Prefill plus KV Transfer should run within 5% of the Prefill-only baseline. The agent found gaps exceeding 20% in some scenarios and traced them to DeepEP v1.2.1, in which neither the intranodedispatch nor the intranodecombine call explicitly released the Python GIL. As long as those calls held the lock, the Mooncake Transfer Python thread in the same process could not acquire the GIL in time, so scheduling and submission of transfer tasks slipped and overlap with computation shrank. The account notes that internode_dispatch in the same version already released the GIL, with a code comment stating the intent was to avoid blocking KV Transfer in other threads while the CPU waited. After the fix released the GIL during the relevant C++ execution intervals, the account reports, the gap fell below 1% under the same test conditions.

The third case concerns kernel performance. Z.ai had the agent distill techniques from handwritten kernels in projects including SGLang, Flash Linear Attention, and DeepGEMM into reusable optimization skeletons carrying applicability conditions, transformation methods, resource constraints, and validation evidence. On a representative KDA Decode kernel, the company reports that the agent’s division optimization cut execution time by 9.6%. After feedback identified computation as the primary bottleneck, the agent merged the kernel’s V-dimension tiles, which had repeated the same FP32 normalization and gating computations four times, into a single thread block with register-resident intermediate results and one warp-level reduction, producing what the company reports as a 1.71× speedup over the prior version.

Stated Results and Recursive Self-Improvement

Z.ai reports that GLM-5.3-Flash went from initial model adaptation to production readiness in less than two weeks, with end-to-end throughput ultimately tripling relative to the initial baseline. The company also said hardware utilization efficiency and per-token cost reached levels comparable to mainstream NVIDIA GPUs.

The company frames the effort as an early example of recursive self-improvement, noting that the model participated in optimizing the inference system on which it runs. At the same time, Z.ai states it has not yet reached recursive self-improvement, and that choosing objectives, setting boundaries, and assessing risk remain human responsibilities it believes humans should continue to hold.

Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
AI & Technology

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

September 17, 2026
An iOS 27 Bug Can Temporarily Freeze Your iPhone
AI & Technology

An iOS 27 Bug Can Temporarily Freeze Your iPhone

September 17, 2026
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
AI & Technology

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Next Post
Meet the Press NOW — September 4

Meet the Press NOW — September 4

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
At least 10 killed in central Mexico fireworks blast

At least 10 killed in central Mexico fireworks blast

September 16, 2026
House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Shoulder Innovations, Inc. (SI) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

Shoulder Innovations, Inc. (SI) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!