• bitcoinBitcoin(BTC)$64,222.001.80%
  • ethereumEthereum(ETH)$1,904.391.20%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$604.57-0.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.000.10%
  • solanaSolana(SOL)$75.700.70%
  • tronTRON(TRX)$0.331002-0.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.50%
  • HyperliquidHyperliquid(HYPE)$58.692.10%
  • dogecoinDogecoin(DOGE)$0.0701600.40%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.0131331.30%
  • zcashZcash(ZEC)$515.144.90%
  • leo-tokenLEO Token(LEO)$9.420.90%
  • moneroMonero(XMR)$410.800.00%
  • chainlinkChainlink(LINK)$9.480.80%
  • whitebitWhiteBIT Coin(WBT)$55.481.60%
  • cardanoCardano(ADA)$0.173424-1.60%
  • stellarStellar(XLM)$0.1572920.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$204.100.20%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.33-1.10%
  • CantonCanton(CC)$0.090463-5.20%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • litecoinLitecoin(LTC)$44.420.00%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • hedera-hashgraphHedera(HBAR)$0.0656730.90%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • suiSui(SUI)$0.680.00%
  • avalanche-2Avalanche(AVAX)$6.340.00%
  • tether-goldTether Gold(XAUT)$4,394.100.90%
  • shiba-inuShiba Inu(SHIB)$0.000004-0.20%
  • crypto-com-chainCronos(CRO)$0.047128-0.80%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • okbOKB(OKB)$101.19-3.30%
  • nearNEAR Protocol(NEAR)$1.620.60%
  • uniswapUniswap(UNI)$3.28-0.20%
  • pax-goldPAX Gold(PAXG)$4,409.420.80%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0605941.30%
  • BittensorBittensor(TAO)$196.03-0.50%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • AsterAster(ASTER)$0.600.20%
  • OndoOndo(ONDO)$0.3321800.60%
  • HTX DAOHTX DAO(HTX)$0.000002-0.10%
  • MemeCoreMemeCore(M)$1.162.00%
  • usddUSDD(USDD)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI

August 13, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI
ShareShareShareShareShare

Cerebras is now running OpenAI’s flagship model at a speed no GPU cloud has publicly matched. On August 13, 2026, the wafer-scale chipmaker announced it powers GPT-5.6 Sol on a new OpenAI service tier called Ultrafast, delivering up to 750 output tokens per second and, by OpenAI’s account, running the model up to 14× faster than Standard processing. Ultrafast launches first in the OpenAI API as a limited preview for a select group of customers, with access expanding as capacity grows.

YOU MAY ALSO LIKE

Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs

Why 4K Blu-Ray Always Beats 4K Streaming For Picture Quality

The claim at the center is a specific one: frontier intelligence without the speed penalty. GPT-5.6 Sol is OpenAI’s most capable model, and on Cerebras silicon it generates tokens fast enough to sit inside real-time products rather than behind an overnight batch job. Both companies frame the tier as removing the tradeoff between a model smart enough for high-stakes work and one fast enough to use while the work is still happening.

The Speed Numbers Behind Ultrafast

The 750 output tokens per second figure is the headline, but the more telling comparisons are the head-to-head runs Cerebras published. Against output speeds reported by Artificial Analysis, the company says GPT-5.6 Sol on Ultrafast runs 11× faster than Claude Fable 5 and 5× faster than Opus 4.8 on Fast mode.

Cerebras also put the tier through a full pass of Humanity’s Last Exam, a 2,500-question benchmark pitched at PhD-level difficulty. GPT-5.6 Sol on Ultrafast answered all 2,500 questions in 11 hours and 11 minutes; Claude Fable 5 needed 78 hours and 27 minutes to reach comparable conclusions: nearly 7× slower, in Cerebras’s telling. On GDP-Val, a benchmark for economically valuable knowledge work, the company reports a 5.6× end-to-end speedup over Standard processing with no quality degradation.

These are vendor-run evaluations, and Cerebras is explicit about that: the Humanity’s Last Exam comparison was benchmarked by Cerebras on July 10 and July 13–15, 2026, and the GDP-Val figure comes from its own July 31, 2026 testing. Treat them as the company’s own measurements, not independent results.

Why the Wafer-Scale Chip Wins on Latency

The mechanism matters here, because it explains why a relatively small chipmaker is serving OpenAI’s biggest model at speeds the GPU incumbents haven’t matched. Fast inference on a large model is fundamentally a data-movement problem: on GPUs, model weights must be shuttled repeatedly between on-chip memory and off-chip storage to generate each successive token, and memory bandwidth becomes the bottleneck.

Cerebras’s answer is to eliminate that movement. Its Wafer-Scale Engine packs 44 GB of SRAM onto a single wafer-sized chip, so the model’s weights stay on-chip and tokens flow through layers pipelined across wafers without interruption. Because the weights never leave the silicon, the approach scales with model size — which is the company’s argument that the speed advantage holds as frontier models grow. For inference economics, that is the whole game: the cost and latency of serving a model are dominated by how fast you can feed weights to the compute, and keeping 44 GB resident on one die attacks that directly.

A $10 Billion Partnership Reaches the Flagship

Ultrafast is the most visible product yet of a relationship that has been building for months. OpenAI tapped Cerebras for $10 billion in low-latency compute earlier in 2026, and this launch puts that capacity behind the company’s top model rather than a smaller or specialized one. OpenAI describes Ultrafast as “the next step” in the partnership to bring ultra-low-latency inference to its platform.

For Cerebras, the placement is significant. The startup has long argued its wafer-scale architecture is the right shape for inference even as the market’s center of gravity sits with GPU suppliers — a contest playing out across the accelerator business as incumbents move to bake models directly into their silicon. Landing the serving layer for OpenAI’s flagship gives Cerebras a production reference account at the top of the market. The model itself anchors the GPT-5.6 family OpenAI launched (Sol as the flagship, alongside the balanced Terra and the cost-efficient Luna), so Ultrafast attaches Cerebras to the front of that lineup.

Who Gets It and What It’s For

OpenAI is positioning the tier at time-sensitive, high-stakes work: incident response while an outage is still unfolding, financial research while market conditions are moving, real-time customer support and voice, commerce, and live research loops that used to run overnight. Early access has gone to companies across coding, commerce, and finance, including Jane Street, Podium, Basis, and Rogo.

> “The increase in speed brought by Cerebras is impressive,” said John Crepezzi, AI Assistants at Jane Street, in OpenAI’s announcement. “It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”

OpenAI is keeping the rollout narrow on purpose. The company says it is using the preview period to learn where an order-of-magnitude speed change creates the most value, and will expand access as capacity grows. Both OpenAI and Cerebras are taking sign-ups for updates as the preview widens.

What Happens Next

The near-term observable is capacity, not capability. Ultrafast is a limited preview, and both companies tie any broader availability to capacity growth rather than a fixed date — so the pace of expansion is the thing to watch. The longer question is whether Cerebras’s on-chip-memory advantage holds as OpenAI’s models scale, which is exactly the bet the wafer-scale architecture is built on.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs
AI & Technology

Developing an End-to-End Document Intelligence Pipeline with docTR for OCR, Layout Analysis, KIE, Benchmarking, and Searchable PDFs

August 17, 2026
Why 4K Blu-Ray Always Beats 4K Streaming For Picture Quality
AI & Technology

Why 4K Blu-Ray Always Beats 4K Streaming For Picture Quality

August 17, 2026
One AI module faked 86% of a pipeline’s accuracy gains by feeding another the answers
AI & Technology

One AI module faked 86% of a pipeline’s accuracy gains by feeding another the answers

August 17, 2026
Copilot Autofix Opened a Shell Injection in Snowflake’s CI/CD Pipeline – Unite.AI
AI & Technology

Copilot Autofix Opened a Shell Injection in Snowflake’s CI/CD Pipeline – Unite.AI

August 17, 2026
Next Post
Heat wave threatens millions in Europe

Heat wave threatens millions in Europe

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE: Steve Kornacki analyzes midterm primary election results | Kornacki Cam | NBC News

LIVE: Steve Kornacki analyzes midterm primary election results | Kornacki Cam | NBC News

August 13, 2026
Constellation Software Q2 2026: Soft Organic Growth, But Encouraging Deployment Pace

Constellation Software Q2 2026: Soft Organic Growth, But Encouraging Deployment Pace

August 13, 2026
LIVE NOW: CPI DATA INFLATION REPORT AUGUST 2026

LIVE NOW: CPI DATA INFLATION REPORT AUGUST 2026

August 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!