• bitcoinBitcoin(BTC)$78,447.00-0.59%
  • ethereumEthereum(ETH)$2,482.240.19%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$749.251.33%
  • rippleXRP(XRP)$1.412.01%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.44-0.38%
  • tronTRON(TRX)$0.3390091.47%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,174.511.54%
  • HyperliquidHyperliquid(HYPE)$83.57-1.95%
  • dogecoinDogecoin(DOGE)$0.0900650.68%
  • RainRain(RAIN)$0.0166461.95%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.61-2.33%
  • whitebitWhiteBIT Coin(WBT)$79.789.61%
  • moneroMonero(XMR)$499.26-6.52%
  • cardanoCardano(ADA)$0.2295745.13%
  • leo-tokenLEO Token(LEO)$9.200.54%
  • stellarStellar(XLM)$0.1904800.41%
  • bitcoin-cashBitcoin Cash(BCH)$256.93-1.22%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$55.02-1.24%
  • uniswapUniswap(UNI)$6.84-0.25%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.105447-0.62%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-0.52%
  • hedera-hashgraphHedera(HBAR)$0.080369-1.28%
  • avalanche-2Avalanche(AVAX)$8.050.07%
  • suiSui(SUI)$0.821.36%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • shiba-inuShiba Inu(SHIB)$0.0000051.10%
  • nearNEAR Protocol(NEAR)$2.423.79%
  • crypto-com-chainCronos(CRO)$0.0602425.62%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,396.96-0.29%
  • MemeCoreMemeCore(M)$1.185.58%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$259.631.21%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.12-0.51%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.18%
  • mantleMantle(MNT)$0.62-1.35%
  • AsterAster(ASTER)$0.76-1.27%
  • aaveAave(AAVE)$129.76-1.16%
  • polkadotPolkadot(DOT)$1.188.88%
  • pax-goldPAX Gold(PAXG)$4,400.02-0.27%
  • OndoOndo(ONDO)$0.380289-1.54%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

On-device AI agents hit a hard memory limit. Apple’s new architecture routes around it.

June 9, 2026
in AI & Technology
Reading Time: 5 mins read
A A
On-device AI agents hit a hard memory limit. Apple’s new architecture routes around it.
ShareShareShareShareShare

On-device AI models have stayed small because the entire weight set has to live in DRAM, capping practical parameter counts well below what server-side deployments use. Enterprise architects evaluating agentic workloads have had to choose between capable cloud-dependent models and limited on-device ones. Apple’s third-generation foundation models, announced at WWDC26, break that constraint by moving the weight set off DRAM entirely.

The AFM 3 family was developed in collaboration with Google and spans five models: two on-device and three server-based, all running within Apple’s Private Cloud Compute boundary. The server-side models, including AFM 3 Cloud Pro for agentic tool use and complex reasoning, run on Nvidia GPUs in Google Cloud. The on-device architecture is Apple’s own. AFM 3 Core Advanced is a 20-billion-parameter model that stores weights in NAND flash rather than DRAM.

YOU MAY ALSO LIKE

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

What Is Considered Good Speed For Home Internet And How Can You Test It?

“Instead of forcing the entire model into DRAM, the full model is stored in flash memory,” Apple’s research team wrote. “Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt.”

How the architecture actually works

The memory wall Apple is working around is one every local AI developer runs into.

“You can’t put 20B parameters in RAM at any reasonable precision,” Awni Hannun, a researcher at Anthropic and former Apple research scientist, posted on X. “To make it work they are using pretty exotic architecture by today’s standards. A small model predicts from the query (or prompt) which experts to load from NAND into RAM.”

That prediction-and-load mechanism has three distinct components, each driven by the hardware constraints of consumer silicon.

The full 20B weight set lives in flash, not DRAM. AFM 3 Core Advanced stores its entire parameter set in NAND flash rather than active memory. Standard on-device deployments require the full model to fit in DRAM, which is what caps their parameter counts. Apple’s approach, which it calls Instruction-Following Pruning (IFP) and developed with its own researchers, treats flash as the model’s permanent home and DRAM as a working buffer for whichever experts a given prompt requires.

Expert routing happens once per prompt, not per token. In a conventional Mixture of Experts model, a router selects different experts for every token generated — which would require continuous weight movement between flash and DRAM at inference speed. NAND-to-DRAM bandwidth cannot support that. AFM 3 Core Advanced routes once at prompt time, selects a fixed expert set, loads it into DRAM alongside always-active shared experts, and generates all tokens from that same configuration.

“The key distinction from a typical MoE is that you do this once per query and then generate all the tokens with the same experts,” Hannun wrote.

Source: Apple Machine Learning Research, June 8, 2026.

Active parameter count scales from 1B to 4B depending on task complexity. Rather than running a fixed model size for every request, AFM 3 Core Advanced adjusts how many parameters it activates based on what the task requires — 1 billion for simpler operations, up to 4 billion for harder ones, all drawn from the 20-billion-parameter pool in flash.

What Apple has and hasn’t disclosed

The architecture paper is detailed on the memory design and sparse activation mechanism. It is less forthcoming on practical deployment constraints.

Apple’s profiling tools expose timing but not the metrics that decide production viability. “Energy, memory bandwidth, thermal? Not in the docs,” Marco Abis, who is building Ziraph, a profiler for local AI on Apple silicon, posted on X. “A notable gap, given those decide most of on-device performance.” 

Abis also did not find a statement in Apple’s documentation — across the Core AI docs, the Foundation Models docs or the Private Cloud Compute security post — of when an on-device request transparently offloads, or whether that routing is visible to the developer or the user. For enterprises that need to document where inference runs, that is a direct compliance problem.

Not all the information is currently available. Apple has indicated a full technical report with benchmarks is coming later this summer.

What this means for enterprise architects

Regulated industries evaluating agentic AI deployments now have a concrete architectural decision to make.

  • The DRAM wall for on-device agents just moved. Enterprises evaluating agents that need to run without a cloud round-trip now have a 20-billion-parameter local option to evaluate. The constraint shifts from model capability to device hardware.

  • The private/cloud boundary is now an architectural decision, not a default. Simpler requests stay on-device; complex agentic tasks route to AFM 3 Cloud Pro on Private Cloud Compute. Apple has not publicly specified when a request offloads or whether that routing is visible to the developer — a gap that complicates policy decisions for organizations that need to document where inference runs.

  • The agentic server tier depends on Google Cloud. AFM 3 Cloud Pro runs on Nvidia GPUs in Google Cloud. The Private Cloud Compute guarantee covers data privacy. It does not eliminate the Google Cloud dependency for server-side inference.

AFM 3 Core Advanced gives enterprises a 20-billion-parameter on-device option that did not exist before WWDC26. Whether it is deployable at scale depends on answers Apple has not yet published. Those details are due in the summer technical report.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI
AI & Technology

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

September 8, 2026
What Is Considered Good Speed For Home Internet And How Can You Test It?
AI & Technology

What Is Considered Good Speed For Home Internet And How Can You Test It?

September 8, 2026
What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?
AI & Technology

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

September 8, 2026
What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI
AI & Technology

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

September 8, 2026
Next Post
Fresh blow for California’s Carl’s Jr as iconic burger chain to close more stores

Fresh blow for California’s Carl’s Jr as iconic burger chain to close more stores

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
American doctor performs 1,000 eye surgeries at South Sudan border

American doctor performs 1,000 eye surgeries at South Sudan border

September 4, 2026
HPE CEO Neri on Oracle Deal, AI Adoption and Earnings Outlook

HPE CEO Neri on Oracle Deal, AI Adoption and Earnings Outlook

September 3, 2026
Trump asks Supreme Court to let executive order on mail-in voting proceed

Trump asks Supreme Court to let executive order on mail-in voting proceed

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!