• bitcoinBitcoin(BTC)$84,423.000.45%
  • ethereumEthereum(ETH)$2,686.33-0.07%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$775.530.34%
  • rippleXRP(XRP)$1.52-1.12%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$121.860.56%
  • tronTRON(TRX)$0.333803-0.70%
  • zcashZcash(ZEC)$1,581.031.29%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.08%
  • HyperliquidHyperliquid(HYPE)$91.47-0.81%
  • dogecoinDogecoin(DOGE)$0.096681-1.22%
  • chainlinkChainlink(LINK)$14.10-1.28%
  • moneroMonero(XMR)$545.41-2.02%
  • whitebitWhiteBIT Coin(WBT)$84.260.40%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.253711-1.08%
  • RainRain(RAIN)$0.012567-5.65%
  • leo-tokenLEO Token(LEO)$9.061.12%
  • stellarStellar(XLM)$0.215427-1.45%
  • nearNEAR Protocol(NEAR)$5.216.84%
  • bitcoin-cashBitcoin Cash(BCH)$334.54-0.97%
  • uniswapUniswap(UNI)$9.700.98%
  • litecoinLitecoin(LTC)$71.07-1.81%
  • CantonCanton(CC)$0.134934-1.96%
  • suiSui(SUI)$1.244.70%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • avalanche-2Avalanche(AVAX)$10.930.33%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.689.44%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.093521-0.87%
  • BittensorBittensor(TAO)$323.44-2.44%
  • shiba-inuShiba Inu(SHIB)$0.000006-1.75%
  • crypto-com-chainCronos(CRO)$0.0670491.86%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BitwayBitway(BTW)$1.1911.90%
  • EthenaEthena(ENA)$0.2784042.65%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • quant-networkQuant(QNT)$185.7852.46%
  • MemeCoreMemeCore(M)$1.18-3.15%
  • tether-goldTether Gold(XAUT)$4,279.680.02%
  • OndoOndo(ONDO)$0.551.86%
  • okbOKB(OKB)$121.03-0.26%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.64-0.50%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.07%
  • Pump.funPump.fun(PUMP)$0.0048629.23%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Agentic Context Engineering (ACE): Self-Improving LLMs via Evolving Contexts, Not Fine-Tuning

October 10, 2025
in AI & Technology
Reading Time: 7 mins read
A A
Agentic Context Engineering (ACE): Self-Improving LLMs via Evolving Contexts, Not Fine-Tuning
ShareShareShareShareShare

TL;DR: A team of researchers from Stanford University, SambaNova Systems and UC Berkeley introduce ACE framework that improves LLM performance by editing and growing the input context instead of updating model weights. Context is treated as a living “playbook” maintained by three roles—Generator, Reflector, Curator—with small delta items merged incrementally to avoid brevity bias and context collapse. Reported gains: +10.6% on AppWorld agent tasks, +8.6% on finance reasoning, and ~86.9% average latency reduction vs strong context-adaptation baselines. On the AppWorld leaderboard snapshot (Sept 20, 2025), ReAct+ACE (59.4%) ≈ IBM CUGA (60.3%, GPT-4.1) while using DeepSeek-V3.1.

https://arxiv.org/pdf/2510.04618

What ACE changes?

ACE positions “context engineering” as a first-class alternative to parameter updates. Instead of compressing instructions into short prompts, ACE accumulates and organizes domain-specific tactics over time, arguing that higher context density improves agentic tasks where tools, multi-turn state, and failure modes matter.

YOU MAY ALSO LIKE

How To Improve Your Router’s Security In 10 Minutes

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

Method: Generator → Reflector → Curator

  • Generator executes tasks and produces trajectories (reasoning/tool calls), exposing helpful vs harmful moves.
  • Reflector distills concrete lessons from those traces.
  • Curator converts lessons into typed delta items (with helpful/harmful counters) and merges them deterministically, with de-duplication and pruning to keep the playbook targeted.

Two design choices—incremental delta updates and grow-and-refine—preserve useful history and prevent “context collapse” from monolithic rewrites. To isolate context effects, the research team fixes the same base LLM (non-thinking DeepSeek-V3.1) across all three roles.

Benchmarks

AppWorld (agents): Built on the official ReAct baseline, ReAct+ACE outperforms strong baselines (ICL, GEPA, Dynamic Cheatsheet), with +10.6% average over selected baselines and ~+7.6% over Dynamic Cheatsheet in online adaptation. On the Sept 20, 2025 leaderboard, ReAct+ACE 59.4% vs IBM CUGA 60.3% (GPT-4.1); ACE surpasses CUGA on the harder test-challenge split, while using a smaller open-source base model.

Finance (XBRL): On FiNER token tagging and XBRL Formula numerical reasoning, ACE reports +8.6% average over baselines with ground-truth labels for offline adaptation; it also works with execution-only feedback, though quality of signals matters.

https://arxiv.org/pdf/2510.04618
https://arxiv.org/pdf/2510.04618

Cost and latency

ACE’s non-LLM merges plus localized updates reduce adaptation overhead substantially:

  • Offline (AppWorld): −82.3% latency and −75.1% rollouts vs GEPA.
  • Online (FiNER): −91.5% latency and −83.6% token cost vs Dynamic Cheatsheet.
https://arxiv.org/pdf/2510.04618

Key Takeaways

  • ACE = context-first adaptation: Improves LLMs by incrementally editing an evolving “playbook” (delta items) curated by Generator→Reflector→Curator, using the same base LLM (non-thinking DeepSeek-V3.1) to isolate context effects and avoid collapse from monolithic rewrites.
  • Measured gains: ReAct+ACE reports +10.6% over strong baselines on AppWorld and achieves 59.4% vs IBM CUGA 60.3% (GPT-4.1) on the Sept 20, 2025 leaderboard snapshot; finance benchmarks (FiNER + XBRL Formula) show +8.6% average over baselines.
  • Lower overhead than reflective-rewrite baselines: ACE reduces adaptation latency by ~82–92% and rollouts/token cost by ~75–84%, contrasting with Dynamic Cheatsheet’s persistent memory and GEPA’s Pareto prompt evolution approaches.

Conclusion

ACE positions context engineering as a first-class alternative to weight updates: maintain a persistent, curated playbook that accumulates task-specific tactics, yielding measurable gains on AppWorld and finance reasoning while cutting adaptation latency and token rollouts versus reflective-rewrite baselines. The approach is practical—deterministic merges, delta items, and long-context–aware serving—and its limits are clear: outcomes track feedback quality and task complexity. If adopted, agent stacks may “self-tune” primarily through evolving context rather than new checkpoints.


Check out the PAPER here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Agentic Context Engineering (ACE): Self-Improving LLMs via Evolving Contexts, Not Fine-Tuning appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Improve Your Router’s Security In 10 Minutes
AI & Technology

How To Improve Your Router’s Security In 10 Minutes

September 27, 2026
Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)
AI & Technology

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

September 27, 2026
AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
AI & Technology

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

September 27, 2026
Next Post
A four-pack of AirTags is cheaper than ever right now

A four-pack of AirTags is cheaper than ever right now

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Meet the Press NOW — August 27

Meet the Press NOW — August 27

September 22, 2026
Israel launches deadly airstrike on Gaza

Israel launches deadly airstrike on Gaza

September 25, 2026
How Trump turned a refugee bureau into a 0 million deportation operation – The Washington Post

How Trump turned a refugee bureau into a $410 million deportation operation – The Washington Post

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!