• bitcoinBitcoin(BTC)$82,649.00-2.43%
  • ethereumEthereum(ETH)$2,641.80-2.57%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$759.41-2.40%
  • rippleXRP(XRP)$1.48-3.43%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$117.91-4.99%
  • tronTRON(TRX)$0.3340300.04%
  • zcashZcash(ZEC)$1,544.34-7.00%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$89.01-4.34%
  • dogecoinDogecoin(DOGE)$0.092518-5.15%
  • chainlinkChainlink(LINK)$13.59-5.13%
  • moneroMonero(XMR)$528.16-5.41%
  • whitebitWhiteBIT Coin(WBT)$82.50-2.44%
  • USDSUSDS(USDS)$1.00-0.03%
  • cardanoCardano(ADA)$0.243860-4.76%
  • RainRain(RAIN)$0.012535-1.33%
  • leo-tokenLEO Token(LEO)$9.06-0.02%
  • stellarStellar(XLM)$0.211260-2.89%
  • nearNEAR Protocol(NEAR)$5.10-2.91%
  • bitcoin-cashBitcoin Cash(BCH)$307.73-9.53%
  • CantonCanton(CC)$0.1434494.94%
  • uniswapUniswap(UNI)$8.88-12.12%
  • litecoinLitecoin(LTC)$69.71-3.14%
  • hedera-hashgraphHedera(HBAR)$0.11303919.11%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.39-6.30%
  • suiSui(SUI)$1.17-6.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.652.89%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.02%
  • BitwayBitway(BTW)$1.3625.11%
  • BittensorBittensor(TAO)$301.12-9.40%
  • shiba-inuShiba Inu(SHIB)$0.000006-5.83%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.063676-6.83%
  • quant-networkQuant(QNT)$201.9314.45%
  • tether-goldTether Gold(XAUT)$4,151.05-3.01%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • MemeCoreMemeCore(M)$1.16-4.51%
  • EthenaEthena(ENA)$0.258401-4.37%
  • OndoOndo(ONDO)$0.52-5.35%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$116.72-4.10%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.11%
  • aaveAave(AAVE)$147.05-5.82%
  • Pump.funPump.fun(PUMP)$0.0048038.57%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MIT Researchers Enhanced Artificial Intelligence (AI) 64x Better at Planning, Achieving 94% Accuracy

September 22, 2025
in AI & Technology
Reading Time: 6 mins read
A A
MIT Researchers Enhanced Artificial Intelligence (AI) 64x Better at Planning, Achieving 94% Accuracy
ShareShareShareShareShare

Can a 8B-parameter language model produce provably valid multi-step plans instead of plausible guesses? MIT CSAIL researchers introduce PDDL-INSTRUCT, an instruction-tuning framework that couples logical chain-of-thought with external plan validation (VAL) to lift symbolic planning performance of LLMs. On PlanBench, a tuned Llama-3-8B reaches 94% valid plans on Blocksworld, with large jumps on Mystery Blocksworld and Logistics; across domains they report up to a 66% absolute improvement over baselines.

https://arxiv.org/pdf/2509.13351

But What’s new?

The research team tackles a well-known failure mode: LLMs often generate “plausible-sounding” but logically invalid multi-step plans. PDDL-INSTRUCT couples explicit state/action semantics with ground-truth checking:

YOU MAY ALSO LIKE

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

You Can Now Preorder The Tiny Boox Picco Ereader

  • Error education: Models are trained to explain why candidate plans fail (unsatisfied preconditions, wrong effects, frame violations, or goal not reached).
  • Logical chain-of-thought (CoT): Prompts require step-by-step inference over preconditions and add/del effects, yielding state→action→state traces ⟨sᵢ, aᵢ₊₁, sᵢ₊₁⟩.
  • External verification (VAL): Every step is validated with the classic VAL plan validator; feedback can be binary (valid/invalid) or detailed (which precondition/effect failed). Detailed feedback yielded the strongest gains.
  • Two-stage optimization:
    • Stage-1 optimizes the reasoning chains (penalizing state-transition errors);
    • Stage-2 optimizes end-task planning accuracy.
for similar explainer infographics for other articles please subscribe to our newsletter

How Good is it? Benchmarks

Evaluation follows PlanBench—Blocksworld, Mystery Blocksworld (predicate names obfuscated to break pattern-matching), and Logistics—established stress tests where generic LLMs historically underperform on plan generation. The authors highlight that Mystery Blocksworld is particularly challenging; prior studies often report <5% validity without tool support.

  • Blocksworld: up to 94% valid plans with Llama-3-8B under PDDL-INSTRUCT.
  • Mystery Blocksworld: large relative gains; the paper reports dramatic improvement versus a near-zero baseline (reported as orders-of-magnitude, e.g., 64× in their summary figures/tables).
  • Logistics: substantial increases in valid plans.

Across domains, the research team showcase up to 66% absolute improvement over untuned baselines. Detailed validator feedback outperforms binary signals, and longer feedback budgets further help.

https://arxiv.org/pdf/2509.13351

Summary

PDDL-INSTRUCT shows that coupling logical chain-of-thought with external plan validation can materially improve LLM planning, but its current scope is classical PDDL domains (Blocksworld, Mystery Blocksworld, Logistics) and relies on VAL as an external oracle; the reported gains—e.g., 94% valid plans on Blocksworld and large relative improvements on Mystery Blocksworld with Llama-3-8B—demonstrate a viable path for neuro-symbolic training where reasoning steps are grounded in formal semantics and checked automatically, suggesting immediate utility for agent pipelines that can tolerate a verifier in the loop while longer-horizon, temporal/numeric, and cost-sensitive planning remain open extensions.


Check out the PAPER. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

The post MIT Researchers Enhanced Artificial Intelligence (AI) 64x Better at Planning, Achieving 94% Accuracy appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
AI & Technology

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

September 28, 2026
Next Post
Congress returns to tackle issues from government funding to the Epstein case

Congress returns to tackle issues from government funding to the Epstein case

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
SpaceX: Louisiana Is Huge, But The Revenue Math Still Doesn’t Work (NASDAQ:SPCX)

SpaceX: Louisiana Is Huge, But The Revenue Math Still Doesn’t Work (NASDAQ:SPCX)

September 22, 2026
Meta social media settlement marks ‘era of holding big tech accountable’: NJ Attorney General

Meta social media settlement marks ‘era of holding big tech accountable’: NJ Attorney General

September 23, 2026
Morning News NOW Full Episode – Aug. 25

Morning News NOW Full Episode – Aug. 25

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!