• bitcoinBitcoin(BTC)$81,182.006.29%
  • ethereumEthereum(ETH)$2,626.317.33%
  • tetherTether(USDT)$1.000.04%
  • binancecoinBNB(BNB)$763.834.17%
  • rippleXRP(XRP)$1.408.38%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$113.3012.04%
  • tronTRON(TRX)$0.3385291.12%
  • zcashZcash(ZEC)$1,498.592.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.29%
  • HyperliquidHyperliquid(HYPE)$91.449.88%
  • dogecoinDogecoin(DOGE)$0.0882108.38%
  • moneroMonero(XMR)$564.5510.16%
  • whitebitWhiteBIT Coin(WBT)$83.305.77%
  • USDSUSDS(USDS)$1.000.03%
  • RainRain(RAIN)$0.0134975.49%
  • chainlinkChainlink(LINK)$12.369.16%
  • cardanoCardano(ADA)$0.22475311.86%
  • leo-tokenLEO Token(LEO)$8.91-0.20%
  • stellarStellar(XLM)$0.1939545.43%
  • uniswapUniswap(UNI)$8.9816.21%
  • bitcoin-cashBitcoin Cash(BCH)$258.4211.09%
  • nearNEAR Protocol(NEAR)$3.7424.00%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$57.146.13%
  • CantonCanton(CC)$0.11108510.97%
  • USD1USD1(USD1)$1.000.05%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.382.76%
  • avalanche-2Avalanche(AVAX)$8.258.72%
  • hedera-hashgraphHedera(HBAR)$0.0794866.27%
  • suiSui(SUI)$0.8211.54%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000068.03%
  • MemeCoreMemeCore(M)$1.3211.51%
  • crypto-com-chainCronos(CRO)$0.0598994.49%
  • BittensorBittensor(TAO)$249.969.12%
  • paypal-usdPayPal USD(PYUSD)$1.000.08%
  • tether-goldTether Gold(XAUT)$4,376.920.77%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.904.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.34%
  • aaveAave(AAVE)$138.497.95%
  • mantleMantle(MNT)$0.629.91%
  • AsterAster(ASTER)$0.761.98%
  • Pump.funPump.fun(PUMP)$0.00438710.51%
  • OndoOndo(ONDO)$0.4000198.36%
  • polkadotPolkadot(DOT)$1.147.09%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta’s SPICE framework lets AI systems teach themselves to reason

November 11, 2025
in AI & Technology
Reading Time: 3 mins read
A A
Meta’s SPICE framework lets AI systems teach themselves to reason
ShareShareShareShareShare

Researchers at Meta FAIR and the National University of Singapore have developed a new reinforcement learning framework for self-improving AI systems.

YOU MAY ALSO LIKE

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI

The New Resident Evil Movie Captures The Survival Horror Magic Of The Games

Called Self-Play In Corpus Environments (SPICE), the framework pits two AI agents against each other, creating its own challenges and gradually improving without human supervision.

While currently a proof-of-concept, this self-play mechanism could provide a basis for future AI systems that can dynamically adapt to their environments, making them more robust against the unpredictability of real-world applications.

The challenge of self-improving AI

The goal of self-improving AI is to create systems that can enhance their capabilities by interacting with their environment.

A common approach is reinforcement learning with verifiable rewards (RLVR), where models are rewarded for providing the correct answers to problems. This is often limited by its reliance on human-curated problem sets and domain-specific reward engineering, which makes it difficult to scale.

Self-play, where a model improves by competing against itself, is another promising paradigm. But existing self-play methods for language models are often limited by two critical factors.

  1. Factual errors in generated questions and answers compound, leading to a feedback loop of hallucinations.

  2. When the problem generator and solver have information symmetry (i.e., share the same knowledge base) they fail to generate genuinely new challenges and fall into repetitive patterns. 

As the researchers note in their paper, “These systematic empirical failures indicate that self-improvement requires interaction with an external source providing diverse, verifiable feedback, rather than closed-loop pure introspection.”

How SPICE works

SPICE is a self-play framework where a single model acts in two distinct roles.

  • A "Challenger" constructs a curriculum of challenging problems from a large corpus of documents.

  • A "Reasoner" then attempts to solve these problems without access to the source documents.

This setup breaks the information symmetry that limits other self-play methods, as the Reasoner does not have access to the documents and knowledge that the Challenger uses to generate the problems.

Grounding the tasks in a vast and diverse corpus of documents prevents hallucination by anchoring questions and answers in real-world content. This is important because for AI systems to reliably self-improve, they need external grounding sources. Therefore, LLM agents should learn from interactions with humans and the real world, not just their own outputs, to avoid compounding errors.

The adversarial dynamic between the two roles creates an automatic curriculum.

The Challenger is rewarded for generating problems that are both diverse and at the frontier of the Reasoner's capability (not too easy and also not impossible).

The Reasoner is rewarded for answering correctly. This symbiotic interaction pushes both agents to continuously discover and overcome new challenges. 

Because the system uses raw documents instead of pre-defined question-answer pairs, it can generate diverse task formats, such as multiple-choice and free-form questions.

This flexibility allows SPICE to be applied to any domain, breaking the bottleneck that has confined previous methods to narrow fields like math and code. It also reduces dependence on expensive human-curated datasets for specialized domains like legal or medical analysis.

SPICE in action

The researchers evaluated SPICE on several base models, including Qwen3-4B-Base and OctoThinker-3B-Hybrid-Base.

They compared its performance against baselines such as the base model with no training, a Reasoner model trained with a fixed "Strong Challenger" (Qwen3-32B-Instruct), and pure self-play methods like R-Zero and Absolute Zero. The evaluation covered a wide range of mathematical and general reasoning benchmarks.

Across all models, SPICE consistently outperformed the baselines, delivering significant improvements in both mathematical and general reasoning tasks.

The results show that the reasoning capabilities developed through corpus-grounded self-play transfer broadly across different models, thanks to the diverse external knowledge corpus they used.

A key finding is that the adversarial dynamic creates an effective automatic curriculum. As training progresses, the Challenger learns to generate increasingly difficult problems.

In one experiment, the Reasoner's pass rate on a fixed set of problems increased from 55% to 85% over time, showing its improved capabilities.

Meanwhile, later versions of the Challenger were able to generate questions that dropped the pass rate of an early-stage Reasoner from 55% to 35%, confirming that both roles co-evolve successfully.

The researchers conclude that this approach presents a paradigm shift in self-improving reasoning methods from “closed-loop self-play that often stagnates due to hallucination drift, to open-ended improvement through interaction with the vast, verifiable knowledge embedded in web document corpora.”

Currently, the corpus used for SPICE represents human experience captured in text. The ultimate goal is for self-improving systems to generate questions based on interactions with reality, including the physical world, the internet, and human interactions across multiple modalities like video, audio, and sensor data.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI
AI & Technology

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI

September 18, 2026
The New Resident Evil Movie Captures The Survival Horror Magic Of The Games
AI & Technology

The New Resident Evil Movie Captures The Survival Horror Magic Of The Games

September 18, 2026
AWS Reworks Bedrock AgentCore Runtime for Elastic Memory, Fast Cold Starts – Unite.AI
AI & Technology

AWS Reworks Bedrock AgentCore Runtime for Elastic Memory, Fast Cold Starts – Unite.AI

September 18, 2026
Still The Best (And It’s Not Close)
AI & Technology

Still The Best (And It’s Not Close)

September 18, 2026
Next Post
LIVE: Trump meets with Chinese President Xi Jinping | NBC News

LIVE: Trump meets with Chinese President Xi Jinping | NBC News

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

September 18, 2026
Mail-in ballot fight heads to Supreme Court for third time

Mail-in ballot fight heads to Supreme Court for third time

September 16, 2026
Newly declassified briefings show clear warnings to presidents before 9/11

Newly declassified briefings show clear warnings to presidents before 9/11

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!