• bitcoinBitcoin(BTC)$78,191.00-1.22%
  • ethereumEthereum(ETH)$2,472.87-1.08%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$718.31-4.92%
  • rippleXRP(XRP)$1.38-3.69%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.61-2.88%
  • tronTRON(TRX)$0.3399080.47%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.84%
  • zcashZcash(ZEC)$1,216.40-0.60%
  • HyperliquidHyperliquid(HYPE)$83.46-3.46%
  • dogecoinDogecoin(DOGE)$0.085532-5.50%
  • RainRain(RAIN)$0.0163301.70%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$511.792.31%
  • whitebitWhiteBIT Coin(WBT)$80.75-1.45%
  • chainlinkChainlink(LINK)$11.80-6.20%
  • leo-tokenLEO Token(LEO)$9.190.09%
  • cardanoCardano(ADA)$0.213021-3.47%
  • stellarStellar(XLM)$0.180167-5.42%
  • bitcoin-cashBitcoin Cash(BCH)$249.26-3.99%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.104011-4.10%
  • litecoinLitecoin(LTC)$52.64-3.28%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-1.96%
  • uniswapUniswap(UNI)$6.02-13.02%
  • avalanche-2Avalanche(AVAX)$7.82-2.57%
  • hedera-hashgraphHedera(HBAR)$0.076773-3.29%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.463.18%
  • suiSui(SUI)$0.77-6.63%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.44%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057826-3.35%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.233.32%
  • tether-goldTether Gold(XAUT)$4,409.220.30%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$252.68-2.79%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.17-1.16%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.14%
  • mantleMantle(MNT)$0.60-5.66%
  • AsterAster(ASTER)$0.72-4.91%
  • aaveAave(AAVE)$124.40-4.15%
  • pax-goldPAX Gold(PAXG)$4,413.640.32%
  • polkadotPolkadot(DOT)$1.10-7.48%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0567111.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

LLMs Outperform Reinforcement Learning- Meet SPRING: An Innovative Prompting Framework for LLMs Designed to Enable in-Context Chain-of-Thought Planning and Reasoning

May 28, 2023
in AI & Technology
Reading Time: 7 mins read
A A
LLMs Outperform Reinforcement Learning- Meet SPRING: An Innovative Prompting Framework for LLMs Designed to Enable in-Context Chain-of-Thought Planning and Reasoning
ShareShareShareShareShare

SPRING is an LLM-based policy that outperforms Reinforcement Learning algorithms in an interactive environment requiring multi-task planning and reasoning. 

A group of researchers from Carnegie Mellon University, NVIDIA, Ariel University, and Microsoft have investigated the use of Large Language Models (LLMs) for understanding and reasoning with human knowledge in the context of games. They propose a two-stage approach called SPRING, which involves studying an academic paper and then using a Question-Answer (QA) framework to justify the knowledge obtained.

More details about SPRING

In the first stage, the authors read the LaTeX source code of the original paper by Hafner (2021) to extract prior knowledge. They employed an LLM to extract relevant information, including game mechanics and desirable behaviors documented in the paper. They then utilized a QA summarization framework similar to Wu et al. (2023) to generate QA dialogue based on the extracted knowledge, enabling SPRING to handle diverse contextual information.

🚀 JOIN the fastest ML Subreddit Community

The second stage focused on in-context chain-of-thought reasoning using LLMs to solve complex games. They constructed a directed acyclic graph (DAG) as a reasoning module, where questions are nodes and dependencies between questions are represented as edges. For example, the question “For each action, are the requirements met?” is linked to the question “What are the top 5 actions?” within the DAG, establishing a dependency from the latter question to the former.

LLM answers are computed for each node/question by traversing the DAG in topological order. The final node in the DAG represents the question about the best action to take, and the LLM’s answer is directly translated into an environmental action.

Experiments and Results

The Crafter Environment, introduced by Hafner (2021), is an open-world survival game with 22 achievements organized in a tech tree of depth 7. The game is represented as a grid world with top-down observations and a discrete action space consisting of 17 options. The observations also provide information about the player’s current inventory state, including health points, food, water, rest levels, and inventory items.

The authors compared SPRING and popular RL methods on the Crafter benchmark. Subsequently, experiments and analysis were carried out on different components of their architecture to examine the impact of each part on the in-context “reasoning” abilities of the LLM.

Source: https://arxiv.org/pdf/2305.15486.pdf

The authors compared the performance of various RL baselines to SPRING with GPT-4, conditioned on the environment paper by Hafner (2021). SPRING surpasses previous state-of-the-art (SOTA) methods by a significant margin, achieving an 88% relative improvement in-game score and a 5% improvement in reward compared to the best-performing RL method by Hafner et al. (2023).

Notably, SPRING leverages prior knowledge from reading the paper and requires zero training steps, while RL methods typically necessitate millions of training steps.

Source: https://arxiv.org/pdf/2305.15486.pdf

The above figure represents a plot of unlock rates for different tasks, comparing SPRING to popular RL baselines. SPRING, empowered by prior knowledge, outperforms RL methods by more than ten times on achievements such as “Make Stone Pickaxe,” “Make Stone Sword,” and “Collect Iron,” which are deeper in the tech tree (up to depth 5) and challenging to reach through random exploration. 

Moreover, SPRING performs perfectly on achievements like “Eat Cow” and “Collect Drink.” At the same time, model-based RL frameworks like Dreamer-V3 have significantly lower unlock rates (over five times lower) for “Eat Cow” due to the challenge of reaching moving cows through random exploration. Importantly, SPRING does not take action “Place Stone” since it was not discussed as beneficial for the agent in the paper by Hafner (2021), even though it could be easily achieved through random exploration.

Limitations

One limitation of using an LLM for interacting with the environment is the need for object recognition and grounding. However, this limitation doesn’t exist in environments that provide accurate object information, such as contemporary games and virtual reality worlds. While pre-trained visual backbones struggle with games, they perform reasonably well in real-world-like environments. Recent advancements in visual-language models indicate potential for reliable solutions in visual-language understanding in the future.

Conclusion

In summary, the SPRING framework showcases the potential of Language Models (LLMs) for game understanding and reasoning. By leveraging prior knowledge from academic papers and employing in-context chain-of-thought reasoning, SPRING outperforms previous state-of-the-art methods on the Crafter benchmark, achieving substantial improvements in-game score and reward. The results highlight the power of LLMs in complex game tasks and suggest future advancements in visual-language models could address existing limitations, paving the way for reliable and generalizable solutions.


Check out the Paper. Don’t forget to join our 22k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI
AI & Technology

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

September 10, 2026
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Next Post
High Yield Bond Run Far From Done Says Fund Manager

High Yield Bond Run Far From Done Says Fund Manager

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
‘no steering wheel, no pedals’

‘no steering wheel, no pedals’

September 3, 2026
My Wife Is Blowing All Of Our Money On Parties

My Wife Is Blowing All Of Our Money On Parties

September 9, 2026
Playdate Season 3 Kicks Off On October 8

Playdate Season 3 Kicks Off On October 8

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!