• bitcoinBitcoin(BTC)$78,165.000.35%
  • ethereumEthereum(ETH)$2,447.330.22%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$687.320.24%
  • rippleXRP(XRP)$1.380.96%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.10-0.40%
  • tronTRON(TRX)$0.325964-1.83%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.38%
  • HyperliquidHyperliquid(HYPE)$83.833.80%
  • zcashZcash(ZEC)$855.964.06%
  • dogecoinDogecoin(DOGE)$0.0828420.79%
  • RainRain(RAIN)$0.016617-2.46%
  • USDSUSDS(USDS)$1.000.02%
  • moneroMonero(XMR)$508.750.30%
  • leo-tokenLEO Token(LEO)$9.36-2.87%
  • chainlinkChainlink(LINK)$11.421.92%
  • whitebitWhiteBIT Coin(WBT)$72.000.21%
  • cardanoCardano(ADA)$0.1994712.65%
  • stellarStellar(XLM)$0.1781061.81%
  • bitcoin-cashBitcoin Cash(BCH)$248.221.52%
  • CantonCanton(CC)$0.117028-0.29%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.472.68%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.362.25%
  • uniswapUniswap(UNI)$5.7012.39%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0745931.63%
  • avalanche-2Avalanche(AVAX)$7.291.90%
  • shiba-inuShiba Inu(SHIB)$0.0000053.61%
  • suiSui(SUI)$0.732.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.0562311.09%
  • tether-goldTether Gold(XAUT)$4,360.16-1.19%
  • nearNEAR Protocol(NEAR)$2.029.72%
  • MemeCoreMemeCore(M)$1.06-1.41%
  • okbOKB(OKB)$111.090.08%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.18%
  • BittensorBittensor(TAO)$226.56-0.02%
  • aaveAave(AAVE)$127.003.42%
  • AsterAster(ASTER)$0.700.80%
  • pax-goldPAX Gold(PAXG)$4,369.36-1.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056797-1.26%
  • mantleMantle(MNT)$0.54-5.13%
  • Pump.funPump.fun(PUMP)$0.0044824.93%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Dream First, Learn Later: DECKARD is an AI Approach That Uses LLMs for Training Reinforcement learning (RL) Agents

May 4, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Dream First, Learn Later: DECKARD is an AI Approach That Uses LLMs for Training Reinforcement learning (RL) Agents
ShareShareShareShareShare

Reinforcement learning (RL) is a popular approach to training autonomous agents that can learn to perform complex tasks by interacting with their environment. RL enables them to learn the best action in different conditions and adapt to their environment using a reward system.

A major challenge in RL is how to explore the vast state space of many real-world problems efficiently. This challenge arises due to the fact that in RL, agents learn by interacting with their environment via exploration. Think of an agent that tries to play Minecraft. If you heard about it before, you know how complicated Minecraft crafting tree looks. You have hundreds of craftable objects, and you might need to craft one to craft another, etc. So, it is a really complex environment.

As the environment can have a large number of possible states and actions, it can become difficult for the agent to find the optimal policy through random exploration alone. The agent must balance between exploiting the current best policy and exploring new parts of the state space to find a better policy potentially. Finding efficient exploration methods that can balance exploration and exploitation is an active area of research in RL.

🚀 JOIN the fastest ML Subreddit Community

It’s known that practical decision-making systems need to use prior knowledge about a task efficiently. By having prior information about the task itself, the agent can better adapt its policy and can avoid getting stuck in sub-optimal policies. However, most reinforcement learning methods currently train without any previous training or external knowledge. 

But why is that the case? In recent years, there has been growing interest in using large language models (LLMs) to aid RL agents in exploration by providing external knowledge. This approach has shown promise, but there are still many challenges to overcome, such as grounding the LLM knowledge in the environment and dealing with the accuracy of LLM outputs.

So, should we give up on using LLMs to aid RL agents? If not, how can we fix those problems and then use them again to guide RL agents? The answer has a name, and it’s DECKARD.

DECKARD is trained for Minecraft, as crafting a specific item in Minecraft can be a challenging task if one lacks expert knowledge of the game. This has been demonstrated by studies that have shown that achieving a goal in Minecraft can be made easier through the use of dense rewards or expert demonstrations. As a result, item crafting in Minecraft has become a persistent challenge in the field of AI.

DECKARD utilizes a few-shot prompting technique on a large language model (LLM) to generate an Abstract World Model (AWM) for subgoals. It uses the LLM to hypothesize an AWM, which means it dreams about the task and the steps to solve it. Then, it wakes up and learns a modular policy of subgoals that it generates during dreaming. Since this is done in the real environment, DECKARD can verify the hypothesized AWM. The AWM is corrected during the waking phase, and discovered nodes are marked as verified to be used again in the future.

Experiments show us that LLM guidance is essential to exploration in DECKARD, with a version of the agent without LLM guidance taking over twice as long to craft most items during open-ended exploration. When exploring a specific task, DECKARD improves sample efficiency by orders of magnitude compared to comparable agents, demonstrating the potential for robustly applying LLMs to RL.


Check out the Research Paper, Code, and Project. Don’t forget to join our 20k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

The Pros And Cons Of Using A TV As Your Computer Monitor

Enovis Makes Binding Offer to Acquire eCential Robotics – Unite.AI

Ekrem Çetinkaya received his B.Sc. in 2018 and M.Sc. in 2019 from Ozyegin University, Istanbul, Türkiye. He wrote his M.Sc. thesis about image denoising using deep convolutional networks. He is currently pursuing a Ph.D. degree at the University of Klagenfurt, Austria, and working as a researcher on the ATHENA project. His research interests include deep learning, computer vision, and multimedia networking.


Credit: Source link

ShareTweetSendSharePin

Related Posts

The Pros And Cons Of Using A TV As Your Computer Monitor
AI & Technology

The Pros And Cons Of Using A TV As Your Computer Monitor

September 1, 2026
Enovis Makes Binding Offer to Acquire eCential Robotics – Unite.AI
AI & Technology

Enovis Makes Binding Offer to Acquire eCential Robotics – Unite.AI

September 1, 2026
Mech-Mind Robotics Lists on Hong Kong Exchange in Embodied AI IPO – Unite.AI
AI & Technology

Mech-Mind Robotics Lists on Hong Kong Exchange in Embodied AI IPO – Unite.AI

September 1, 2026
Dyson’s CameraJet Is A 0 Three-In-One Toothbrush That Scans Your Maw
AI & Technology

Dyson’s CameraJet Is A $500 Three-In-One Toothbrush That Scans Your Maw

September 1, 2026
Next Post
How to Avoid Investing in Companies that go Bust

How to Avoid Investing in Companies that go Bust

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
3 killed in shooting at North Carolina home

3 killed in shooting at North Carolina home

August 28, 2026
Hugging Face Always Attracts Buyers, Co-Founder Says

Hugging Face Always Attracts Buyers, Co-Founder Says

August 29, 2026
‘Major casualties’ feared after flash floods along Nepal-Tibet border – The Guardian

‘Major casualties’ feared after flash floods along Nepal-Tibet border – The Guardian

August 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!