• bitcoinBitcoin(BTC)$84,824.000.91%
  • ethereumEthereum(ETH)$2,708.780.75%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$780.730.84%
  • rippleXRP(XRP)$1.54-0.39%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$123.552.01%
  • tronTRON(TRX)$0.333986-0.77%
  • zcashZcash(ZEC)$1,659.167.51%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.68%
  • HyperliquidHyperliquid(HYPE)$92.900.86%
  • dogecoinDogecoin(DOGE)$0.0984411.04%
  • chainlinkChainlink(LINK)$14.22-0.07%
  • moneroMonero(XMR)$552.180.30%
  • whitebitWhiteBIT Coin(WBT)$84.670.93%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2568880.37%
  • RainRain(RAIN)$0.0126882.33%
  • leo-tokenLEO Token(LEO)$9.071.11%
  • stellarStellar(XLM)$0.216861-0.75%
  • bitcoin-cashBitcoin Cash(BCH)$338.040.93%
  • nearNEAR Protocol(NEAR)$5.196.93%
  • uniswapUniswap(UNI)$9.832.63%
  • litecoinLitecoin(LTC)$71.56-1.82%
  • CantonCanton(CC)$0.136364-0.56%
  • suiSui(SUI)$1.267.48%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • avalanche-2Avalanche(AVAX)$10.931.06%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.619.39%
  • USD1USD1(USD1)$1.000.02%
  • hedera-hashgraphHedera(HBAR)$0.0948701.22%
  • BittensorBittensor(TAO)$332.640.53%
  • shiba-inuShiba Inu(SHIB)$0.0000061.19%
  • crypto-com-chainCronos(CRO)$0.0683244.80%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BitwayBitway(BTW)$1.1424.07%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.273279-2.50%
  • MemeCoreMemeCore(M)$1.20-0.61%
  • tether-goldTether Gold(XAUT)$4,280.340.01%
  • OndoOndo(ONDO)$0.550.03%
  • okbOKB(OKB)$121.910.62%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.410.66%
  • quant-networkQuant(QNT)$160.1454.60%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
  • mantleMantle(MNT)$0.69-2.60%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Imperial College London and DeepMind Designed an AI Framework that Uses Language as the Core Reasoning Tool of an RL Agent

July 28, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from Imperial College London and DeepMind Designed an AI Framework that Uses Language as the Core Reasoning Tool of an RL Agent
ShareShareShareShareShare

In recent years, there have been significant breakthroughs in the field of Deep Learning, particularly in the popular sub-fields of Artificial Intelligence, including Natural Language Processing (NLP), Natural Language Understanding (NLU) and Computer Vision (CV). Large Language Models (LLMs) have been created in the framework of NLP and demonstrate outstanding language processing and text production skills that are on par with human talents. On the other hand, without any explicit guidance, CV’s Vision Transformers (ViTs) have been able to learn meaningful representations from photos and videos. Vision-linguistic Models (VLMs) have also been developed, which can connect visual inputs with linguistic descriptions or the other way around.

Foundation Models behind a wide range of downstream applications involving various input modalities have been pre-trained on vast amounts of textual and visual data, leading to the emergence of significant attributes like common sense reasoning, proposing and sequencing sub-goals, and visual understanding. The prospect of utilizing Foundation Models’ capabilities to create more effective and all-encompassing reinforcement learning (RL) agents is the topic of research for researchers. RL agents often pick up knowledge through interacting with their surroundings and getting rewards as feedback, but this method of learning by trial and error can be time-consuming and unworkable.

To address the limitations, a team of researchers has proposed a framework that places language at the core of reinforcement learning robotic agents, particularly in scenarios where learning from scratch is required. The core contribution of their work is to demonstrate that by utilizing LLMs and VLMs, they can effectively address several fundamental problems in particularly four RL settings.

  1. Efficient Exploration in Sparse-Reward Settings: It is difficult for RL agents to learn the best behavior because they frequently find it difficult to explore settings with few rewards. The suggested approach makes exploration and learning in these contexts more effective by utilizing the knowledge kept in Foundation Models.
  1. Reusing gathered Data for Sequential Learning: The framework allows RL agents to build on previously gathered data rather than beginning from scratch each time a new task is met, aiding the sequential learning of new tasks.
  1. Scheduling learned abilities for NewTasks: The framework supports the scheduling of learned abilities, enabling agents to handle novel tasks with their current knowledge efficiently.
  1. Learning from Observations of Expert Agents: By using Foundation Models to learn from observations of expert agents, learning processes can become more efficient and quick.

The team has summarized the main contributions as follows –

  1. The framework has been made in a way that enables the RL agent to reason and make judgments more effectively based on textual information by using language models and vision language models as the fundamental reasoning tools. The agent’s capacity to comprehend challenging tasks and settings is improved by this method.
  1. The proposed framework shows its efficiency in resolving fundamental RL problems that in the past needed distinct, specially created algorithms.
  1. The new framework outperforms conventional baseline techniques in the sparse-reward robotic manipulation setting.
  2. The framework also shows that it can efficiently use previously taught skills to complete tasks. The RL agent’s generalization and adaptability are enhanced by the ability to transfer learned information to new situations.
  1. It demonstrates how the RL agent may accurately learn from observable demonstrations by imitating films of human experts.

In conclusion, the study shows that language models and vision language models have the ability to serve as the core components of reinforcement learning agents’ reasoning.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🔥 Gain a competitive
edge with data: Actionable market intelligence for global brands, retailers, analysts, and investors. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
AI & Technology

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

September 27, 2026
Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?
AI & Technology

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

September 27, 2026
Next Post
Facebook Is Working to Combat Online Sex Trafficking, Says Zuckerberg

Facebook Is Working to Combat Online Sex Trafficking, Says Zuckerberg

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Judge warns prosecution not to bring up Clancy’s Catholic faith

Judge warns prosecution not to bring up Clancy’s Catholic faith

September 24, 2026
Hayden Panettiere died from an overdose involving fentanyl. She had just gone to rehab, coroner says – AP News

Hayden Panettiere died from an overdose involving fentanyl. She had just gone to rehab, coroner says – AP News

September 22, 2026
Bodycam video shows arrest of 49ers owner

Bodycam video shows arrest of 49ers owner

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!