• bitcoinBitcoin(BTC)$79,831.000.18%
  • ethereumEthereum(ETH)$2,482.791.18%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$767.046.44%
  • rippleXRP(XRP)$1.421.19%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.331.42%
  • tronTRON(TRX)$0.3340630.80%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.061.94%
  • HyperliquidHyperliquid(HYPE)$85.351.03%
  • zcashZcash(ZEC)$1,026.90-0.06%
  • dogecoinDogecoin(DOGE)$0.0899866.17%
  • RainRain(RAIN)$0.0169852.49%
  • moneroMonero(XMR)$553.876.13%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.053.70%
  • whitebitWhiteBIT Coin(WBT)$73.470.32%
  • leo-tokenLEO Token(LEO)$9.280.75%
  • cardanoCardano(ADA)$0.2195153.94%
  • stellarStellar(XLM)$0.1844122.84%
  • bitcoin-cashBitcoin Cash(BCH)$257.043.69%
  • daiDai(DAI)$1.000.03%
  • uniswapUniswap(UNI)$7.0714.40%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • CantonCanton(CC)$0.1090951.91%
  • litecoinLitecoin(LTC)$54.838.04%
  • USD1USD1(USD1)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.432.78%
  • hedera-hashgraphHedera(HBAR)$0.0802612.73%
  • avalanche-2Avalanche(AVAX)$7.603.04%
  • suiSui(SUI)$0.805.66%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000054.80%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.190.08%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0567451.65%
  • tether-goldTether Gold(XAUT)$4,426.73-0.03%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.130.87%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.484.92%
  • BittensorBittensor(TAO)$236.244.92%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.09%
  • AsterAster(ASTER)$0.785.00%
  • aaveAave(AAVE)$134.693.39%
  • mantleMantle(MNT)$0.592.15%
  • pax-goldPAX Gold(PAXG)$4,432.74-0.01%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0570721.74%
  • OndoOndo(ONDO)$0.3714014.11%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Researchers Introduce RT-2: A Novel Vision-Language-Action (VLA) Model that Learns from both Web and Robotics Data and Turns it into Action

August 4, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Google DeepMind Researchers Introduce RT-2: A Novel Vision-Language-Action (VLA) Model that Learns from both Web and Robotics Data and Turns it into Action
ShareShareShareShareShare

Large language models can enable fluent text generation, emergent problem-solving, and creative generation of prose and code. In contrast, vision-language models enable open-vocabulary visual recognition and can even make complex inferences about object-agent interactions in images. The best way for robots to learn new skills needs to be clarified. Compared to the billions of tokens and photos used to train the most advanced language and vision-language models on the web, the amount of data collected from robots is unlikely to be comparable. However, it is also challenging to immediately adapt such models to robotic activities since these models reason about semantics, labels, and textual prompts. In contrast, robots must be instructed in low-level actions, such as those using the Cartesian end-effector.

Google Deepmind’s research aims to improve generalization and enable emergent semantic reasoning by directly incorporating vision-language models trained on Internet-scale data into end-to-end robotic control. With the help of web-based language and vision-language data, we aim to make a single, comprehensively trained model to learn to link robot observations to actions. They propose fine-tuning state-of-the-art vision-language models together using data from robot trajectories and large-scale visual question-answering exercises conducted over the Internet. In contrast to other methods, they propose a straightforward, all-purpose recipe: express robotic actions as text tokens and incorporate them directly into the model’s training set as natural language tokens would. Researchers study vision-language-action models (VLA), and RT-2 instantiates one such model. Through rigorous testing (6k assessment trials), they could ascertain that RT-2 acquired various emergent skills through Internet-scale training and that the technique led to performant robotic policies.

Google DeepMind unveiled RT-2, a Transformer-based model trained on the web-sourced text and images that can directly perform robotic operations, as a follow-up to its Robotics Transformer model 1. They use robot actions to represent a second language that can be converted into text tokens and taught alongside large-scale vision-language datasets available online. Inference involves de-tokenizing text tokens into robot behaviors that can then be controlled via a feedback loop. This permits transferring some of the generalization, semantic comprehension, and reasoning of vision-language models to learning robotic policies. On the project website, accessible at https://robotics-transformer2.github.io/, the team behind RT-2 provides live demonstrations of its use. 

The model retains the ability to deploy its physical skills in ways consistent with the distribution found in the robot data. Still, it also learns to use those skills in novel contexts by reading visuals and linguistic commands using knowledge gathered from the web. Even though semantic cues like precise numbers or icons aren’t included in the robot data, the model can repurpose its learned pick-and-place skills. No such relations were supplied in the robot demos, yet the model could pick the correct object and position it in the correct location. In addition, the model can make even more complex semantic inferences if the command is supplemented with a chain of thought prompting, such as knowing that a rock is the best choice for an improvised hammer or an energy drink is the best choice for someone tired.

Google DeepMind’s key contribution is RT-2, a family of models created by fine-tuning huge vision-language models trained on web-scale data to serve as generalizable and semantically aware robotic rules. Experiments probe models with as much as 55B parameters, learned from publicly available data and annotated with robotic motion commands. Across 6,000 robotic evaluations, they demonstrate that RT-2 enables considerable advances in generalization over objects, scenes, and instructions and displays a range of emergent abilities that are a byproduct of web-scale vision-language pretraining. 

Key Features

  • The reasoning, symbol interpretation, and human identification capabilities of RT-2 can be used in a wide range of practical scenarios. 
  • The results of RT-2 demonstrate that pretraining VLMs using robotic data can turn them into powerful vision-language-action (VLA) models that can directly control a robot.
  • A promising direction to pursue is to construct a general-purpose physical robot that can think, problem-solve, and interpret information for completing various activities in the actual world, like RT-2.
  • Its adaptability and efficiency in handling various tasks are displayed in RT-2’s capacity to transfer information from language and visual training data to robot movements.

Limitations

Despite its encouraging generalization properties, RT-2 suffers from several drawbacks. Although studies suggest that incorporating web-scale pretraining through VLMs improves generalization across semantic and visual concepts, this does not give the robot any new abilities regarding its capacity to perform motions. Though the model can only use the physical abilities found in the robot data in novel ways, it does learn to make better use of its abilities. They attribute this to a need for more diversity in the sample along the dimensions of competence. New data-gathering paradigms, such as films of humans, present an intriguing opportunity for future research into acquiring new skills.

To sum it up, Google DeepMind researchers demonstrated that big VLA models could be run in real-time, but this was at a considerable computational expense. As these methods are applied to situations requiring high-frequency control, real-time inference risks become a significant bottleneck. Quantization and distillation approaches that could let such models operate faster or on cheaper hardware are attractive areas for future study. This is related to another existing restriction in that relatively few VLM models can be utilized to develop RT-2.

Researchers from Google DeepMind summarized the process of training vision-language-action (VLA) models by integrating pretraining with vision-language models (VLMs) and data from robotics. They then introduced two variants of VLAs (RT-2-PaLM-E and RT-2-PaLI-X) that PaLM-E and PaLI-X, respectively inspired. These models are fine-tuned with data on robotic trajectories to generate robot actions, which are tokenized as text. More crucially, they demonstrated that the technique improves generalization performance and emergent capabilities inherited from web-scale vision-language pretraining, leading to very effective robotic policies. According to Google DeepMind, the discipline of robot learning is now strategically positioned to profit from improvements in other fields thanks to this straightforward and universal methodology. 


Check out the Paper and Reference Article. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 27k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Is The Steam Deck Still Worth It In 2026?

The Reasons Rugged Laptops Are Rarely Bought By Consumers

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🔥 Use SQL to predict the future (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Is The Steam Deck Still Worth It In 2026?
AI & Technology

Is The Steam Deck Still Worth It In 2026?

September 5, 2026
The Reasons Rugged Laptops Are Rarely Bought By Consumers
AI & Technology

The Reasons Rugged Laptops Are Rarely Bought By Consumers

September 5, 2026
GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI
AI & Technology

GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

September 5, 2026
You’re Probably Wasting These Keys On Your Keyboard — Here’s How To Remap Them
AI & Technology

You’re Probably Wasting These Keys On Your Keyboard — Here’s How To Remap Them

September 5, 2026
Next Post
Markets’ Swoon Lingered Into the Afternoon With the Dow Racking Up Big Losses

Markets' Swoon Lingered Into the Afternoon With the Dow Racking Up Big Losses

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Modding Platform Nexus Mods Now Owns SteamDB

Modding Platform Nexus Mods Now Owns SteamDB

September 2, 2026
Full Episode: TODAY Show – July 29

Full Episode: TODAY Show – July 29

September 3, 2026
Nolan Wells’ state autopsy completed 

Nolan Wells’ state autopsy completed 

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!