• bitcoinBitcoin(BTC)$83,402.00-2.46%
  • ethereumEthereum(ETH)$2,642.77-2.97%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$770.44-1.19%
  • rippleXRP(XRP)$1.47-6.11%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$113.27-3.28%
  • tronTRON(TRX)$0.339717-0.64%
  • zcashZcash(ZEC)$1,479.20-9.69%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.45%
  • HyperliquidHyperliquid(HYPE)$91.20-4.17%
  • dogecoinDogecoin(DOGE)$0.092386-7.23%
  • moneroMonero(XMR)$544.32-3.64%
  • whitebitWhiteBIT Coin(WBT)$83.43-2.86%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.25-3.85%
  • cardanoCardano(ADA)$0.235740-5.73%
  • RainRain(RAIN)$0.011989-6.45%
  • leo-tokenLEO Token(LEO)$8.91-0.72%
  • stellarStellar(XLM)$0.199702-6.47%
  • bitcoin-cashBitcoin Cash(BCH)$334.13-3.75%
  • nearNEAR Protocol(NEAR)$4.39-7.05%
  • uniswapUniswap(UNI)$8.99-7.95%
  • litecoinLitecoin(LTC)$66.086.19%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • daiDai(DAI)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.03-6.90%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.107494-4.37%
  • hedera-hashgraphHedera(HBAR)$0.090113-5.23%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-3.13%
  • suiSui(SUI)$0.95-6.07%
  • shiba-inuShiba Inu(SHIB)$0.000006-6.79%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BittensorBittensor(TAO)$280.34-9.09%
  • crypto-com-chainCronos(CRO)$0.060535-7.35%
  • MemeCoreMemeCore(M)$1.22-3.77%
  • BitwayBitway(BTW)$1.037.96%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,267.91-1.02%
  • okbOKB(OKB)$117.85-3.08%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.07%
  • OndoOndo(ONDO)$0.4594305.86%
  • mantleMantle(MNT)$0.67-0.61%
  • aaveAave(AAVE)$136.80-7.39%
  • EthenaEthena(ENA)$0.206552-3.78%
  • polkadotPolkadot(DOT)$1.11-4.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Researchers Present Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

July 15, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Google DeepMind Researchers Present Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
ShareShareShareShareShare

Technological advancements in sensors, AI, and processing power have propelled robot navigation to new heights in the last several decades. To take robotics to the next level and make them a regular part of our lives, many studies suggest transferring the natural language space of ObjNav and VLN to the multimodal space so the robot can follow commands in both text and images at the same time. Researchers call this type of maritime activity Multimodal Instruction Navigation (MIN).

MIN encompasses a wide range of activities, including exploring the surroundings and following instructions for navigation. However, the use of a demonstration tour film that covers the entire region allows one to avoid investigation often altogether. 

YOU MAY ALSO LIKE

Revolut Is Piloting Facial Recognition At Store Checkouts In The UK

Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev

A Google DeepMind study presents and investigates a class of tasks called Multimodal Instruction Navigation with Tours (MINT). MINT uses demonstration tours and is concerned with carrying out multimodal user instructions. The remarkable capabilities of massive Vision-Language Models (VLMs) in language and picture interpretation and common-sense reasoning have recently demonstrated considerable promise in addressing MINT. On the other hand, VLMs on their own aren’t up to the task of solving MINT because of the following reasons:

  1. Many VLMs have a very limited quantity of input photos because of context-length limits. Because of this, an accurate understanding of huge environments is quite limited.
  2. Computed robot actions are necessary for solving MINT. The queries used to request these kinds of activities from robots are usually separate from the distribution that VLMs are (pre)trained to handle. Consequently, zero-shot navigation performance could be better. 

To address MINT, the team provides Mobility VLA, a hierarchical Vision-Language-Action (VLA) navigation policy that integrates the knowledge of the surroundings and the ability to reason intuitively from long-context VLMs with a strong low-level navigation policy built on topological networks. The high-level VLM uses the demonstration tour video and multimodal user guidance to locate the desired frame in the tour film. Following this, a conventional low-level policy takes the goal frame and constructs a topological graph offline from the tour frames at each time step. This graph is then used to create robot actions, also called waypoints. The fidelity issue with environment understanding was tackled by employing long-context VLMs, and the topological graph connected the VLM training distribution to the robot actions needed to solve MINT.

The team’s testing of Mobility VLA in a realistic (836m2) office setting and a more residential one yielded promising results. On complex MINT problems requiring intricate thinking, Mobility VLA achieved success rates of 86% and 90%, respectively, which is significantly higher than the baseline techniques. These findings reassure us about the capabilities of Mobility VLA in real-world scenarios.

Rather than exploring its surroundings autonomously, the present version of Mobility VLA depends on a demonstration trip. On the other hand, the demonstration tour provides a great opportunity to incorporate preexisting exploration methods like frontier or diffusion-based exploration.

The researchers highlight that unnatural user interactions are hindered by long VLM inference times. Users have to endure uncomfortable waiting times for robot responses due to the inference time of high-level VLMs, which is approximately 10-30 seconds. Caching the demonstration tour—which uses up around 99.9 percent of the input tokens—can greatly enhance inference speed. 

Given the light onboard compute demand (VLMs run on clouds) and the requirement of only RGB camera observations, Mobility VLA can be implemented on numerous robot incarnations. This potential for widespread deployment of Mobility VLA is a cause for optimism and a step forward in the field of robotics and AI. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Revolut Is Piloting Facial Recognition At Store Checkouts In The UK
AI & Technology

Revolut Is Piloting Facial Recognition At Store Checkouts In The UK

September 24, 2026
Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev
AI & Technology

Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev

September 24, 2026
A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
AI & Technology

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

September 24, 2026
Everything Announced At Meta Connect 2026
AI & Technology

Everything Announced At Meta Connect 2026

September 24, 2026
Next Post
Boeing’s Starliner docks at International Space Station

Boeing’s Starliner docks at International Space Station

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
D4vd abruptly parts ways with legal team in murder case

D4vd abruptly parts ways with legal team in murder case

September 20, 2026
1989: Dolly Parton on her ‘Steel Magnolias’ character

1989: Dolly Parton on her ‘Steel Magnolias’ character

September 24, 2026
Marine One radio issue while Trump onboard causes concern

Marine One radio issue while Trump onboard causes concern

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!