• bitcoinBitcoin(BTC)$83,383.00-0.02%
  • ethereumEthereum(ETH)$2,683.460.37%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$767.030.59%
  • rippleXRP(XRP)$1.49-0.85%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.91-1.45%
  • tronTRON(TRX)$0.3379610.80%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.28%
  • zcashZcash(ZEC)$1,407.91-0.65%
  • HyperliquidHyperliquid(HYPE)$88.973.07%
  • dogecoinDogecoin(DOGE)$0.0948190.91%
  • chainlinkChainlink(LINK)$14.41-0.45%
  • moneroMonero(XMR)$544.20-0.20%
  • whitebitWhiteBIT Coin(WBT)$83.350.01%
  • USDSUSDS(USDS)$1.000.02%
  • cardanoCardano(ADA)$0.2477060.77%
  • RainRain(RAIN)$0.012258-2.73%
  • leo-tokenLEO Token(LEO)$8.83-2.30%
  • stellarStellar(XLM)$0.2265852.34%
  • nearNEAR Protocol(NEAR)$5.225.31%
  • bitcoin-cashBitcoin Cash(BCH)$305.31-0.72%
  • uniswapUniswap(UNI)$8.80-0.12%
  • litecoinLitecoin(LTC)$67.04-0.31%
  • CantonCanton(CC)$0.1265950.00%
  • avalanche-2Avalanche(AVAX)$10.99-3.30%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.171.07%
  • daiDai(DAI)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.1041480.46%
  • USD1USD1(USD1)$1.00-0.02%
  • quant-networkQuant(QNT)$292.372.96%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.500.81%
  • BitwayBitway(BTW)$1.32-6.29%
  • BittensorBittensor(TAO)$302.28-0.16%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.82%
  • crypto-com-chainCronos(CRO)$0.0679870.75%
  • tether-goldTether Gold(XAUT)$4,156.07-0.64%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • EthenaEthena(ENA)$0.2664236.41%
  • Pump.funPump.fun(PUMP)$0.005719-2.28%
  • okbOKB(OKB)$121.530.30%
  • aaveAave(AAVE)$162.301.31%
  • OndoOndo(ONDO)$0.500.75%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.050.72%
  • mantleMantle(MNT)$0.694.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.04%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes an Effective Paradigm for Large Scale Vision-and-Language Navigation (VLN) Training and Quantitatively Evaluates the Influence of Each Component in the Pipeline

August 3, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Proposes an Effective Paradigm for Large Scale Vision-and-Language Navigation (VLN) Training and Quantitatively Evaluates the Influence of Each Component in the Pipeline
ShareShareShareShareShare

Several human demos have been collected for learning visual navigation, and recent huge datasets contain hundreds of interactive scenarios, both of which have led to significant improvements in agent performance. However, getting to such massive training requires solving a number of key sub-problems, such as how to construct navigation graphs, restore corrupted rendered images, and generate navigational instructions. All of this has a major impact on the quality of the data collected and thus should be thoroughly explored. 

It is necessary to research how to efficiently utilize large-scale data to benefit the training of navigational agents appropriately, and an agent that can understand human natural language and navigate in photorealistic surroundings is a sophisticated and modularized system.

To train large-scale vision-and-language navigation networks (VLNs), researchers from the Australian National University, OpenGVLab, Shanghai AI Laboratory, UNC, Chapel Hill, University of Adelaide, and Adobe Research offer a new paradigm by statistically assessing the impact of each component in the pipeline. Using the Habitat simulator, they use environments from the HM3D and Gibson datasets and construct navigation graphs for the environments. They sample new trajectories, create instructions, and train agents to solve downstream navigation problems. 

In contrast to prior methods like AutoVLN and MARVAL, these navigation graphs are constructed with an excessive viewpoint sampling and aggregation procedure, employing the graph creation heuristic introduced in. This approach yields fully-connected networks with extensive outdoor coverage. 

The researchers also train the Co-Modulated GAN to generate photorealistic images from the broken, deformed, or missing sections in corrupted generated images from HM3D and Gibson settings, reducing visual data noise’s impact. In contrast to MARVAL, this large-scale training regime is fully reproducible and straightforward to execute while significantly improving the agent’s performance.

Extensive experiments show that if the agent is to perform better on downstream tasks with specific instructions, such as R2R, the navigation graph must be fully traversable. Furthermore, they demonstrate the benefits of recovering photorealistic images from generated images, particularly for the low-quality 3D scans from the Gibson habitats. Findings also indicate that agents can generally use more diverse visual data and can improve their generalization to novel contexts by learning from new scenes rather than just more data. 

Additionally, the team verifies that an agent trained with augmented instructions provided by a basic LSTM-based model can perform well on various navigation tasks. They conclude that the agent’s generalization capacity can be improved by integrating the augmented data with the original data during pre-training and fine-tuning. 

Surprisingly, by using the above analysis as guidelines for data augmentation and agent training, the proposed VLN model can achieve 80% SR on the R2R test split via simple imitation learning without pre-exploration, beam search, or model ensembling and eliminate the navigation gap between seen and unseen environments. This result is a huge improvement over the previous best approach (73%), bringing the performance gap to within 6 percentage points of human levels. The approach to several language-guided visual navigation challenges, such as CVDN and REVERIE, has pushed the state-of-the-art forward. The VLN performance is improved by 5% SR in the continuous environments (R2R-CE), a more realistic yet challenging scenario, even though the enhanced data is discrete. 


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 27k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

How To Watch NASA’s Crew-13 Launch

Elon Musk And Palmer Luckey Will Advise The Government On The Future Of Warfare

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🔥 Use SQL to predict the future (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Watch NASA’s Crew-13 Launch
AI & Technology

How To Watch NASA’s Crew-13 Launch

October 1, 2026
Elon Musk And Palmer Luckey Will Advise The Government On The Future Of Warfare
AI & Technology

Elon Musk And Palmer Luckey Will Advise The Government On The Future Of Warfare

September 30, 2026
Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense
AI & Technology

Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense

September 30, 2026
OpenAI Releases GPT-6.1 Sol: Near-Astra Coding and Computer Use at One-Fifth of Astra’s Token Price
AI & Technology

OpenAI Releases GPT-6.1 Sol: Near-Astra Coding and Computer Use at One-Fifth of Astra’s Token Price

September 30, 2026
Next Post
SpaceX Launches Biological Printer and Nickelodeon Slime Aboard Falcon 9 Rocket

SpaceX Launches Biological Printer and Nickelodeon Slime Aboard Falcon 9 Rocket

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

September 26, 2026
New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

September 26, 2026
Candy Crush Maker Signs Collective Bargaining Agreement, Averting Strike

Candy Crush Maker Signs Collective Bargaining Agreement, Averting Strike

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!