• bitcoinBitcoin(BTC)$78,480.000.12%
  • ethereumEthereum(ETH)$2,485.500.06%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$740.04-1.16%
  • rippleXRP(XRP)$1.42-0.64%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$103.41-0.08%
  • tronTRON(TRX)$0.3399640.33%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.61%
  • zcashZcash(ZEC)$1,266.849.06%
  • HyperliquidHyperliquid(HYPE)$85.982.09%
  • dogecoinDogecoin(DOGE)$0.088814-0.77%
  • RainRain(RAIN)$0.016134-2.96%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$510.872.32%
  • whitebitWhiteBIT Coin(WBT)$81.09-0.23%
  • chainlinkChainlink(LINK)$11.93-5.73%
  • leo-tokenLEO Token(LEO)$9.18-0.25%
  • cardanoCardano(ADA)$0.217266-2.07%
  • stellarStellar(XLM)$0.184918-1.98%
  • bitcoin-cashBitcoin Cash(BCH)$257.970.47%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$54.120.17%
  • uniswapUniswap(UNI)$6.60-3.10%
  • CantonCanton(CC)$0.103870-2.74%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-1.45%
  • avalanche-2Avalanche(AVAX)$7.93-0.59%
  • hedera-hashgraphHedera(HBAR)$0.078032-1.81%
  • nearNEAR Protocol(NEAR)$2.6011.03%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.80-1.38%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.67%
  • crypto-com-chainCronos(CRO)$0.059285-1.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.19-0.61%
  • tether-goldTether Gold(XAUT)$4,402.880.69%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$261.291.19%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.08-0.69%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.14%
  • mantleMantle(MNT)$0.62-2.01%
  • AsterAster(ASTER)$0.75-1.04%
  • aaveAave(AAVE)$128.29-0.52%
  • Pump.funPump.fun(PUMP)$0.0047348.34%
  • polkadotPolkadot(DOT)$1.12-6.78%
  • pax-goldPAX Gold(PAXG)$4,406.340.75%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes LLM-Grounder: A Zero-Shot, Open-Vocabulary Approach to 3D Visual Grounding for Next-Gen Household Robots

September 29, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Proposes LLM-Grounder: A Zero-Shot, Open-Vocabulary Approach to 3D Visual Grounding for Next-Gen Household Robots
ShareShareShareShareShare

Understanding their surroundings in three dimensions (3D vision) is essential for domestic robots to perform tasks like navigation, manipulation, and answering queries. At the same time, current methods can need help to deal with complicated language queries or rely excessively on large amounts of labeled data.

ChatGPT and GPT-4 are just two examples of large language models (LLMs) with amazing language understanding skills, such as planning and tool use. By breaking down large problems into smaller ones and learning when, what, and how to employ a tool to finish sub-tasks, LLMs can be deployed as agents to solve complicated problems. Parsing the compositional language into smaller semantic constituents, interacting with tools and environment to collect feedback, and reasoning with spatial and commonsense knowledge to iteratively ground the language to the target object are all necessary for 3D visual grounding with complex natural language queries.

Nikhil Madaan and researchers from the University of Michigan and New York University present LLM-Grounder, a novel zero-shot LLM-agent-based 3D visual grounding process that uses an open vocabulary. While a visual grounder excels at grounding basic noun phrases, the team hypothesizes that an LLM can help mitigate the “bag-of-words” limitation of a CLIP-based visual grounder by taking on the challenging language deconstruction, spatial, and commonsense reasoning tasks itself.

LLM-Grounder relies on an LLM to coordinate the grounding procedure. After receiving a natural language query, the LLM breaks it down into its parts or semantic ideas, such as the type of object sought, its properties (including color, shape, and material), landmarks, and geographical relationships. To locate each concept in the scene, these sub-queries are sent to a visual grounder tool supported by OpenScene or LERF, both of which are CLIP-based open-vocabulary 3D visual grounding approaches. The visual grounder suggests a few bounding boxes based on where the most promising candidates for a notion are located in the scene. The visual grounder tools compute spatial information, such as object volumes and distances to landmarks, and feed that data back to the LLM agent, allowing the latter to make a more well-rounded assessment of the situation in terms of spatial relation and common sense and ultimately choose a candidate that best matches all criteria in the original query. The LLM agent will continue to cycle through these steps until it reaches a decision. The researchers take a step beyond existing neural-symbolic methods by using the surrounding context in their analysis.

The team highlights that the method doesn’t require labeled data for training. Given the semantic variety of 3D settings and the scarcity of 3D-text labeled data, its open-vocabulary and zero-shot generalization to novel 3D scenes and arbitrary text queries is an attractive feature. Using the ScanRefer benchmark, the researchers conduct experimental evaluations of LLM-Grounder. The ability to interpret compositional visual referential expressions is important to evaluating grounding in 3D vision language in this benchmark. The results show that the method outperforms state-of-the-art zero-shot grounding accuracy on ScanRefer with no labeled data. It also enhances the grounding capacity of open-vocabulary approaches like OpenScene and LERF. Based on their erasure research, LLM improves grounding capabilities in proportion to the complexity of the language query. These show the efficiency of the LLM-Grounder method for 3D vision language problems, making it ideal for robotics applications where awareness of context and the ability to quickly and accurately react to changing questions are crucial.


Check out the Paper and Demo. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI
AI & Technology

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

September 9, 2026
Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI
AI & Technology

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
Everything Announced During Nintendo Direct
AI & Technology

Everything Announced During Nintendo Direct

September 9, 2026
Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Next Post
World Wide Web Founder on Trump, Fake News, AI

World Wide Web Founder on Trump, Fake News, AI

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Emmy Awards Night 1 Winners List (Updating Live) – Deadline

Emmy Awards Night 1 Winners List (Updating Live) – Deadline

September 6, 2026
Morning News NOW Full Episode – July 24

Morning News NOW Full Episode – July 24

September 5, 2026
TODAY co-host opens up about husband’s battle with cancer: ‘He encouraged me to keep writing’

TODAY co-host opens up about husband’s battle with cancer: ‘He encouraged me to keep writing’

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!