• bitcoinBitcoin(BTC)$78,967.00-1.37%
  • ethereumEthereum(ETH)$2,486.55-1.18%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$739.80-1.39%
  • rippleXRP(XRP)$1.39-1.63%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.36-2.60%
  • tronTRON(TRX)$0.335009-0.07%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,130.68-4.94%
  • HyperliquidHyperliquid(HYPE)$84.35-3.28%
  • dogecoinDogecoin(DOGE)$0.090181-0.17%
  • RainRain(RAIN)$0.016275-2.74%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$511.89-3.97%
  • chainlinkChainlink(LINK)$12.64-3.96%
  • whitebitWhiteBIT Coin(WBT)$76.453.59%
  • leo-tokenLEO Token(LEO)$9.20-0.43%
  • cardanoCardano(ADA)$0.218201-2.13%
  • stellarStellar(XLM)$0.1898140.21%
  • bitcoin-cashBitcoin Cash(BCH)$258.04-0.19%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • uniswapUniswap(UNI)$6.96-2.26%
  • litecoinLitecoin(LTC)$55.341.71%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.105616-4.67%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-3.26%
  • hedera-hashgraphHedera(HBAR)$0.0818690.41%
  • avalanche-2Avalanche(AVAX)$8.031.91%
  • suiSui(SUI)$0.821.33%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.45%
  • nearNEAR Protocol(NEAR)$2.31-4.30%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056764-1.98%
  • tether-goldTether Gold(XAUT)$4,419.420.29%
  • MemeCoreMemeCore(M)$1.163.88%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BittensorBittensor(TAO)$258.23-6.03%
  • okbOKB(OKB)$116.882.32%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.35%
  • AsterAster(ASTER)$0.76-4.04%
  • mantleMantle(MNT)$0.62-1.18%
  • aaveAave(AAVE)$131.00-2.55%
  • pax-goldPAX Gold(PAXG)$4,422.820.29%
  • OndoOndo(ONDO)$0.380170-2.36%
  • polkadotPolkadot(DOT)$1.067.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CMU Researchers Introduce BUTD-DETR: An Artificial Intelligence (AI) Model That Conditions Directly On A Language Utterance And Detects All Objects That The Utterance Mentions

July 23, 2023
in AI & Technology
Reading Time: 5 mins read
A A
CMU Researchers Introduce BUTD-DETR: An Artificial Intelligence (AI) Model That Conditions Directly On A Language Utterance And Detects All Objects That The Utterance Mentions
ShareShareShareShareShare

Finding all of the “objects” in a given image is the groundwork of computer vision. By creating a vocabulary of categories and training a model to recognize instances of this vocabulary, one may avoid the question, “What is an Object?” The situation worsens when one tries to use these object detectors as practical home agents. Models often learn to pick the referenced item from a pool of object suggestions a pre-trained detector offers when requested to ground referential utterances in 2D or 3D settings. As a result, the detector may miss utterances that relate to finer-grained visual things, such as the chair, the chair leg, or the chair leg’s front tip.

The research team presents a Bottom-up, Top-Down DEtection TRansformer (BUTD-DETR pron. Beauty-DETER) as a model that conditions directly on a spoken utterance and finds all mentioned items. BUTD-DETR functions as a normal object detector when the utterance is a list of object categories. It is trained on image-language pairings tagged with the bounding boxes for all items alluded to in the speech, as well as fixed-vocab object detection datasets. However, with a few tweaks, BUTD-DETR may also anchor language phrases in 3D point clouds and 2D pictures.

Instead of randomly picking them from a pool, BUTD-DETR decodes object boxes by paying attention to verbal and visual input. The bottom-up, task-agnostic attention can overlook some details when locating an item, but language-directed attention fills in the gaps. A scene and a spoken utterance are used as input for the model. Suggestions for boxes are extracted using a detector that has already been trained. Next, visual, box, and linguistic tokens are extracted from the scene, boxes, and speech using per-modality-specific encoders. These tokens gain meaning within their context by paying attention to one another. Refined visual tickets kick off object queries that decode boxes and span over many streams.

🚀 Build high-quality training datasets with Kili Technology and solve NLP machine learning challenges to develop powerful ML applications

The practice of object detection is an example of grounded referential language, where the utterance is the category label for the thing being detected. Researchers use object detection as the referential grounding of detection prompts by randomly selecting certain object categories from the detector’s vocabulary and generating synthetic utterances by sequencing them (for example, “Couch. Person. Chair.”). These detection cues are used as supplemental supervision information, with the goal being to find all occurrences of the category labels specified in the cue inside the scene. The model is instructed to avoid making box associations for category labels for which there are no visual input examples (such as “person” in the example above). In this approach, a single model can ground language and recognize objects while sharing the same training data for both tasks.

Outcomes

The developed MDETR-3D equivalent performs poorly compared to earlier models, whereas BUTD-DETR achieves state-of-the-art performance on 3D language grounding.

BUTD-DETR also functions in the 2D domain, and with architectural enhancements like deformable attention, it achieves performance on par with MDETR while converging twice as quickly. The approach takes a step toward unifying grounding models for 2D and 3D since it can be easily adapted to function in both dimensions with minor adjustments.

For all 3D language grounding benchmarks, BUTD-DETR demonstrates significant performance gains over state-of-the-art methods (SR3D, NR3D, ScanRefer). In addition, it was the best submission at the ECCV workshop on Language for 3D Scenes, where the ReferIt3D competition was conducted. However, when trained on massive data, BUTD-DETR may compete with the best existing approaches for 2D language grounding benchmarks. Specifically, researchers’ efficient deformable attention to the 2D model allows the model to converge twice as rapidly as state-of-the-art MDETR.

The video below describes the complete workflow.


Check out the Paper, Github, and CMU Blog. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our Reddit Page, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Uber, Wayve Unleash Supervised Robotaxis in London

Chip Suppliers Bullish on AI Buildout

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🔥 Gain a competitive
edge with data: Actionable market intelligence for global brands, retailers, analysts, and investors. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Uber, Wayve Unleash Supervised Robotaxis in London
AI & Technology

Uber, Wayve Unleash Supervised Robotaxis in London

September 8, 2026
Chip Suppliers Bullish on AI Buildout
AI & Technology

Chip Suppliers Bullish on AI Buildout

September 8, 2026
Anthropic’s  Billion Credit Line Sets Stage for IPO
AI & Technology

Anthropic’s $15 Billion Credit Line Sets Stage for IPO

September 8, 2026
How To Reset The Camera Settings On Your iPhone
AI & Technology

How To Reset The Camera Settings On Your iPhone

September 8, 2026
Next Post
Formula 1 picks, odds, race time: 2023 Hungarian Grand Prix predictions, F1 best bets from proven model

Formula 1 picks, odds, race time: 2023 Hungarian Grand Prix predictions, F1 best bets from proven model

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The Pros And Cons Of Using Wireless Android Auto

The Pros And Cons Of Using Wireless Android Auto

September 6, 2026
Trump encouraged Army secretary to stay in job – The Washington Post

Trump encouraged Army secretary to stay in job – The Washington Post

September 3, 2026
Enterprises put non-Nvidia chips 14 points ahead of Nvidia’s next-gen GPUs on their evaluation lists

Enterprises put non-Nvidia chips 14 points ahead of Nvidia’s next-gen GPUs on their evaluation lists

September 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!