• bitcoinBitcoin(BTC)$77,266.000.04%
  • ethereumEthereum(ETH)$2,521.360.40%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$728.23-0.77%
  • rippleXRP(XRP)$1.370.35%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.990.31%
  • tronTRON(TRX)$0.3400680.17%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.07%
  • zcashZcash(ZEC)$1,135.460.43%
  • HyperliquidHyperliquid(HYPE)$79.411.00%
  • dogecoinDogecoin(DOGE)$0.0849840.72%
  • RainRain(RAIN)$0.0157472.75%
  • moneroMonero(XMR)$542.583.33%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.290.20%
  • chainlinkChainlink(LINK)$11.520.43%
  • leo-tokenLEO Token(LEO)$9.06-0.83%
  • cardanoCardano(ADA)$0.207622-0.49%
  • stellarStellar(XLM)$0.180233-0.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$225.36-2.20%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.70-0.37%
  • uniswapUniswap(UNI)$6.374.24%
  • CantonCanton(CC)$0.097982-0.95%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.69%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0750520.89%
  • avalanche-2Avalanche(AVAX)$7.42-0.28%
  • shiba-inuShiba Inu(SHIB)$0.0000051.64%
  • nearNEAR Protocol(NEAR)$2.36-0.63%
  • suiSui(SUI)$0.730.12%
  • crypto-com-chainCronos(CRO)$0.0599254.60%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.190.33%
  • tether-goldTether Gold(XAUT)$4,350.630.03%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$114.600.00%
  • BittensorBittensor(TAO)$234.920.42%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.20%
  • aaveAave(AAVE)$127.561.81%
  • pax-goldPAX Gold(PAXG)$4,355.990.03%
  • AsterAster(ASTER)$0.701.95%
  • mantleMantle(MNT)$0.55-4.47%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0570154.72%
  • polkadotPolkadot(DOT)$1.01-3.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers Shanghai AI Lab and SenseTime Propose MM-Grounding-DINO: An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

January 17, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Researchers Shanghai AI Lab and SenseTime Propose MM-Grounding-DINO: An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
ShareShareShareShareShare

Object detection plays a vital role in multi-modal understanding systems, where images are input into models to generate proposals aligned with text. This process is crucial for state-of-the-art models handling Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). OVD models are trained on base categories in zero-shot scenarios but must predict both base and novel categories within a broad vocabulary. PG provides a phrase to describe candidate categories and output corresponding boxes, while REC accurately identifies a target from text and outlines its position using a bounding box. Grounding-DINO addresses OVD, PG, and REC, gaining widespread adoption for diverse applications. 

https://arxiv.org/abs/2401.02361v2

Researchers from Shanghai AI Lab and SenseTime Research have developed MM-Grounding-DINO, a user-friendly and open-source pipeline created using the MMDetection toolbox. It utilizes diverse vision datasets for pre-training and a range of detection and grounding datasets for fine-tuning. A comprehensive analysis of reported results and detailed settings for reproducibility are provided. Through extensive experiments on benchmarks, MM-Grounding-DINO-Tiny surpasses the performance of the Grounding-DINO-Tiny baseline. 

https://arxiv.org/abs/2401.02361v2

MM-Grounding-DINO builds upon the foundation of Grounding-DINO. It operates by aligning textual descriptions with corresponding generated bounding boxes in images with varied shapes. The main components of the MM-Grounding-DINO include a text backbone responsible for extracting features from text, an image backbone for extracting features from images, a feature enhancer for thorough fusion of image and text features, a language-guided query selection module for initializing queries, and a cross-modality decoder for refining bounding boxes.

When presented with an image-text pair, MM-Grounding-DINO employs an image backbone to extract features from the image at various scales. Simultaneously, a text backbone extracts features from the accompanying text. These extracted features are input into a feature enhancer module, facilitating cross-modality fusion. Within this module, text and image features undergo fusion through a Bi-Attention Block, encompassing text-to-image and image-to-text cross-attention layers. Subsequently, the fused features undergo further enhancement through vanilla self-attention and deformable self-attention layers, followed by a Feedforward Network (FFN) layer.

The study presents an open, comprehensive pipeline for unified object grounding and detection covering OVD, PG, and REC tasks. The model’s performance is evaluated through a visualization-based analysis, which reveals inaccuracies in the ground-truth annotations of the evaluation dataset. The MM-Grounding-DINO model achieves state-of-the-art performance in zero-shot settings on COCO, with a mean average precision (mAP) of 52.5. The MM-Grounding-DINO model also outperforms fine-tuned models in various domains, including marine objects, brain tumor detection, urban street scenes, and people in paintings, setting new benchmarks for mAP. 

https://arxiv.org/abs/2401.02361v2

In conclusion, The study introduces a comprehensive and open pipeline for unified object grounding and detection, addressing tasks like OVD, PG, and REC. The model exhibits notable improvements in mAP across various datasets, such as COCO and LVIS, through fine-tuning. The model’s predictions’ precision surpasses existing annotations for specific objects. The authors propose an extensive evaluation framework facilitating systematic assessment across diverse datasets, including COCO, LVIS, RefCOCOg, Flickr30k Entities, ODinW1335, and Description Detection Dataset (D3).


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Is There Any Benefit To Restarting Your Gaming Handheld Regularly?
AI & Technology

Is There Any Benefit To Restarting Your Gaming Handheld Regularly?

September 12, 2026
Next Post
Meet the Press NOW — Aug.11

Meet the Press NOW — Aug.11

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How To Reset The Camera Settings On Your iPhone

How To Reset The Camera Settings On Your iPhone

September 8, 2026
Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

September 9, 2026
Artemis II moonshot commander and pilot are hanging up their astronaut suits at NASA – KSL.com

Artemis II moonshot commander and pilot are hanging up their astronaut suits at NASA – KSL.com

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!