• bitcoinBitcoin(BTC)$79,417.00-0.62%
  • ethereumEthereum(ETH)$2,489.63-0.46%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$744.71-1.74%
  • rippleXRP(XRP)$1.40-1.53%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.93-1.45%
  • tronTRON(TRX)$0.3363650.41%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,193.350.75%
  • HyperliquidHyperliquid(HYPE)$87.75-0.90%
  • dogecoinDogecoin(DOGE)$0.089788-0.30%
  • RainRain(RAIN)$0.016476-3.18%
  • moneroMonero(XMR)$533.74-0.92%
  • chainlinkChainlink(LINK)$13.187.37%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$73.16-0.68%
  • leo-tokenLEO Token(LEO)$9.13-2.20%
  • cardanoCardano(ADA)$0.219644-0.53%
  • stellarStellar(XLM)$0.1911692.23%
  • bitcoin-cashBitcoin Cash(BCH)$256.12-1.76%
  • daiDai(DAI)$1.00-0.01%
  • uniswapUniswap(UNI)$7.03-0.61%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$55.822.26%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.106530-3.72%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-0.80%
  • hedera-hashgraphHedera(HBAR)$0.080727-0.91%
  • avalanche-2Avalanche(AVAX)$7.862.22%
  • suiSui(SUI)$0.810.67%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.33%
  • nearNEAR Protocol(NEAR)$2.35-1.89%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0573580.14%
  • tether-goldTether Gold(XAUT)$4,383.60-0.90%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.13-0.19%
  • BittensorBittensor(TAO)$264.329.12%
  • okbOKB(OKB)$115.471.33%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.49%
  • AsterAster(ASTER)$0.801.34%
  • mantleMantle(MNT)$0.645.38%
  • aaveAave(AAVE)$133.78-1.30%
  • pax-goldPAX Gold(PAXG)$4,386.65-0.98%
  • OndoOndo(ONDO)$0.3848241.30%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056620-0.54%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MIT Researchers Created a New Annotated Synthetic Dataset of Images that Depict a Wide Range of Scenarios to Help Machine-Learning Models Understand the Concepts in a Scene

September 16, 2023
in AI & Technology
Reading Time: 5 mins read
A A
MIT Researchers Created a New Annotated Synthetic Dataset of Images that Depict a Wide Range of Scenarios to Help Machine-Learning Models Understand the Concepts in a Scene
ShareShareShareShareShare

Large-scale pre-trained Vision and language models have demonstrated remarkable performance in numerous applications, allowing for the replacement of a fixed set of supported classes with zero-shot open vocabulary reasoning over (nearly arbitrary) natural language queries. However, recent research has revealed a fundamental flaw in these models. For instance, their inability to comprehend Visual Language Concepts (VLC) that extend “beyond nouns,” such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or their difficulty with compositional reasoning, such as comprehending the significance of the word order in a sentence.

Vision and language models, powerful machine-learning algorithms that learn to match text with images, have demonstrated remarkable results when requested to generate video captions or summaries. While these models excel at distinguishing objects, they frequently need help comprehending concepts, such as the attributes of things or the arrangement of items in a scene. For example, a vision and language model may perceive the cup and table in an image but fail to comprehend that the cup is atop the table.

Researchers from MIT have demonstrated a new technique that employs computer-generated data to assist vision and language models in overcoming this deficiency. In particular, they propose enhancing the VLC and compositionality aspects of the generated visual and text data and then using these data to fine-tune VL models by instructing them to pay closer attention to these characteristics. Moreover, in addition to being essentially free and infinitely scalable, synthetic data can also be free of the privacy concerns that always accompany real data. Creating synthetic data that can be effectively used to enhance VLC and compositionality aspects of VL models pre-trained on massive amounts of real data presents additional technical challenges. In contrast to most prior work on generating synthetic visual data, they must develop images and text describing a scene’s compositional elements. In addition, they generate synthetic videos that utilize genuine physical 3D simulation, such as diverse 3D environments and diverse 3D objects, human motions, and action assets], added interaction with things, and various camera angles.

Previous works utilized motion assets to generate synthetic data, but the visual data was not accompanied by textual captions and needed to be designed with compositionality in mind. Researchers contribute to Synthetic Visual Concepts (SyViC), a large (million-scale) generated synthetic VL dataset with rich textual captions readily extendable via the data synthesis code and all previously generated million-scale synthetic data.

Making Contributions

  • Researchers contribute SyViC – a million-scale synthetic dataset with rich textual annotations designed to enhance VLC comprehension and compositional reasoning in VL models, as well as the methodology and generation of codebase 2 for its synthesis and potential extensibility.
  • Effective general VL model fine-tuning that leverages SyViC data to improve the characteristics of strong pre-trained VL models without compromising their zero-shot performance.
  • Experimental results and a comprehensive ablation study demonstrate a significant (over 10% in some cases) improvement in VLC comprehension and compositional reasoning, as measured on the most recent VL-Checklist, ARO, and Winoground benchmarks and validated on the most popular CLIP model and its derivatives (e.g., the most recent CyCLIP.

The Results

Variants of all models were generated using the proposed method and SyViC synthetic data. Before fine-tuning on SyViC, each model is compared to its respective source model trained on large-scale real data. According to the findings of the researchers, both the SyViC synthetic data and the proposed fine-tuning recipe demonstrate significant enhancements over their respective source baselines. In addition, researchers illustrate the individual VLC metrics improvements acquired for CLIP in VL-Checklist and ARO benchmarks in showing up to 9.1% and 12.6% respective absolute improvements. This demonstrates the efficiency and potential of the method and SyViC synthetic data for enhancing VLC comprehension and compositional reasoning in VL models.

Try here https://synthetic-vic.github.io/ 

Limitations

While researchers have obtained quite promising results in three different benchmarks, there are limitations to their work. As an illustration, the graphics simulator has a simplified model of lighting, sensor noise, and reflectance functions compared to the actual world, which may affect color constancy robustness. More sophisticated domain adaptation and rendering techniques are likely required to enhance the outcomes further. In addition, a more in-depth examination of the scaling laws for synthetic data would be an excellent way to completely realize the work’s potential. 

To summarize 

Large vision and language models have dictated the status quo in computer vision and multimodal perception, achieving cutting-edge results in several difficult benchmarks. However, existing models need help with compositional reasoning and comprehending concepts beyond object nouns, such as attributes and relationships. This is the first investigation into whether synthetic data can mitigate these deficiencies. Researchers at MIT proposed a data generation pipeline to create a million-scale dataset of synthetic images and accompanying captions and an efficient fine-tuning strategy with a comprehensive analysis to improve the compositional and concept understanding capabilities of multimodal models without compromising their zero-shot classification performance.


Check out the Project Page and Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Proteomic Aging Clocks Track Biological Age Reversal in Rentosertib Trial – Unite.AI

Google, Cathay Pacific Expand Contrail Avoidance Trials in Asia-Pacific – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Proteomic Aging Clocks Track Biological Age Reversal in Rentosertib Trial – Unite.AI
AI & Technology

Proteomic Aging Clocks Track Biological Age Reversal in Rentosertib Trial – Unite.AI

September 7, 2026
Google, Cathay Pacific Expand Contrail Avoidance Trials in Asia-Pacific – Unite.AI
AI & Technology

Google, Cathay Pacific Expand Contrail Avoidance Trials in Asia-Pacific – Unite.AI

September 7, 2026
IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B
AI & Technology

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

September 7, 2026
Is It Safe To Buy A Refurbished iPhone From Walmart?
AI & Technology

Is It Safe To Buy A Refurbished iPhone From Walmart?

September 7, 2026
Next Post
The Highlights From Fox’s Annual Meeting

The Highlights From Fox's Annual Meeting

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Uncanny and unappetizing: appetites spoil as AI images take over food menus – The Guardian

Uncanny and unappetizing: appetites spoil as AI images take over food menus – The Guardian

September 6, 2026
Texas’ power grid meets record demand with investments in solar energy

Texas’ power grid meets record demand with investments in solar energy

September 6, 2026
Young adults under 21 – banned from sports betting – traded .4B on Kalshi this year

Young adults under 21 – banned from sports betting – traded $5.4B on Kalshi this year

August 31, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!