• bitcoinBitcoin(BTC)$81,158.005.06%
  • ethereumEthereum(ETH)$2,505.464.86%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$724.835.30%
  • rippleXRP(XRP)$1.457.25%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.783.60%
  • tronTRON(TRX)$0.3306811.82%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.79%
  • HyperliquidHyperliquid(HYPE)$87.586.84%
  • zcashZcash(ZEC)$952.2716.93%
  • dogecoinDogecoin(DOGE)$0.0876397.30%
  • RainRain(RAIN)$0.0171402.33%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$519.175.58%
  • chainlinkChainlink(LINK)$11.826.57%
  • whitebitWhiteBIT Coin(WBT)$73.994.56%
  • leo-tokenLEO Token(LEO)$9.351.18%
  • cardanoCardano(ADA)$0.2209049.08%
  • stellarStellar(XLM)$0.1843994.94%
  • bitcoin-cashBitcoin Cash(BCH)$256.435.10%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.1127313.41%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$51.453.46%
  • uniswapUniswap(UNI)$6.318.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.373.03%
  • hedera-hashgraphHedera(HBAR)$0.0792966.25%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.504.56%
  • suiSui(SUI)$0.784.43%
  • shiba-inuShiba Inu(SHIB)$0.0000054.10%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • crypto-com-chainCronos(CRO)$0.0576625.94%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,473.081.98%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.953.76%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.04-2.12%
  • okbOKB(OKB)$109.183.73%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • BittensorBittensor(TAO)$225.473.87%
  • aaveAave(AAVE)$132.654.04%
  • AsterAster(ASTER)$0.72-0.67%
  • pax-goldPAX Gold(PAXG)$4,481.981.98%
  • mantleMantle(MNT)$0.571.02%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572802.36%
  • OndoOndo(ONDO)$0.3628734.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NTU and Microsoft Researchers Propose MIMIC-IT: A Large-Scale Multi-Modal in-Context Instruction Tuning Dataset

June 15, 2023
in AI & Technology
Reading Time: 5 mins read
A A
NTU and Microsoft Researchers Propose MIMIC-IT: A Large-Scale Multi-Modal in-Context Instruction Tuning Dataset
ShareShareShareShareShare

Recent developments in artificial intelligence have concentrated on conversational assistants with great comprehension capabilities who can then act. The noteworthy successes of these conversational assistants may be ascribed to the practice of instruction adjustment in addition to the large language models’ (LLMs) high generalization capacity. It entails optimizing LLMs for a variety of activities that are described by varied and excellent instructions. By including instruction adjustment, LLMs get a deeper understanding of user intentions, improving their zero-shot performance even in newly unexplored tasks. 

Instruction tuning internalizes the context, which is desirable in user interactions, especially when user input bypasses obvious context, which may be one explanation for the zero-shot speed improvement. Conversational assistants have had amazing progress in linguistic challenges. An ideal casual assistant, however, must be able to handle jobs requiring several modalities. An extensive and top-notch multimodal instruction-following dataset is needed for this. The original vision-language instruction-following dataset is called LLaVAInstruct-150K or LLaVA. It is built utilizing COCO pictures, instructions, and data from GPT-4 based on item bounding boxes and image descriptions. 

LLaVA-Instruct-150K is inspirational, yet it has three drawbacks. (1) Limited visual diversity: Because the dataset only uses the COCO picture, its visual diversity is limited. (2) It uses a single image as visual input, but a multimodal conversational assistant should be able to handle several photos or even lengthy films. For instance, when a user asks for assistance in coming up with an album title for a set of photographs (or an image sequence, such as a video), the system needs to respond properly. (3) Language-only in-context information: While a multimodal conversational assistant should use multimodal in-context information to understand better user instructions, language-only in-context information relies entirely on language. 

🚀 JOIN the fastest ML Subreddit Community

For instance, if a human user offers a specific visual sample of the required features, an assistant can more properly align its description of an image with the tone, style, or other elements. Researchers from S-Lab, Nanyang Technological University, Singapore and Microsoft Research, Redmond provide MIMICIT (Multimodal In-Context Instruction Tuning), which addresses these restrictions. (1) Diverse visual scenes, integrating photos and videos from general scenes, egocentric view scenes, and indoor RGB-D images across different datasets, are a feature of MIMIC-IT. (2) Multiple pictures (or a video) used as visual data to support instruction-response pairings that various images or movies may accompany. (3) Multimodal in-context infor consists of in-context data presented in various instruction-response pairs, photos, or videos (for more details on data format, see Fig. 1). 

They provide Sythus, an automated pipeline for instruction-response annotation inspired by the self-instruct approach, to effectively create instruction-response pairings. Targeting the three core functions of vision-language models—perception, reasoning, and planning—Sythus uses system message, visual annotation, and in-context examples to guide the language model (GPT-4 or ChatGPT) in generating instruction-response pairs based on visual context, including timestamps, captions, and object information. Instructions and replies are also translated from English into seven other languages to allow multilingual usage. They train a multimodal model named Otter based on OpenFlamingo on MIMIC-IT. 

Figure 1: MIMIC-IT vs. LLaVA-Instruct-150K Data Format Comparison. (a) LLaVA-Instruct150K is made up of a single picture and the necessary in-context linguistic information (yellow box). (b) MIMIC-IT provides multi-modal in-context information and can accommodate several pictures or videos inside the input data, i.e., it treats both visual and linguistic inputs as in-context information.

Otter’s multimodal talents are assessed in two ways: (1) Otter performs best in the ChatGPT evaluation on the MMAGIBenchmark, which compares Otter’s perceptual and reasoning skills to other current vision-language models (VLMs). (2) Human assessment in the Multi-Modality Arena, where Otter performs better than other VLMs and receives the highest Elo score. Otter outperforms OpenFlamingo in all few-shot conditions, according to our evaluation of its few-shot in-context learning capabilities using the COCO Caption dataset.

Specifically, they provided: • The Multimodal In-Context Instruction Tuning (MIMIC-IT) dataset contains 2.8 million multimodal in-context instruction-response pairings with 2.2 million distinct instructions in various real-world settings. • Syphus, an automated process created with LLMs to produce instruction-response pairs that are high-quality and multilingual depending on visual context. • Otter, a multimodal model, exhibits skilful in-context learning and strong multimodal perception and reasoning ability, successfully following human intent.


Check Out The Paper and GitHub link. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How Apple’s New CEO Will Shape the iPhone Maker

Nvidia Deepens Chip Ties With $3.5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


➡️ Try: Criminal IP: AI-based Phishing Link Checker Chrome Extension

Credit: Source link

ShareTweetSendSharePin

Related Posts

How Apple’s New CEO Will Shape the iPhone Maker
AI & Technology

How Apple’s New CEO Will Shape the iPhone Maker

September 3, 2026
Nvidia Deepens Chip Ties With .5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026
AI & Technology

Nvidia Deepens Chip Ties With $3.5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026

September 3, 2026
Anthropic IPO Could Open AI Listings Floodgates: Madrona
AI & Technology

Anthropic IPO Could Open AI Listings Floodgates: Madrona

September 3, 2026
What Is The Best Waterproof Rating For Bluetooth Speakers?
AI & Technology

What Is The Best Waterproof Rating For Bluetooth Speakers?

September 3, 2026
Next Post
Struggling Oil Not Market Volatility Will Keep Fed on Hold

Struggling Oil Not Market Volatility Will Keep Fed on Hold

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The story behind Emeril’s iconic ‘BAM!’

The story behind Emeril’s iconic ‘BAM!’

August 31, 2026
Video shows police rescuing child involved in carjacking

Video shows police rescuing child involved in carjacking

August 30, 2026
FIFA boss Gianni Infantino cancels planned speech at swanky Hamptons bash amid legal threat

FIFA boss Gianni Infantino cancels planned speech at swanky Hamptons bash amid legal threat

August 28, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!