• bitcoinBitcoin(BTC)$81,181.005.21%
  • ethereumEthereum(ETH)$2,497.034.73%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$724.195.48%
  • rippleXRP(XRP)$1.457.86%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.934.08%
  • tronTRON(TRX)$0.3306051.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.79%
  • HyperliquidHyperliquid(HYPE)$86.936.74%
  • zcashZcash(ZEC)$947.8616.48%
  • dogecoinDogecoin(DOGE)$0.0876007.75%
  • RainRain(RAIN)$0.0171342.44%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$520.203.64%
  • chainlinkChainlink(LINK)$11.796.39%
  • whitebitWhiteBIT Coin(WBT)$73.964.61%
  • leo-tokenLEO Token(LEO)$9.351.17%
  • cardanoCardano(ADA)$0.22070810.26%
  • stellarStellar(XLM)$0.1838105.22%
  • bitcoin-cashBitcoin Cash(BCH)$257.015.54%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.1126812.09%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$51.253.37%
  • uniswapUniswap(UNI)$6.338.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.373.13%
  • hedera-hashgraphHedera(HBAR)$0.0789876.18%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.504.67%
  • suiSui(SUI)$0.784.95%
  • shiba-inuShiba Inu(SHIB)$0.0000054.33%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0577896.62%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,465.541.92%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.964.33%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.04-1.69%
  • okbOKB(OKB)$109.014.43%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.44%
  • BittensorBittensor(TAO)$225.314.15%
  • aaveAave(AAVE)$132.674.61%
  • AsterAster(ASTER)$0.72-0.83%
  • pax-goldPAX Gold(PAXG)$4,474.321.92%
  • mantleMantle(MNT)$0.571.71%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572942.40%
  • OndoOndo(ONDO)$0.3624394.93%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NTU and Microsoft Researchers Propose MIMIC-IT: A Large-Scale Multi-Modal in-Context Instruction Tuning Dataset

June 15, 2023
in AI & Technology
Reading Time: 5 mins read
A A
NTU and Microsoft Researchers Propose MIMIC-IT: A Large-Scale Multi-Modal in-Context Instruction Tuning Dataset
ShareShareShareShareShare

Recent developments in artificial intelligence have concentrated on conversational assistants with great comprehension capabilities who can then act. The noteworthy successes of these conversational assistants may be ascribed to the practice of instruction adjustment in addition to the large language models’ (LLMs) high generalization capacity. It entails optimizing LLMs for a variety of activities that are described by varied and excellent instructions. By including instruction adjustment, LLMs get a deeper understanding of user intentions, improving their zero-shot performance even in newly unexplored tasks. 

Instruction tuning internalizes the context, which is desirable in user interactions, especially when user input bypasses obvious context, which may be one explanation for the zero-shot speed improvement. Conversational assistants have had amazing progress in linguistic challenges. An ideal casual assistant, however, must be able to handle jobs requiring several modalities. An extensive and top-notch multimodal instruction-following dataset is needed for this. The original vision-language instruction-following dataset is called LLaVAInstruct-150K or LLaVA. It is built utilizing COCO pictures, instructions, and data from GPT-4 based on item bounding boxes and image descriptions. 

LLaVA-Instruct-150K is inspirational, yet it has three drawbacks. (1) Limited visual diversity: Because the dataset only uses the COCO picture, its visual diversity is limited. (2) It uses a single image as visual input, but a multimodal conversational assistant should be able to handle several photos or even lengthy films. For instance, when a user asks for assistance in coming up with an album title for a set of photographs (or an image sequence, such as a video), the system needs to respond properly. (3) Language-only in-context information: While a multimodal conversational assistant should use multimodal in-context information to understand better user instructions, language-only in-context information relies entirely on language. 

🚀 JOIN the fastest ML Subreddit Community

For instance, if a human user offers a specific visual sample of the required features, an assistant can more properly align its description of an image with the tone, style, or other elements. Researchers from S-Lab, Nanyang Technological University, Singapore and Microsoft Research, Redmond provide MIMICIT (Multimodal In-Context Instruction Tuning), which addresses these restrictions. (1) Diverse visual scenes, integrating photos and videos from general scenes, egocentric view scenes, and indoor RGB-D images across different datasets, are a feature of MIMIC-IT. (2) Multiple pictures (or a video) used as visual data to support instruction-response pairings that various images or movies may accompany. (3) Multimodal in-context infor consists of in-context data presented in various instruction-response pairs, photos, or videos (for more details on data format, see Fig. 1). 

They provide Sythus, an automated pipeline for instruction-response annotation inspired by the self-instruct approach, to effectively create instruction-response pairings. Targeting the three core functions of vision-language models—perception, reasoning, and planning—Sythus uses system message, visual annotation, and in-context examples to guide the language model (GPT-4 or ChatGPT) in generating instruction-response pairs based on visual context, including timestamps, captions, and object information. Instructions and replies are also translated from English into seven other languages to allow multilingual usage. They train a multimodal model named Otter based on OpenFlamingo on MIMIC-IT. 

Figure 1: MIMIC-IT vs. LLaVA-Instruct-150K Data Format Comparison. (a) LLaVA-Instruct150K is made up of a single picture and the necessary in-context linguistic information (yellow box). (b) MIMIC-IT provides multi-modal in-context information and can accommodate several pictures or videos inside the input data, i.e., it treats both visual and linguistic inputs as in-context information.

Otter’s multimodal talents are assessed in two ways: (1) Otter performs best in the ChatGPT evaluation on the MMAGIBenchmark, which compares Otter’s perceptual and reasoning skills to other current vision-language models (VLMs). (2) Human assessment in the Multi-Modality Arena, where Otter performs better than other VLMs and receives the highest Elo score. Otter outperforms OpenFlamingo in all few-shot conditions, according to our evaluation of its few-shot in-context learning capabilities using the COCO Caption dataset.

Specifically, they provided: • The Multimodal In-Context Instruction Tuning (MIMIC-IT) dataset contains 2.8 million multimodal in-context instruction-response pairings with 2.2 million distinct instructions in various real-world settings. • Syphus, an automated process created with LLMs to produce instruction-response pairs that are high-quality and multilingual depending on visual context. • Otter, a multimodal model, exhibits skilful in-context learning and strong multimodal perception and reasoning ability, successfully following human intent.


Check Out The Paper and GitHub link. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Anthropic IPO Could Open AI Listings Floodgates: Madrona

Why the US Military Is Betting on Portable Nuclear

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


➡️ Try: Criminal IP: AI-based Phishing Link Checker Chrome Extension

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic IPO Could Open AI Listings Floodgates: Madrona
AI & Technology

Anthropic IPO Could Open AI Listings Floodgates: Madrona

September 3, 2026
Why the US Military Is Betting on Portable Nuclear
AI & Technology

Why the US Military Is Betting on Portable Nuclear

September 3, 2026
AI Boom Fuels a New Tech Debt Binge
AI & Technology

AI Boom Fuels a New Tech Debt Binge

September 3, 2026
Ternus Takes Apple Helm Ahead of Foldable iPhone Debut
AI & Technology

Ternus Takes Apple Helm Ahead of Foldable iPhone Debut

September 3, 2026
Next Post
Struggling Oil Not Market Volatility Will Keep Fed on Hold

Struggling Oil Not Market Volatility Will Keep Fed on Hold

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
SBAR: 85% In A Money Market ETF, But That’s Not Where The 12% Yield Comes From (NYSEARCA:SBAR)

SBAR: 85% In A Money Market ETF, But That’s Not Where The 12% Yield Comes From (NYSEARCA:SBAR)

August 29, 2026
Two dead after Greek firefighting helicopters collide

Two dead after Greek firefighting helicopters collide

August 30, 2026
Full special: News NOW coverage of midterm primaries in Michigan, Missouri and more

Full special: News NOW coverage of midterm primaries in Michigan, Missouri and more

August 29, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!