• bitcoinBitcoin(BTC)$79,949.000.11%
  • ethereumEthereum(ETH)$2,510.681.22%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$753.09-2.25%
  • rippleXRP(XRP)$1.420.36%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$106.082.27%
  • tronTRON(TRX)$0.3360730.42%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.065.05%
  • zcashZcash(ZEC)$1,230.2021.50%
  • HyperliquidHyperliquid(HYPE)$87.972.84%
  • dogecoinDogecoin(DOGE)$0.090406-0.88%
  • RainRain(RAIN)$0.016784-1.43%
  • moneroMonero(XMR)$531.00-3.22%
  • USDSUSDS(USDS)$1.000.03%
  • chainlinkChainlink(LINK)$13.118.58%
  • whitebitWhiteBIT Coin(WBT)$73.750.37%
  • leo-tokenLEO Token(LEO)$9.331.19%
  • cardanoCardano(ADA)$0.2218140.85%
  • stellarStellar(XLM)$0.1861780.63%
  • bitcoin-cashBitcoin Cash(BCH)$259.530.91%
  • daiDai(DAI)$1.000.00%
  • uniswapUniswap(UNI)$7.140.11%
  • CantonCanton(CC)$0.1107501.49%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • litecoinLitecoin(LTC)$54.951.51%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.05%
  • hedera-hashgraphHedera(HBAR)$0.0816700.78%
  • avalanche-2Avalanche(AVAX)$7.832.96%
  • suiSui(SUI)$0.811.24%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.45%
  • nearNEAR Protocol(NEAR)$2.4211.70%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0577152.11%
  • tether-goldTether Gold(XAUT)$4,419.30-0.17%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BittensorBittensor(TAO)$269.6514.96%
  • MemeCoreMemeCore(M)$1.12-0.02%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.630.54%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • AsterAster(ASTER)$0.78-0.62%
  • aaveAave(AAVE)$134.920.40%
  • mantleMantle(MNT)$0.602.90%
  • pax-goldPAX Gold(PAXG)$4,423.08-0.22%
  • OndoOndo(ONDO)$0.3849813.80%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056638-1.31%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Make ChatGPT See Again: This AI Approach Explores Link-Context Learning to Enable Multimodal Learning

September 7, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Make ChatGPT See Again: This AI Approach Explores Link-Context Learning to Enable Multimodal Learning
ShareShareShareShareShare

Language models have revolutionized the way we communicate with computers by their ability to generate coherent and contextually relevant text. Large Language Models (LLMs) have been at the forefront of this progress, trained on massive amounts of text data to learn the patterns and nuances of human language. ChatGPT, the pioneer of the LLM revolution, is extremely popular among people in different disciplines.

LLMs have made various tasks easier to tackle thanks to their extreme ability. We use them to summarize texts, help us write emails, automate coding tasks, explain documents, etc. All these tasks were quite time-consuming just a year ago, but nowadays, they take just a couple of minutes to complete.

However, with the increasing demand for multimodal understanding, where models need to process and generate content across different modalities like text, images, and even videos, the need for Multimodal Large Language Models (MLLMs) has emerged. MLLMs combine the power of language models with visual understanding, enabling machines to comprehend and generate content in a more comprehensive and contextually aware manner.

Once the ChatGPT craze settled down a bit, MLLMs took the AI world by storm, enabling machines to understand and generate content across different modalities like text and images. These models have shown remarkable performance in tasks like image recognition, visual grounding, and instruction understanding. However, training these models effectively remains a challenge. The biggest challenge is when an MLLM encounters entirely novel scenarios where both the image and the label are unseen.

Moreover, MLLMs tend to get “lost in the middle” when processing longer contexts. These models heavily rely on the beginning and middle positions, which explains the plateau in accuracy as the number of shots increases. Therefore, MLLMs struggle with longer inputs.

Time to meet Link-context-learning (LCL) that tackles various challenges in MLLM.

In MLLM, there are two key training strategies. Multimodal Prompt Tuning (M-PT) and Multimodal Instruction Tuning (M-IT). M-PT involves fine-tuning only a small portion of the model’s parameters while keeping the rest frozen. This approach helps achieve similar results to full fine-tuning while minimizing computational resources. On the other hand, M-IT enhances the zero-shot capability of MLLMs by fine-tuning them on datasets that include instruction descriptions. This strategy improves the model’s ability to understand and respond to new tasks without prior training. These work fine, but they both sacrifice certain aspects. 

Instead, LCL explores different training strategies: mix strategy, 2-way strategy, 2-way-random, and 2-way-weight. The mixed strategy stands out by significantly boosting zero-shot accuracy and achieving impressive results at 6-shot. However, its performance slightly decreases at 16-shot. On the contrary, the 2-way strategy shows a gradual increase in accuracy from 2-shot to 16-shot, indicating a closer alignment with the trained pattern.

Unlike traditional in-context learning, LCL goes a step further by empowering the model to establish a mapping between the source and target, enhancing its overall performance. By providing demonstrations with causal links, LCL enables MLLMs to discern not only analogies but also the underlying causal associations between data points, allowing them to recognize unseen images and understand novel concepts more effectively. The ISEKAI dataset serves as a crucial resource for evaluating and advancing the capabilities of MLLMs in the context of link-context learning.

Moreover, LCL introduces the ISEKAI dataset, a novel and comprehensive dataset specifically designed to evaluate the capabilities of MLLMs. The ISEKAI dataset comprises entirely generated images and fabricated concepts. It challenges MLLMs to assimilate new concepts from ongoing conversations and retain this knowledge for accurate question-answering. 

In conclusion, LCL provides valuable insights into the training strategies employed for multimodal language models. The mixed strategy and 2-way strategy offer different approaches to enhance the performance of MLLMs, each with its own strengths and limitations. The contextual analysis sheds light on the challenges faced by MLLMs when processing longer inputs, emphasizing the importance of further research in this area. 


Check out the Paper and Code. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

How To Send High-Quality Images And Videos From Android To iPhone

Ekrem Çetinkaya received his B.Sc. in 2018, and M.Sc. in 2019 from Ozyegin University, Istanbul, Türkiye. He wrote his M.Sc. thesis about image denoising using deep convolutional networks. He received his Ph.D. degree in 2023 from the University of Klagenfurt, Austria, with his dissertation titled “Video Coding Enhancements for HTTP Adaptive Streaming Using Machine Learning.” His research interests include deep learning, computer vision, video encoding, and multimedia networking.


🚀 Check out Hostinger AI Website Builder (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
AI & Technology

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

September 6, 2026
How To Send High-Quality Images And Videos From Android To iPhone
AI & Technology

How To Send High-Quality Images And Videos From Android To iPhone

September 6, 2026
What Is Vibe Coding And Why Does It Get So Much Hate?
AI & Technology

What Is Vibe Coding And Why Does It Get So Much Hate?

September 6, 2026
My Content Tracker Idea Became a Real App – Unite.AI
AI & Technology

My Content Tracker Idea Became a Real App – Unite.AI

September 6, 2026
Next Post
Tech IPOs To Watch

Tech IPOs To Watch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
JapanFold Keeps Open-Source Drug Discovery Computations Inside Japan – Unite.AI

JapanFold Keeps Open-Source Drug Discovery Computations Inside Japan – Unite.AI

September 3, 2026
Extended Interview: Gov. Kathy Hochul on N.Y.’s Data Center Moratorium

Extended Interview: Gov. Kathy Hochul on N.Y.’s Data Center Moratorium

September 6, 2026
HPE CEO Neri on Oracle Deal, AI Adoption and Earnings Outlook

HPE CEO Neri on Oracle Deal, AI Adoption and Earnings Outlook

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!