• bitcoinBitcoin(BTC)$80,977.00-0.27%
  • ethereumEthereum(ETH)$2,623.66-0.37%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$758.68-0.73%
  • rippleXRP(XRP)$1.410.91%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$110.68-2.45%
  • tronTRON(TRX)$0.3392350.21%
  • zcashZcash(ZEC)$1,476.61-0.89%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.21%
  • HyperliquidHyperliquid(HYPE)$91.56-0.10%
  • dogecoinDogecoin(DOGE)$0.087737-0.41%
  • moneroMonero(XMR)$537.96-6.26%
  • whitebitWhiteBIT Coin(WBT)$82.66-0.84%
  • RainRain(RAIN)$0.0137842.24%
  • USDSUSDS(USDS)$1.00-0.03%
  • chainlinkChainlink(LINK)$12.38-0.14%
  • cardanoCardano(ADA)$0.2271971.36%
  • leo-tokenLEO Token(LEO)$8.90-0.04%
  • stellarStellar(XLM)$0.1955250.85%
  • uniswapUniswap(UNI)$8.52-5.99%
  • bitcoin-cashBitcoin Cash(BCH)$251.41-1.96%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • nearNEAR Protocol(NEAR)$3.54-3.66%
  • daiDai(DAI)$1.00-0.01%
  • litecoinLitecoin(LTC)$57.260.03%
  • USD1USD1(USD1)$1.00-0.03%
  • CantonCanton(CC)$0.109598-1.09%
  • avalanche-2Avalanche(AVAX)$9.7117.36%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.380.38%
  • hedera-hashgraphHedera(HBAR)$0.0806331.35%
  • suiSui(SUI)$0.864.89%
  • MemeCoreMemeCore(M)$1.5013.60%
  • Global DollarGlobal Dollar(USDG)$1.00-0.03%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.25%
  • BittensorBittensor(TAO)$261.875.02%
  • crypto-com-chainCronos(CRO)$0.059155-1.26%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • tether-goldTether Gold(XAUT)$4,371.97-0.10%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$118.091.14%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.02%
  • aaveAave(AAVE)$140.911.00%
  • AsterAster(ASTER)$0.770.87%
  • mantleMantle(MNT)$0.62-0.23%
  • OndoOndo(ONDO)$0.4164794.37%
  • EthenaEthena(ENA)$0.20017017.27%
  • Pump.funPump.fun(PUMP)$0.004158-5.33%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Stanford University Evaluates the Performance of Multimodal Foundation Models Scaling from Few-Shot to Many-Shot-In-Context Learning ICL

May 19, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper from Stanford University Evaluates the Performance of Multimodal Foundation Models Scaling from Few-Shot to Many-Shot-In-Context Learning ICL
ShareShareShareShareShare

Incorporating demonstrating examples, known as in-context learning (ICL), significantly enhances large language models (LLMs) and large multimodal models (LMMs) without requiring parameter updates. Recent studies confirm the efficacy of few-shot multimodal ICL, particularly in improving LMM performance on out-of-domain tasks. With longer context windows in advanced models like GPT-4o and Gemini 1.5 Pro, researchers can now investigate the impact of increasing demonstrating examples, a factor previously constrained by context window limitations.

Some researchers observed enhanced performance in LLMs with increased in-context examples, albeit constrained by context size. Recent studies extended this exploration, demonstrating improvements with over 1,000 examples, besides in text-only benchmarks. Multimodal ICL research remains emerging, with studies showing benefits for models like GPT-4V and Gemini in out-domain tasks. Batch querying strategies offer efficiency gains in inference, with recent variations proposed to optimize performance, utilizing larger context windows in recent models.

To examine the potential of advanced multimodal foundation models in many-shot ICL, researchers from Stanford execute an extensive array of experiments to assess model efficacy across 10 datasets covering various domains and image classification tasks. This involves significantly increasing the number of demonstrating examples to gauge model performance.

The Key findings of this study include:

1. Increased demonstrating examples significantly enhance model performance, with Gemini 1.5 Pro showing consistent log-linear improvements compared to GPT-4o.

2. Gemini 1.5 Pro demonstrates higher ICL data efficiency compared to GPT-4o across most datasets.

3. Combining multiple queries into a single request can deliver comparable or superior performance to individual queries in a many-shot scenario. This approach also reduces per-example latency significantly and offers a more cost-effective inference process.

4. Batched questioning notably enhances performance in zero-shot scenarios, attributed to domain and class calibrated and self-generated demonstrating examples through autoregressive decoding.

Three advanced multimodal foundation models—GPT-4o, GPT4(V)-Turbo, and Gemini 1.5 Pro—are employed, with GPT-4o and Gemini 1.5 Pro emphasized due to superior performance. Claude3-Opus is excluded from experiments due to its 20-image limit per request. Each model is accessed through specific endpoints, with OpenAI’s API service for GPT-4o and GPT-4(V)-Turbo, and Google Cloud’s Vertex AI for Gemini 1.5 Pro. Zero temperature is set for all models, and a random seed ensures deterministic responses. Sampling strategies ensure class balance in demonstration and test sets across 10 datasets spanning various domains and classification tasks, with demonstration examples scaled up while maintaining balance for evaluation.

Gemini 1.5 Pro consistently demonstrates significant performance enhancements across most datasets as demonstrating examples increase, except for DrugOOD Assay. Particularly significant improvements are observed in HAM10000 (+23% accuracy compared to zero-shot), FIVES (+29% accuracy), and EuroSAT (+38% accuracy).  for 5 out of the 10 datasets (FIVES, UCMerced, EuroSAT, Oxford Pets, and

DTD), Gemini 1.5 Pro performance continues to improve up to the highest number of demonstrating

examples considered (~1,000 examples). Conversely, GPT-4o exhibits performance improvements on most datasets but with less consistency, showing V-shaped scaling curves on many datasets. GPT-4o’s performance on DrugOOD Assay also displays high variance, similar to Gemini 1.5 Pro, with peak performance at 50 demo examples.

To recapitulate, This study assesses many-shot ICL of state-of-the-art multimodal foundation models across 10 datasets, revealing consistent performance enhancements. Batching queries with many-shot ICL significantly reduces per-example latency and inference costs without sacrificing performance. These findings suggest the potential of utilizing large numbers of demonstrating examples to adapt models quickly to new tasks and domains, circumventing the need for traditional fine-tuning. Future research should investigate the comparative effectiveness and data efficiency of traditional fine-tuning versus many-shot ICL. Also, examining issues like hallucinations and biases in the context of many-shot ICL and batched queries is crucial for model refinement and mitigating biases across diverse sub-groups.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

SpaceX Targets September 28 For Starship’s First Orbital Flight

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text
AI & Technology

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

September 19, 2026
Why Is Your iPad Not Charging (And How To Fix It)
AI & Technology

Why Is Your iPad Not Charging (And How To Fix It)

September 19, 2026
How To Block And Unblock A Number On Your Android Phone
AI & Technology

How To Block And Unblock A Number On Your Android Phone

September 19, 2026
Next Post
Video shows Israeli forces in disguise inside a West Bank hospital

Video shows Israeli forces in disguise inside a West Bank hospital

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Canon EOS R8 Mark II Announced: Our thoughts and reaction – Canon Rumors

Canon EOS R8 Mark II Announced: Our thoughts and reaction – Canon Rumors

September 16, 2026
Lindsay Clancy lawyer, jurors speak out after mistrial

Lindsay Clancy lawyer, jurors speak out after mistrial

September 14, 2026
Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!