• bitcoinBitcoin(BTC)$79,076.000.75%
  • ethereumEthereum(ETH)$2,505.870.93%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$744.33-0.81%
  • rippleXRP(XRP)$1.431.04%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.770.51%
  • tronTRON(TRX)$0.338685-0.33%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.22%
  • zcashZcash(ZEC)$1,281.618.67%
  • HyperliquidHyperliquid(HYPE)$86.213.15%
  • dogecoinDogecoin(DOGE)$0.0901810.58%
  • RainRain(RAIN)$0.016422-1.56%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$517.193.29%
  • whitebitWhiteBIT Coin(WBT)$81.762.54%
  • chainlinkChainlink(LINK)$12.10-3.81%
  • leo-tokenLEO Token(LEO)$9.19-0.18%
  • cardanoCardano(ADA)$0.218766-0.87%
  • stellarStellar(XLM)$0.187920-1.15%
  • bitcoin-cashBitcoin Cash(BCH)$258.770.74%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.31-0.41%
  • CantonCanton(CC)$0.1052670.42%
  • uniswapUniswap(UNI)$6.64-4.12%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-0.54%
  • nearNEAR Protocol(NEAR)$2.6412.93%
  • hedera-hashgraphHedera(HBAR)$0.078587-1.85%
  • avalanche-2Avalanche(AVAX)$7.96-0.29%
  • suiSui(SUI)$0.81-0.37%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.37%
  • crypto-com-chainCronos(CRO)$0.059779-1.33%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,416.990.69%
  • MemeCoreMemeCore(M)$1.180.28%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$263.242.27%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.94-0.25%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • mantleMantle(MNT)$0.642.55%
  • AsterAster(ASTER)$0.75-1.34%
  • aaveAave(AAVE)$130.150.79%
  • Pump.funPump.fun(PUMP)$0.0046745.45%
  • polkadotPolkadot(DOT)$1.133.17%
  • pax-goldPAX Gold(PAXG)$4,420.480.67%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet Prismer: An Open Source Vision-Language Model with An Ensemble of Experts

July 22, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Meet Prismer: An Open Source Vision-Language Model with An Ensemble of Experts
ShareShareShareShareShare

Several recent vision-language models have demonstrated remarkable multi-modal generation abilities. But typically, they call for training enormous models on enormous datasets. Researchers introduce Prismer, a data- and parameter-efficient vision-language model that uses an ensemble of domain experts, as a scalable alternative. By inheriting most of the network weights from publicly available, pre-trained domain experts and freezing them during training, Prismer only requires training a few components.

The generalization abilities of large pre-trained models are exceptional across many different tasks. However, these features come at a high price, necessitating a lot of training data and computational resources for training and inference. Models with hundreds of billions of trainable parameters are common in the language domain, and they typically necessitate a computing budget on the yottaFLOP scale.

Issues related to visual language learning are more difficult to solve. Even though this field is a superset of language processing, it also necessitates visual and multi-modal thinking expertise. Using its projected multi-modal signals, Prismer is a data-efficient vision-language model that uses a wide range of pre-trained experts. It can handle visual question answering and picture captioning, two examples of vision-language reasoning tasks. Using a prism as an example, Prismer divides a general reasoning job into several smaller, more manageable chunks.

🚀 Build high-quality training datasets with Kili Technology and solve NLP machine learning challenges to develop powerful ML applications

Researchers developed a visually conditioned autoregressive text generation model toTwo of Prismer’s most important design features are I vision-only. Language-only models for web-scale knowledge to construct our core network backbones, and (ii) modalities-specific vision experts encoding multiple types of visual information, from low-level vision signals like depth to high-level vision signals like instance and semantic labels, as auxiliary knowledge, directly from their corresponding network outputs. Researchers developed a visually conditioned autoregressive text generation model to better use various pre-trained domain experts for exploratory vision-language reasoning tasks.

Even though Prismer was only trained on 13M examples of publicly available image/alt-text data, it shows strong multi-modal reasoning performance in tasks like image captioning, image classification, and visual question answering, which is competitive with many state-of-the-art vision language models. Researchers conclude with a thorough investigation of Prismer’s learning habits, where researchers find several good features.

Model Design:

The Prismer model, shown in its encoder-decoder transformer version, draws on a large pool of already-trained subject matter experts to speed up the training process. A visual encoder plus an autoregressive language decoder make up this system. The vision encoder receives a sequence of RGB and multi-modal labels (depth, surface normal, and segmentation labels anticipated from the frozen pre-trained experts) as input. It produces a sequence of RGB and multi-modal features as output. As a result of this cross-attention training, the language decoder is conditioned to generate a string of text tokens.

Advantages:

  • The Prismer model has several benefits, but one of the most notable is that it uses data extremely efficiently while being trained. Prismer is constructed on top of pre-trained vision-only and language-only backbone models to achieve this goal with a considerable decrease in GPU hours necessary to attain equivalent performance to other state-of-the-art vision-language models. One may use these pre-trained parameters to use the massive amounts of available web-scale knowledge.
  • Researchers also developed a multi-modal signal input for the vision encoder. The created multi-modal auxiliary knowledge can better capture semantics and information about the input image. Prismer’s architecture is optimized for maximizing the use of trained experts with few trainable parameters.

Researchers have included two varieties of pre-trained specialists in Prismer:

  1. Specialists in the Backbone The pre-trained models responsible for translating text and pictures into a meaningful sequence of tokens are called “vision-only” and “language-only” models, respectively.
  2. Depending on the data used in their training, moderators of Discourse Models may label tasks in various ways.

Properties

  • The more knowledgeable people there are, the better the results. As the number of modality specialists in Prismer grows, its performance enhances.
  • More Skilled Professionals, Higher Results researchers replace some fraction of the predicted depth labels with random noise taken from a Uniform Distribution to create a corrupted depth expert and assess the effect of expert quality on Prismer’s performance.
  • Resistance to Unhelpful Opinions the findings further demonstrate that Prismer’s performance is steady when noise-predicting experts are incorporated.

Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

Lyft Is Now Offering Waymo Rides In Nashville

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🔥 Gain a competitive
edge with data: Actionable market intelligence for global brands, retailers, analysts, and investors. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Lyft Is Now Offering Waymo Rides In Nashville
AI & Technology

Lyft Is Now Offering Waymo Rides In Nashville

September 9, 2026
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Next Post
Zillow’s Economist Explains Why March Is the Best Month to Sell a Home

Zillow's Economist Explains Why March Is the Best Month to Sell a Home

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
5 dead, 5 injured after Amazon cargo plane overruns runway at MIA, sheriff says – NBC 6 South Florida

5 dead, 5 injured after Amazon cargo plane overruns runway at MIA, sheriff says – NBC 6 South Florida

September 7, 2026
How To Send High-Quality Images And Videos From Android To iPhone

How To Send High-Quality Images And Videos From Android To iPhone

September 6, 2026
Miley, Beyoncé, ADÉLA & More: New Music Friday Guide – Billboard

Miley, Beyoncé, ADÉLA & More: New Music Friday Guide – Billboard

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!