• bitcoinBitcoin(BTC)$78,308.00-1.09%
  • ethereumEthereum(ETH)$2,476.59-1.32%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$721.92-4.33%
  • rippleXRP(XRP)$1.39-3.30%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.89-2.38%
  • tronTRON(TRX)$0.3402540.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,245.06-0.17%
  • HyperliquidHyperliquid(HYPE)$84.19-2.24%
  • dogecoinDogecoin(DOGE)$0.085767-5.37%
  • RainRain(RAIN)$0.0163861.98%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$512.381.64%
  • whitebitWhiteBIT Coin(WBT)$80.83-1.44%
  • chainlinkChainlink(LINK)$11.81-6.11%
  • leo-tokenLEO Token(LEO)$9.190.07%
  • cardanoCardano(ADA)$0.213162-3.47%
  • stellarStellar(XLM)$0.179947-5.24%
  • bitcoin-cashBitcoin Cash(BCH)$250.37-3.85%
  • daiDai(DAI)$1.00-0.02%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.103783-4.89%
  • litecoinLitecoin(LTC)$52.70-3.18%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-1.31%
  • uniswapUniswap(UNI)$6.07-12.13%
  • avalanche-2Avalanche(AVAX)$7.83-2.51%
  • hedera-hashgraphHedera(HBAR)$0.076855-3.27%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.527.78%
  • suiSui(SUI)$0.77-6.42%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.03%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.057850-4.37%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.222.97%
  • tether-goldTether Gold(XAUT)$4,403.150.52%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$256.18-1.22%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.42-1.18%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.04%
  • mantleMantle(MNT)$0.60-5.58%
  • AsterAster(ASTER)$0.73-4.46%
  • aaveAave(AAVE)$125.01-3.89%
  • pax-goldPAX Gold(PAXG)$4,407.340.56%
  • polkadotPolkadot(DOT)$1.11-9.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0566431.09%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CMU Researchers Introduce FROMAGe: An AI Model That Efficiently Bootstraps Frozen Large Language Models (LLMs) To Generate Free-Form Text Interleaved With Images

July 3, 2023
in AI & Technology
Reading Time: 5 mins read
A A
CMU Researchers Introduce FROMAGe: An AI Model That Efficiently Bootstraps Frozen Large Language Models (LLMs) To Generate Free-Form Text Interleaved With Images
ShareShareShareShareShare

Enormous large language models (LLMs) can exhibit appealing skills like producing human-like discourse and responding to complicated inquiries because they have been trained at scale on large text corpora. While undoubtedly amazing, most cutting-edge LLMs are trained on text-only data downloaded from the Internet. They frequently cannot absorb concepts based on the actual world because they need to be exposed to rich visual clues. As a result, most language models now in use show limits on tasks that need visual reasoning and grounding and are also unable to generate visuals. In this article, they demonstrate how to effectively use a frozen LLM’s capabilities for multimodal (picture and text) input and output.

They train the language model to learn a new [RET] token that stands in for an image for image-text retrieval. They also know linear mapping using contrastive learning to map the [RET] embeddings for a caption to be close to the visual embeddings for its associated picture. Only the weights of the linear layers and the [RET] token embedding are updated during training, with most of the model remaining frozen. As a result, their suggested approach is highly memory and computationally efficient. Once trained, a model demonstrates several skills. It has a new multimodal conversation and reasoning skills in addition to the original text-only LLM’s ability to create text. Their suggested approach is model-independent and may be used to base future releases of stronger or bigger LLMs.

The language model is trained to learn a new [RET] token representing an image, and contrastive learning is used to know a linear mapping that maps the [RET] embeddings for a caption to be close to the visual embeddings for its matched picture. Only the weights of the linear layers and the [RET] token embedding are updated during training, leaving most of the model fixed. As a result, their suggested approach is highly memory and computationally efficient. 1Once taught, their model demonstrates several skills. It has a new multimodal conversation and reasoning skills in addition to the original text-only LLM’s ability to create text. Their suggested approach is model-independent and may be used to base future releases of stronger or bigger LLMs.

🔥 Join The Fastest Growing ML Subreddit

Showcasing the increased sensitivity of text-to-image retrieval performed by autoregressive LLMs. One of their primary contributions is the Frozen Retrieval Over Multimodal Data for Autoregressive Generation (FROMAGe) model, effectively trained by visually anchoring LLMs through picture captioning and contrastive learning. While previous algorithms require webscale interleaved image-text data, FROMAGe develops strong few-shot multimodal capabilities from image caption pairings alone. Their method is more accurate on lengthy and complicated free-form text than previous models. Demonstrating how pretrained text-only LLMs’ current skills, including in-context learning, input sensitivity, and conversation creation, may be used for tasks that require visual input. 

They show: (1) contextual image retrieval from sequences of interspersed pictures and text; (2) good zero-shot performance on visual conversation; and (3) enhanced discourse context sensitivity for image retrieval. Their results open the door to models that can learn from and produce lengthy, coherent multimodal sequences. They also highlight the capabilities of pretrained text-only LLMs on visually based tasks. To promote more research and development, their code and pretrained models will be made available to the general public soon.

Using an approach, a language model is grounded in the visual domain and is able to handle arbitrarily interspersed image-text inputs and produce coherent image-text outputs. Green speech bubbles are created by the model, whereas grey speech bubbles represent input prompts.

Check out the Paper, Project, and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 13k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Next Post
Shares of IBM Slip After Reporting Latest Earnings Results

Shares of IBM Slip After Reporting Latest Earnings Results

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Grocers file antitrust suit to block Mamdami admin’s plan for NYC-owned supermarkets: ‘Predatory pricing’

Grocers file antitrust suit to block Mamdami admin’s plan for NYC-owned supermarkets: ‘Predatory pricing’

September 9, 2026
A Robo-Advisor Is the Lowest-Friction Way to Start Investing

A Robo-Advisor Is the Lowest-Friction Way to Start Investing

September 7, 2026
Cyclist becomes first person to ride on top of a hot-air balloon

Cyclist becomes first person to ride on top of a hot-air balloon

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!