• bitcoinBitcoin(BTC)$76,911.00-1.40%
  • ethereumEthereum(ETH)$2,451.72-0.37%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$712.65-1.13%
  • rippleXRP(XRP)$1.34-2.68%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$99.29-1.71%
  • tronTRON(TRX)$0.3405540.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.98%
  • zcashZcash(ZEC)$1,078.14-11.39%
  • HyperliquidHyperliquid(HYPE)$78.93-5.26%
  • dogecoinDogecoin(DOGE)$0.083448-2.30%
  • RainRain(RAIN)$0.015723-2.72%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$507.76-0.61%
  • whitebitWhiteBIT Coin(WBT)$79.62-1.10%
  • chainlinkChainlink(LINK)$11.49-2.22%
  • leo-tokenLEO Token(LEO)$9.10-1.09%
  • cardanoCardano(ADA)$0.206702-1.82%
  • stellarStellar(XLM)$0.175651-1.78%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.44-9.39%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.750.15%
  • CantonCanton(CC)$0.097448-5.48%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-0.81%
  • uniswapUniswap(UNI)$6.010.57%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.075197-1.41%
  • avalanche-2Avalanche(AVAX)$7.46-3.58%
  • nearNEAR Protocol(NEAR)$2.42-1.26%
  • suiSui(SUI)$0.73-3.64%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.64%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056351-2.69%
  • tether-goldTether Gold(XAUT)$4,331.64-1.67%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.14-5.40%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$109.06-3.21%
  • BittensorBittensor(TAO)$234.57-6.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • polkadotPolkadot(DOT)$1.131.97%
  • AsterAster(ASTER)$0.70-2.38%
  • mantleMantle(MNT)$0.57-4.07%
  • aaveAave(AAVE)$121.69-1.97%
  • pax-goldPAX Gold(PAXG)$4,337.81-1.63%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0568050.70%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Contextual AI Introduces LENS: An AI Framework for Vision-Augmented Language Models that Outperforms Flamingo by 9% (56->65%) on VQAv2

July 1, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Contextual AI Introduces LENS: An AI Framework for Vision-Augmented Language Models that Outperforms Flamingo by 9% (56->65%) on VQAv2
ShareShareShareShareShare

Large Language Models (LLMs) have transformed natural language understanding in recent years, demonstrating remarkable aptitudes in semantic comprehension, query resolution, and text production, particularly in zero-shot and few-shot environments. As seen in Fig. 1(a), several methods have been put forth for using LLMs on tasks involving vision. An optical encoder may be trained to represent each picture as a series of continuous embeddings, allowing the LLM to understand it. Another uses a contrastively trained frozen vision encoder while adding additional layers to the frozen LLM that are then learned from scratch. 

Another method recommends training a lightweight transformer to align a frozen visual encoder (pre-trained contrastively) and a frozen LLM. Even if they have progressed in the abovementioned research, it is still difficult to justify the additional pretraining stage(s)’ computational cost. In addition, massive databases, including text, photos, and videos, are required to synchronize the visual and linguistic modalities with an existing LLM. Flamingo adds new cross-attention layers into an LLM pre-trained to add visual features. 

Figure 1: Comparing methods for coordinating visual and linguistic modalities There are two options for multimodal pretraining: (a) utilising a paired or web dataset; and (b) LENS, a pretraining-free technique that can be used with any off-the-shelf LLM without the requirement for extra multimodal datasets. Unlike LENS, previous approaches require joint alignment pretraining on substantial multimodal datasets in order to accomplish visual tasks.

The multimodal pretraining stage requires stunning 2 billion picture-text pairs and 43 million websites, which can take up to 15 days, even employing a pretrained image encoder and a pretrained frozen LLM. Instead, using a variety of “vision modules,” they can extract information from visual inputs and produce detailed textual representations (such as tags, attributes, actions, and relationships, among other things), which they can then feed directly to the LLM to avoid the need for additional multimodal pretraining, as shown in Fig. 1(b). Researchers from Contextual AI and Stanford University introduce LENS (Large Language Models ENnhanced to See) a modular strategy that uses an LLM as the “reasoning module” and functions across separate “vision modules.” 

🔥 Join The Fastest Growing ML Subreddit

They first extract rich textual information in the LENS technique using pretrained vision modules, such as contrastive models and image-captioning models. The text is then sent to the LLM, enabling it to carry out tasks, including object recognition, vision, and language (V&L). LENS bridges the gap between the modalities at no expense by eliminating the necessity for additional multimodal pretraining stages or data. Incorporating LENS gives them a model that operates across domains out of the box without the need for additional cross-domain pretraining. Additionally, this integration enables us to immediately use the most recent developments in computer vision and natural language processing, maximizing the advantages associated with both disciplines. 

They provide the following contributions: 

• They present LENS, a modular method that handles computer vision challenges by using language models’ few-shot, in-context learning capabilities through natural language descriptions of visual inputs. 

• LENS gives any off-the-shelf LLM the ability to see without further training or data. 

• They use frozen LLMs to handle object recognition and visual reasoning tasks without additional vision-and-language alignment or multimodal data. Experimental results show that their approach achieves zero-shot performance that is competitive with or superior to end-to-end jointly pre-trained models like Kosmos and Flamingo. A partial implementation of their paper is available on GitHub.


Check Out the Paper, Demo, Github link, and Blog. Don’t forget to join our 25k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools:

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How These XL Phones Compete

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.
AI & Technology

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

September 10, 2026
Next Post
Newell Rubbermaid’s Stock Is Surging Because of Pens and Water Bottles?

Newell Rubbermaid’s Stock Is Surging Because of Pens and Water Bottles?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Unrepentant Elizabeth Holmes already plotting biotech comeback from behind bars

Unrepentant Elizabeth Holmes already plotting biotech comeback from behind bars

September 9, 2026
Maryland school bus crashes into building

Maryland school bus crashes into building

September 7, 2026
Democrats play up Trump investigations, Epstein files during GOP convention in Dallas – The Washington Post

Democrats play up Trump investigations, Epstein files during GOP convention in Dallas – The Washington Post

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!