• bitcoinBitcoin(BTC)$78,052.00-1.81%
  • ethereumEthereum(ETH)$2,469.46-1.88%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$718.20-4.83%
  • rippleXRP(XRP)$1.38-4.21%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.08-3.48%
  • tronTRON(TRX)$0.3393330.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.84%
  • zcashZcash(ZEC)$1,217.39-2.05%
  • HyperliquidHyperliquid(HYPE)$82.99-4.25%
  • dogecoinDogecoin(DOGE)$0.085280-6.40%
  • RainRain(RAIN)$0.0162941.17%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$511.542.47%
  • whitebitWhiteBIT Coin(WBT)$80.60-2.11%
  • chainlinkChainlink(LINK)$11.76-5.87%
  • leo-tokenLEO Token(LEO)$9.190.08%
  • cardanoCardano(ADA)$0.213210-4.06%
  • stellarStellar(XLM)$0.179104-6.03%
  • bitcoin-cashBitcoin Cash(BCH)$247.54-4.85%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.104158-3.88%
  • litecoinLitecoin(LTC)$52.35-4.10%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-3.02%
  • uniswapUniswap(UNI)$5.98-12.66%
  • avalanche-2Avalanche(AVAX)$7.75-3.56%
  • hedera-hashgraphHedera(HBAR)$0.076204-4.55%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • nearNEAR Protocol(NEAR)$2.42-0.57%
  • suiSui(SUI)$0.76-7.69%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.64%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057148-6.47%
  • MemeCoreMemeCore(M)$1.200.60%
  • tether-goldTether Gold(XAUT)$4,401.410.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BittensorBittensor(TAO)$251.90-4.44%
  • okbOKB(OKB)$112.17-2.39%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
  • mantleMantle(MNT)$0.59-8.48%
  • AsterAster(ASTER)$0.72-6.01%
  • aaveAave(AAVE)$124.23-4.67%
  • pax-goldPAX Gold(PAXG)$4,404.640.00%
  • polkadotPolkadot(DOT)$1.11-5.54%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.055964-0.12%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from China Introduce CogVLM: A Powerful Open-Source Visual Language Foundation Model

November 13, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from China Introduce CogVLM: A Powerful Open-Source Visual Language Foundation Model
ShareShareShareShareShare

Models of visual language are strong and flexible. Next, token prediction may be used to create a variety of vision and cross-modality tasks, such as picture captioning, visual question answering, visual grounding, and even segmentation. As VLMs are scaled up, useful skills like in-context learning also appear along with the enhancement of downstream activities. It is more difficult to train a VLM from the start with the same NLP performance as well-trained pure language models like LLaMA2, as introducing a big language model is already a difficult task. Consequently, it makes sense to look at the process of training a VLM using a pre-trained language model that is readily available. 

The widely used shallow alignment techniques, represented by BLIP-2, transfer the image characteristics into the language model’s input embedding space using a trainable Q-Former or a linear layer, which connects a frozen pretrained vision encoder and language model. While this approach converges quickly, it does not perform as well as training the language and vision modules simultaneously, such as PaLI-X. When it comes to chat-style VLM that was taught using shallow alignment techniques, such as MiniGPT-4, LLAVA, and VisualGLM, the poor visual comprehension skills show up as hallucinations. Is it feasible to enhance the big language model’s visual understanding skills without sacrificing its natural language processing (NLP) capabilities? 

CogVLM responds with a “yes.” Researchers from Zhipu AI and Tsinghua University introduced CogVLM. This powerful open-source visual language foundation model believes the lack of deep integration between language and visual information is the primary reason for the shallow alignment approaches’ subpar performance. This idea came from comparing the two approaches to effective finetuning: p-tuning learns a task prefix embedding in the input. LoRA uses a low-rank matrix to adjust the model weights in each layer. LoRA functions more effectively and steadily as a result. Since the picture features in the shallow alignment techniques behave similarly to the prefix embedding in p-tuning, a similar occurrence may also occur in VLM. 

The following are more specific causes of p-tuning and shallow alignment’s decreased performance: 

1. Text tokens train the language model’s frozen weights. The input text area only perfectly matches visual characteristics. The visual characteristics may, therefore, no longer match the input distribution of the weights in the deep layers following multi-layer modifications. 

2. The writing style and caption length of the picture captioning job, for instance, may only be encoded into the visual characteristics in the shallow alignment approaches during pretraining. The coherence between the visual elements and the content could be stronger. Adapting the language model to the image-text combined training, as used by Qwen-VL and PaLI, is one potential remedy.

However, this unnecessarily impairs NLP, which may impact text-centered activities like creating image-based poetry or providing context for pictures. Making the language model trainable during VLM pretraining, according to PaLM-E, will result in catastrophic forgetting and a loss of 87.3% in the NLG performance for the 8B language model. Instead, CogVLM enhances the language model with a trainable visual expert. Each layer uses a separate QKV matrix for the picture features in the sequence and an MLP layer for the text characteristics. The visual expert maintains the same FLOPs but increases the number of parameters. If there isn’t an image in the input sequence, the behaviors are the same as in the original language model since all parameters are fixed. 

On 14 typical cross-modal benchmarks, such as: 1) image captioning datasets (NoCaps, Flicker30k, COCO), 2) VQA datasets (VQAv2, OKVQA, GQA, TextVQA, VizWiz), and 3) image captioning datasets (SecondBest), their CogVLM-17B trained from Vicuna-7B achieves state-of-the-art or the second-best performance. 3) multiple choice datasets (TDIUC, ScienceQA); 4) visual grounding datasets (RefCOCO, RefCOCO+, RefCOCOg, Visual7W). Not included in this study is the CogVLM-28B-zh that they trained from ChatGLM-12B to support both Chinese and English for commercial use. Since the majority of the most well-known VLMs in the past, such as Flamingo, SimVLM, Coca, BEIT-3, GIT2, PaLI, and PaLI-X, are closed-source, it is anticipated that CogVLM’s open-sourcing will have a significant positive impact on visual understanding research and industrial application.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 32k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

2028 Volvo XC40 First Look: Hello new tech, goodbye EV
AI & Technology

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

September 10, 2026
Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI
AI & Technology

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

September 10, 2026
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Next Post
This AI Research Unveils LSS Transformer: A Revolutionary AI Approach for Efficient Long Sequence Training in Transformers

This AI Research Unveils LSS Transformer: A Revolutionary AI Approach for Efficient Long Sequence Training in Transformers

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Kornacki on Trump’s approval rating 100 days before midterms

Kornacki on Trump’s approval rating 100 days before midterms

September 4, 2026
A rare look inside the Pope’s summer residence garden

A rare look inside the Pope’s summer residence garden

September 3, 2026
Staying With Your Employer Is a Financial Choice

Staying With Your Employer Is a Financial Choice

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!