• bitcoinBitcoin(BTC)$78,824.00-1.26%
  • ethereumEthereum(ETH)$2,482.27-1.16%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$752.160.50%
  • rippleXRP(XRP)$1.39-1.35%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.36-2.10%
  • tronTRON(TRX)$0.336297-0.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,138.62-5.66%
  • HyperliquidHyperliquid(HYPE)$84.36-2.65%
  • dogecoinDogecoin(DOGE)$0.090096-0.22%
  • RainRain(RAIN)$0.016257-2.66%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$517.19-3.48%
  • chainlinkChainlink(LINK)$12.71-4.12%
  • whitebitWhiteBIT Coin(WBT)$76.293.61%
  • leo-tokenLEO Token(LEO)$9.21-0.38%
  • cardanoCardano(ADA)$0.218788-0.83%
  • stellarStellar(XLM)$0.191532-0.73%
  • bitcoin-cashBitcoin Cash(BCH)$260.170.67%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • uniswapUniswap(UNI)$7.03-0.10%
  • litecoinLitecoin(LTC)$55.361.47%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.106328-3.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-2.32%
  • hedera-hashgraphHedera(HBAR)$0.0815770.24%
  • avalanche-2Avalanche(AVAX)$8.072.25%
  • suiSui(SUI)$0.822.08%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.23%
  • nearNEAR Protocol(NEAR)$2.28-5.89%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0582071.07%
  • tether-goldTether Gold(XAUT)$4,420.160.37%
  • MemeCoreMemeCore(M)$1.170.62%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BittensorBittensor(TAO)$256.62-6.12%
  • okbOKB(OKB)$116.092.40%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • AsterAster(ASTER)$0.77-2.14%
  • mantleMantle(MNT)$0.62-5.86%
  • aaveAave(AAVE)$132.00-2.77%
  • pax-goldPAX Gold(PAXG)$4,424.680.43%
  • OndoOndo(ONDO)$0.379901-2.01%
  • polkadotPolkadot(DOT)$1.078.52%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Researchers Propose DeepSpeed-VisualChat: A Leap Forward in Scalable Multi-Modal Language Model Training

October 20, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Researchers Propose DeepSpeed-VisualChat: A Leap Forward in Scalable Multi-Modal Language Model Training
ShareShareShareShareShare

Large language models are sophisticated artificial intelligence systems created to understand and produce language similar to humans on a large scale. These models are useful in various applications, such as question-answering, content generation, and interactive dialogues. Their usefulness comes from a long learning process where they analyze and understand massive amounts of online data.

These models are advanced instruments that improve human-computer interaction by encouraging a more sophisticated and effective use of language in various contexts.

Beyond reading and writing text, research is being carried out to teach them how to comprehend and use various forms of information, such as sounds and images. The advancement in multi-modal capabilities is highly fascinating and holds great promise. Contemporary large language models (LLMs), such as GPT, have shown exceptional performance across a range of text-related tasks. These models become very good at different interactive tasks by using extra training methods like supervised fine-tuning or reinforcement learning with human guidance. To reach the level of expertise seen in human specialists, especially in challenges involving coding, quantitative thinking, mathematical reasoning, and engaging in conversations like AI chatbots, it is essential to refine the models through these training techniques.

It is getting closer to allowing these models to understand and create material in various formats, including images, sounds, and videos. Methods, including feature alignment and model modification, are applied. Large vision and language models (LVLMs) are one of these initiatives. However, because of problems with training and data availability, current models have difficulty addressing complicated scenarios, such as multi-image multi-round dialogues, and they are constrained in terms of adaptability and scalability in various interaction contexts.

The researchers of Microsoft have dubbed DeepSpeed-VisualChat. This framework enhances LLMs by incorporating multi-modal capabilities, demonstrating outstanding scalability even with a language model size of 70 billion parameters. This was formulated to facilitate dynamic chats with multi-round and multi-picture dialogues, seamlessly fusing text and image inputs. To increase the adaptability and responsiveness of multi-modal models, the framework uses Multi-Modal Causal Attention (MMCA), a method that separately estimates attention weights across several modalities. The team has used data blending approaches to overcome issues with the available datasets, resulting in a rich and varied training environment.

DeepSpeed-VisualChat is distinguished by its outstanding scalability, which was made possible by thoughtfully integrating the DeepSpeed framework. This framework exhibits exceptional scalability and pushes the limits of what is possible in multi-modal dialogue systems by utilizing a 2 billion parameter visual encoder and a 70 billion parameter language decoder from LLaMA-2. 

The researchers emphasize that DeepSpeed-VisualChat’s architecture is based on MiniGPT4. In this structure, an image is encoded using a pre-trained vision encoder and then aligned with the output of the text embedding layer’s hidden dimension using a linear layer. These inputs are fed into language models like LLaMA2, supported by the ground-breaking Multi-Modal Causal Attention (MMCA) mechanism. It is significant that during this procedure, both the language model and the vision encoder stay frozen.

According to the researchers, classic Cross Attention (CrA) provides new dimensions and problems, but Multi-Modal Causal Attention (MMCA) takes a different approach. For text and image tokens, MMCA uses separate attention weight matrices such that visual tokens focus on themselves and text permits focus on the tokens that came before them.

DeepSpeed-VisualChat is more scalable than previous models, according to real-world outcomes. It enhances adaption in various interaction scenarios without increasing complexity or training costs. With scaling up to a language model size of 70 billion parameters, it delivers particularly excellent scalability. This achievement provides a strong foundation for continued advancement in multi-modal language models and constitutes a significant step forward.


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on WhatsApp. Join our AI Channel on Whatsapp..


YOU MAY ALSO LIKE

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI

Rachit Ranjan is a consulting intern at MarktechPost . He is currently pursuing his B.Tech from Indian Institute of Technology(IIT) Patna . He is actively shaping his career in the field of Artificial Intelligence and Data Science and is passionate and dedicated for exploring these fields.


▶️ Now Watch AI Research Updates On Our Youtube Channel [Watch Now]

Credit: Source link

ShareTweetSendSharePin

Related Posts

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page
AI & Technology

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

September 8, 2026
XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI
AI & Technology

XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI

September 8, 2026
Uber, Wayve Unleash Supervised Robotaxis in London
AI & Technology

Uber, Wayve Unleash Supervised Robotaxis in London

September 8, 2026
Chip Suppliers Bullish on AI Buildout
AI & Technology

Chip Suppliers Bullish on AI Buildout

September 8, 2026
Next Post
FL State Attorney Suspended For Saying He Won’t Enforce Restrictions On Abortion

FL State Attorney Suspended For Saying He Won't Enforce Restrictions On Abortion

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Wife of Chiefs offensive coordinator allegedly shot by son

Wife of Chiefs offensive coordinator allegedly shot by son

September 4, 2026
Wildberries says warehouses struck in drone attack

Wildberries says warehouses struck in drone attack

September 5, 2026
Powerful storms slam the East Coast

Powerful storms slam the East Coast

September 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!