• bitcoinBitcoin(BTC)$77,526.001.61%
  • ethereumEthereum(ETH)$2,482.871.86%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$752.853.90%
  • rippleXRP(XRP)$1.321.79%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$105.445.95%
  • tronTRON(TRX)$0.3357400.14%
  • zcashZcash(ZEC)$1,518.3512.58%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.16%
  • HyperliquidHyperliquid(HYPE)$87.4911.11%
  • dogecoinDogecoin(DOGE)$0.0841864.24%
  • moneroMonero(XMR)$525.215.22%
  • USDSUSDS(USDS)$1.000.04%
  • whitebitWhiteBIT Coin(WBT)$79.861.77%
  • RainRain(RAIN)$0.012641-1.77%
  • chainlinkChainlink(LINK)$11.745.72%
  • leo-tokenLEO Token(LEO)$8.92-0.18%
  • cardanoCardano(ADA)$0.2127928.41%
  • stellarStellar(XLM)$0.1870912.64%
  • uniswapUniswap(UNI)$8.5326.13%
  • bitcoin-cashBitcoin Cash(BCH)$246.7911.72%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.4730.90%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.10946311.74%
  • litecoinLitecoin(LTC)$54.795.48%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.353.33%
  • avalanche-2Avalanche(AVAX)$7.885.00%
  • hedera-hashgraphHedera(HBAR)$0.0763863.68%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.787.85%
  • shiba-inuShiba Inu(SHIB)$0.0000056.86%
  • crypto-com-chainCronos(CRO)$0.0585760.54%
  • MemeCoreMemeCore(M)$1.2814.66%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BittensorBittensor(TAO)$241.007.72%
  • tether-goldTether Gold(XAUT)$4,372.471.83%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$114.432.63%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.05%
  • aaveAave(AAVE)$133.569.33%
  • AsterAster(ASTER)$0.752.79%
  • Pump.funPump.fun(PUMP)$0.00423210.94%
  • mantleMantle(MNT)$0.584.60%
  • polkadotPolkadot(DOT)$1.1211.03%
  • pax-goldPAX Gold(PAXG)$4,371.251.76%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta AI Introduces Chameleon: A New Family of Early-Fusion Token-based Foundation Models that Set a New Bar for Multimodal Machine Learning

May 18, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Meta AI Introduces Chameleon: A New Family of Early-Fusion Token-based Foundation Models that Set a New Bar for Multimodal Machine Learning
ShareShareShareShareShare

Although recent multimodal foundation models are extensively utilized, they tend to segregate various modalities, typically employing specific encoders or decoders for each. This approach constrains their capacity to fuse information across modalities effectively and produce multimodal documents comprising diverse sequences of images and text. Consequently, there’s a limitation in their ability to seamlessly integrate different types of content within a single document.

Meta researchers present Chameleon, a mixed-modal foundation model that facilitates generating and reasoning with interleaved textual and image sequences, enabling comprehensive multimodal document modeling. Unlike traditional models, Chameleon employs a unified architecture, treating both modalities equally by tokenizing images akin to text. This approach, termed early fusion, allows seamless reasoning across modalities but poses optimization challenges. To address these, the researchers propose architectural enhancements and training techniques. By adapting transformer architecture and finetuning strategies.

Researchers developed a novel image tokenizer, encoding 512 × 512 images into 1024 tokens from an 8192-codebook, focusing on licensed images and doubling face-containing images during pre-training. However, their tokenizer struggles with text-heavy image reconstruction. Also, they trained a BPE tokenizer with a 65,536-vocabulary, including image tokens, using the sentencepiece library, over a subset of training data. Chameleon addressed stability issues with QK-Norm, dropout, and z-loss regularization during training, facilitating successful training on Meta’s RSC. Inference streamlined processing for mixed-modal generation using PyTorch and xformers, supporting both streaming and non-streaming modes with token masking for conditional logic.

The alignment stage, fine-tunes on diverse datasets, including Text, Code, Visual Chat, and Safety, aiming to enhance model capabilities and safety. They curate high-quality images for Image Generation using an aesthetic classifier. Supervised Fine-Tuning (SFT) encompasses data balancing across modalities, utilizing a cosine learning rate schedule and a weight decay of 0.1. Each instance in SFT pairs prompts with corresponding answers, optimizing exclusively based on the latter. Dropout of 0.05 is applied, along with z-loss regularization. Images in prompts are resized with border padding, while those in answers are center-cropped for quality image generation.

Chameleon evaluates its text-only capabilities against state-of-the-art models, achieving competitive performance across various tasks like commonsense reasoning and math. It outperforms LLaMa-2 on many tasks, attributing gains to better pre-training and inclusion of code data. In image-to-text tasks, Chameleon excels in image captioning, matching or surpassing larger models like Flamingo-80B and IDEFICS-80B with fewer shots. In visual question answering (VQA), it approaches performance of top models, even though Llava-1.5 slightly outperforms VQA-v2. Chameleon’s versatility and efficiency make it competitive across different tasks, requiring fewer training examples and smaller model sizes.

To recapitulate, This study introduces Chameleon, a token-based model, achieves superior performance in vision-language tasks by integrating image and text tokens seamlessly. Its architecture enables joint reasoning over modalities, surpassing late-fusion models like Flamingo and IDEFICS in tasks like image captioning and visual question answering. Chameleon’s early-fusion approach introduces novel techniques for stable training, addressing previous scalability challenges. It unlocks new multimodal interaction possibilities, evident in its strong performance on mixed-modal open-ended QA benchmarks.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

eGPUs Do Work, But They Come With Some Notable Limitations

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI
AI & Technology

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

September 18, 2026
eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
Next Post
Congressional inaction on border deal will ‘continue to be chaos,’ says Colorado governor

Congressional inaction on border deal will ‘continue to be chaos,’ says Colorado governor

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI

Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI

September 11, 2026
Wild animal attacks three people in New Hampshire

Wild animal attacks three people in New Hampshire

September 18, 2026
Current with Christine Romans – Sept. 7 | NBC News NOW

Current with Christine Romans – Sept. 7 | NBC News NOW

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!