• bitcoinBitcoin(BTC)$85,368.004.57%
  • ethereumEthereum(ETH)$2,730.402.34%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$786.671.79%
  • rippleXRP(XRP)$1.515.92%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.804.35%
  • tronTRON(TRX)$0.3486681.56%
  • zcashZcash(ZEC)$1,494.64-1.57%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$93.83-0.26%
  • dogecoinDogecoin(DOGE)$0.09998612.87%
  • moneroMonero(XMR)$572.95-4.88%
  • whitebitWhiteBIT Coin(WBT)$85.863.05%
  • RainRain(RAIN)$0.013662-3.62%
  • chainlinkChainlink(LINK)$12.912.73%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2435014.92%
  • leo-tokenLEO Token(LEO)$8.950.41%
  • stellarStellar(XLM)$0.2120437.04%
  • nearNEAR Protocol(NEAR)$4.434.82%
  • uniswapUniswap(UNI)$8.994.42%
  • bitcoin-cashBitcoin Cash(BCH)$265.154.52%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$10.71-4.40%
  • litecoinLitecoin(LTC)$60.904.52%
  • CantonCanton(CC)$0.1173544.60%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.027.24%
  • hedera-hashgraphHedera(HBAR)$0.0926836.55%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.71%
  • BittensorBittensor(TAO)$323.0619.82%
  • shiba-inuShiba Inu(SHIB)$0.0000068.72%
  • crypto-com-chainCronos(CRO)$0.0658507.14%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.36-10.36%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,323.13-0.69%
  • okbOKB(OKB)$121.121.33%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • aaveAave(AAVE)$143.323.93%
  • EthenaEthena(ENA)$0.2164231.18%
  • pepePepe(PEPE)$0.00000528.58%
  • BitwayBitway(BTW)$0.803.32%
  • mantleMantle(MNT)$0.645.19%
  • OndoOndo(ONDO)$0.4328870.86%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Apple Releases 4M-21: A Very Effective Multimodal AI Model that Solves Tens of Tasks and Modalities

June 18, 2024
in AI & Technology
Reading Time: 7 mins read
A A
Apple Releases 4M-21: A Very Effective Multimodal AI Model that Solves Tens of Tasks and Modalities
ShareShareShareShareShare

Large language models (LLMs) have made significant strides in handling multiple modalities and tasks, but they still need to improve their ability to process diverse inputs and perform a wide range of tasks effectively. The primary challenge lies in developing a single neural network capable of handling a broad spectrum of tasks and modalities while maintaining high performance across all domains. Current models, such as 4M and UnifiedIO, show promise but are constrained by the limited number of modalities and tasks they are trained on. This limitation hinders their practical application in scenarios requiring truly versatile and adaptable AI systems.

Recent attempts to solve multitask learning challenges in vision have evolved from combining dense vision tasks to integrating numerous tasks into unified multimodal models. Methods like Gato, OFA, Pix2Seq, UnifiedIO, and 4M transform various modalities into discrete tokens and train Transformers using sequence or masked modeling objectives. Some approaches enable a wide range of tasks through co-training on disjoint datasets, while others, like 4M, use pseudo labeling for any-to-any modality prediction on aligned datasets. Masked modeling has proven effective in learning cross-modal representations, crucial for multimodal learning, and enables generative applications when combined with tokenization.

Researchers from Apple and the Swiss Federal Institute of Technology Lausanne (EPFL) build their method upon the multimodal masking pre-training scheme, significantly expanding its capabilities by training on a diverse set of modalities. The approach incorporates over 20 modalities, including SAM segments, 3D human poses, Canny edges, color palettes, and various metadata and embeddings. By using modality-specific discrete tokenizers, the method encodes diverse inputs into a unified format, enabling the training of a single model on multiple modalities without performance degradation. This unified approach expands existing capabilities across several key axes, including increased modality support, improved diversity in data types, effective tokenization techniques, and scaled model size. The resulting model demonstrates new possibilities for multimodal interaction, such as cross-modal retrieval and highly steerable generation across all training modalities.

This method adopts the 4M pre-training scheme, expanding it to handle a diverse set of modalities. It transforms all modalities into sequences of discrete tokens using modality-specific tokenizers. The training objective involves predicting one subset of tokens from another, using random selections from all modalities as inputs and targets. It utilizes pseudo-labeling to create a large pre-training dataset with multiple aligned modalities. The method incorporates a wide range of modalities, including RGB, geometric, semantic, edges, feature maps, metadata, and text. Tokenization plays a crucial role in unifying the representation space across these diverse modalities. This unification enables training with a single pre-training objective, improves training stability, allows full parameter sharing, and eliminates the need for task-specific components. Three main types of tokenizers are employed: ViT-based tokenizers for image-like modalities, MLP tokenizers for human poses and global embeddings, and a WordPiece tokenizer for text and other structured data. This comprehensive tokenization approach allows the model to handle a wide array of modalities efficiently, reducing computational complexity and enabling generative tasks across multiple domains.

The 4M-21 model demonstrates a wide range of capabilities, including steerable multimodal generation, multimodal retrieval, and strong out-of-the-box performance across various vision tasks. It can predict any training modality by iteratively decoding tokens, enabling fine-grained and multimodal generation with improved text understanding. The model performs multimodal retrievals by predicting global embeddings from any input modality, allowing for versatile retrieval capabilities. In out-of-the-box evaluations, 4M-21 achieves competitive performance on tasks such as surface normal estimation, depth estimation, semantic segmentation, instance segmentation, 3D human pose estimation, and image retrieval. It often matches or outperforms specialist models and pseudo-labelers while being a single model for all tasks. The 4M-21 XL variant, in particular, demonstrates strong performance across multiple modalities without sacrificing capability in any single domain.

Researchers examine the scaling characteristics of pre-training any-to-any models on a large set of modalities, comparing three model sizes: B, L, and XL. Evaluating both unimodal (RGB) and multimodal (RGB + Depth) transfer learning scenarios. In unimodal transfers, 4M-21 maintains performance on tasks similar to the original seven modalities while showing improved results on complex tasks like 3D object detection. The model demonstrates better performance with increased size, indicating promising scaling trends. For multimodal transfers, 4M-21 effectively utilizes optional depth inputs, significantly outperforming baselines. The study reveals that training on a broader set of modalities does not compromise performance on familiar tasks and can enhance capabilities on new ones, especially as model size increases.

This research demonstrates the successful training of an any-to-any model on a diverse set of 21 modalities and tasks. This achievement is made possible by employing modality-specific tokenizers to map all modalities to discrete sets of tokens, coupled with a multimodal masked training objective. The model scales to three billion parameters across multiple datasets without compromising performance compared to more specialized models. The resulting unified model exhibits strong out-of-the-box capabilities and opens new avenues for multimodal interaction, generation, and retrieval. However, the study acknowledges certain limitations and areas for future work. These include the need to further explore transfer and emergent capabilities, which remain largely untapped compared to language models. 


Check out the Paper, Project, and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit

We are releasing 4M-21 with a permissive license, including its source code and trained models. It’s a pretty effective multimodal model that solves 10s of tasks & modalities. See the demo code, sample results, and the tokenizers of diverse modalities on the website.

IMO, the… https://t.co/0hY0fHxtzB pic.twitter.com/o0BjwlSmeP

— Amir Zamir (@zamir_ar) June 14, 2024


YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
New Mexico AG sues Meta, alleges it is a ‘breeding ground’ for predators

New Mexico AG sues Meta, alleges it is a ‘breeding ground’ for predators

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
‘We’re done’ with the Democratic & Republican machine: New Hampshire DSA Senate candidate

‘We’re done’ with the Democratic & Republican machine: New Hampshire DSA Senate candidate

September 15, 2026
Miami plane crash victims were in a cleaning-company van

Miami plane crash victims were in a cleaning-company van

September 16, 2026
CNX Resources: Use The Time Left To Head To Industry Leaders And Cash (NYSE:CNX)

CNX Resources: Use The Time Left To Head To Industry Leaders And Cash (NYSE:CNX)

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!