• bitcoinBitcoin(BTC)$76,676.001.24%
  • ethereumEthereum(ETH)$2,467.653.20%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$726.962.20%
  • rippleXRP(XRP)$1.313.16%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.174.25%
  • tronTRON(TRX)$0.334172-0.44%
  • zcashZcash(ZEC)$1,486.7319.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.022.04%
  • HyperliquidHyperliquid(HYPE)$82.514.91%
  • dogecoinDogecoin(DOGE)$0.0818763.63%
  • moneroMonero(XMR)$513.614.35%
  • USDSUSDS(USDS)$1.000.05%
  • whitebitWhiteBIT Coin(WBT)$79.041.58%
  • RainRain(RAIN)$0.012929-1.86%
  • chainlinkChainlink(LINK)$11.366.42%
  • leo-tokenLEO Token(LEO)$8.920.76%
  • cardanoCardano(ADA)$0.2019955.60%
  • stellarStellar(XLM)$0.1862947.41%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$232.807.58%
  • daiDai(DAI)$1.00-0.02%
  • uniswapUniswap(UNI)$7.1014.63%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.486.20%
  • CantonCanton(CC)$0.10131511.99%
  • nearNEAR Protocol(NEAR)$2.8817.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.87%
  • avalanche-2Avalanche(AVAX)$7.595.06%
  • hedera-hashgraphHedera(HBAR)$0.0763535.93%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000058.30%
  • suiSui(SUI)$0.737.17%
  • crypto-com-chainCronos(CRO)$0.0581515.39%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.05%
  • tether-goldTether Gold(XAUT)$4,360.310.27%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$227.946.34%
  • MemeCoreMemeCore(M)$1.141.04%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.242.78%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.04%
  • AsterAster(ASTER)$0.749.60%
  • BitwayBitway(BTW)$0.73-4.07%
  • aaveAave(AAVE)$125.798.50%
  • pax-goldPAX Gold(PAXG)$4,361.930.19%
  • mantleMantle(MNT)$0.574.46%
  • Pump.funPump.fun(PUMP)$0.00397310.23%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Friendship Ended with Single Modality – Now Multi-Modality is My Best Friend: CoDi is an AI Model that can Achieve Any-to-Any Generation via Composable Diffusion

June 21, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Friendship Ended with Single Modality – Now Multi-Modality is My Best Friend: CoDi is an AI Model that can Achieve Any-to-Any Generation via Composable Diffusion
ShareShareShareShareShare

Generative AI is a term we hear almost every day now. I even don’t remember how many papers I’ve read and summarized about generative AI here. They are impressive, what they do seems unreal and magical, and they can be used in many applications. We can generate images, videos, audio, and more by just using text prompts.

The significant progress made in generative AI models in recent years has enabled use cases that were deemed impossible not so long ago. It started with text-to-image models, and once it was seen that they produced incredibly nice results. After that, the demand for AI models capable of handling multiple modalities has increased.

YOU MAY ALSO LIKE

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

Recently, there is a surging demand for models that can take any combination of inputs (e.g., text + audio) and generate various combinations of modal outputs (e.g., video + audio) has increased. Several models have been proposed to tackle this, but these models have limitations regarding real-world applications involving multiple modalities that coexist and interact. 

🚀 JOIN the fastest ML Subreddit Community

While it’s possible to chain together modality-specific generative models in a multi-step process, the generation power of each step remains inherently limited, resulting in a cumbersome and slow approach. Additionally, independently generated unimodal streams may lack consistency and alignment when combined, making post-processing synchronization challenging.

Training a model to handle any mixture of input modalities and flexibly generate any combination of outputs presents significant computational and data requirements. The number of possible input-output combinations scales exponentially, while aligned training data for many groups of modalities is scarce or non-existent. 

Let us meet with CoDi, which is proposed to tackle this challenge. CoDi is a novel neural architecture that enables the simultaneous processing and generation of arbitrary combinations of modalities. 

Overview of CoDi. Source: https://arxiv.org/pdf/2305.11846.pdf

CoDi proposes aligning multiple modalities in both the input conditioning and generation diffusion steps. Additionally, it introduces a “Bridging Alignment” strategy for contrastive learning, enabling it to efficiently model the exponential number of input-output combinations with a linear number of training objectives.

The key innovation of CoDi lies in its ability to handle any-to-any generation by leveraging a combination of latent diffusion models (LDMs), multimodal conditioning mechanisms, and cross-attention modules. By training separate LDMs for each modality and projecting input modalities into a shared feature space, CoDi can generate any modality or combination of modalities without direct training for such settings. 

The development of CoDi requires comprehensive model design and training on diverse data resources. First, the training starts with a latent diffusion model (LDM) for each modality, such as text, image, video, and audio. These models can be trained independently in parallel, ensuring exceptional single-modality generation quality using modality-specific training data. For conditional cross-modality generation, where images are generated using audio+language prompts, the input modalities are projected into a shared feature space, and the output LDM attends to the combination of input features. This multimodal conditioning mechanism prepares the diffusion model to handle any modality or combination of modalities without direct training for such settings.

Overview of CoDi model. Source: https://arxiv.org/pdf/2305.11846.pdf

In the second stage of training, CoDi handles many-to-many generation strategies involving the simultaneous generation of arbitrary combinations of output modalities. This is achieved by adding a cross-attention module to each diffuser and an environment encoder to project the latent variable of different LDMs into a shared latent space. This seamless generation capability allows CoDi to generate any group of modalities without training on all possible generation combinations, reducing the number of training objectives from exponential to linear.


Check Out The Paper, Code, and Project. Don’t forget to join our 24k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools From AI Tools Club

🚀 Check Out 100’s AI Tools in AI Tools Club


Ekrem Çetinkaya received his B.Sc. in 2018, and M.Sc. in 2019 from Ozyegin University, Istanbul, Türkiye. He wrote his M.Sc. thesis about image denoising using deep convolutional networks. He received his Ph.D. degree in 2023 from the University of Klagenfurt, Austria, with his dissertation titled “Video Coding Enhancements for HTTP Adaptive Streaming Using Machine Learning.” His research interests include deep learning, computer vision, video encoding, and multimedia networking.


➡️ Meet Notion: Your Wiki, Docs, & Projects Together

Credit: Source link

ShareTweetSendSharePin

Related Posts

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches
AI & Technology

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches

September 17, 2026
OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI
AI & Technology

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

September 17, 2026
Europe’s EU Kids Act Would Ban Social Media Access For Children Under 13
AI & Technology

Europe’s EU Kids Act Would Ban Social Media Access For Children Under 13

September 17, 2026
NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections
AI & Technology

NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections

September 17, 2026
Next Post
Ford’s Third Quarter Results Show U.S. Sales Are Surging

Ford's Third Quarter Results Show U.S. Sales Are Surging

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
‘There was so much chaos’: 9/11 survivor reflects on the day of the attacks

‘There was so much chaos’: 9/11 survivor reflects on the day of the attacks

September 12, 2026
James Talarico responds to Ken Paxton’s AI deepfake ad

James Talarico responds to Ken Paxton’s AI deepfake ad

September 14, 2026
Anthropic CEO Says It’s Time to Slow AI Model Advances – Bloomberg.com

Anthropic CEO Says It’s Time to Slow AI Model Advances – Bloomberg.com

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!