• bitcoinBitcoin(BTC)$85,782.005.58%
  • ethereumEthereum(ETH)$2,746.273.27%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$791.631.29%
  • rippleXRP(XRP)$1.527.87%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.925.92%
  • tronTRON(TRX)$0.3476831.34%
  • zcashZcash(ZEC)$1,460.98-3.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$93.390.42%
  • dogecoinDogecoin(DOGE)$0.09956513.16%
  • moneroMonero(XMR)$586.042.48%
  • whitebitWhiteBIT Coin(WBT)$86.283.88%
  • RainRain(RAIN)$0.013872-2.14%
  • chainlinkChainlink(LINK)$13.043.45%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2480497.97%
  • leo-tokenLEO Token(LEO)$8.950.25%
  • stellarStellar(XLM)$0.2145918.65%
  • nearNEAR Protocol(NEAR)$4.568.91%
  • uniswapUniswap(UNI)$9.245.99%
  • bitcoin-cashBitcoin Cash(BCH)$266.764.97%
  • avalanche-2Avalanche(AVAX)$11.19-0.48%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$61.053.87%
  • CantonCanton(CC)$0.1180186.08%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.0613.90%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.454.71%
  • hedera-hashgraphHedera(HBAR)$0.0926567.86%
  • BittensorBittensor(TAO)$321.7820.93%
  • shiba-inuShiba Inu(SHIB)$0.0000068.82%
  • crypto-com-chainCronos(CRO)$0.0677239.94%
  • MemeCoreMemeCore(M)$1.44-4.59%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,350.73-0.41%
  • okbOKB(OKB)$123.082.87%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • aaveAave(AAVE)$144.374.79%
  • OndoOndo(ONDO)$0.4431184.12%
  • EthenaEthena(ENA)$0.212556-0.13%
  • mantleMantle(MNT)$0.656.22%
  • pepePepe(PEPE)$0.00000525.38%
  • BitwayBitway(BTW)$0.784.63%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meissonic: A Non-Autoregressive Mask Image Modeling Text-to-Image Synthesis Model that can Generate High-Resolution Images

October 17, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Meissonic: A Non-Autoregressive Mask Image Modeling Text-to-Image Synthesis Model that can Generate High-Resolution Images
ShareShareShareShareShare

Large Language Models (LLMs) have demonstrated remarkable progress in natural language processing tasks, inspiring researchers to explore similar approaches for text-to-image synthesis. At the same time, diffusion models have become the dominant approach in visual generation. However, the operational differences between the two approaches present a significant challenge in developing a unified methodology for language and vision tasks. Recent developments like LlamaGen have ventured into autoregressive image generation using discrete image tokens; however, it is inefficient due to the large number of image tokens compared to text tokens. Non-autoregressive methods like MaskGIT and MUSE have emerged, cutting down on the number of decoding steps, but failing to produce high-quality, high-resolution images.

Existing attempts to solve the challenges in text-to-image synthesis have mainly focused on two approaches: diffusion-based and token-based image generation. Diffusion models, like Stable Diffusion and SDXL, have made significant progress by working within compressed latent spaces and introducing techniques like micro-conditions and multi-aspect training. The integration of transformer architectures, as seen in DiT and U-ViT, has further enhanced the potential of diffusion models. However, these models still face challenges in real-time applications and quantization. Token-based approaches like MaskGIT and MUSE, have introduced masked image modeling (MIM) to overcome the computational demands of autoregressive methods.

YOU MAY ALSO LIKE

Why It’s Important To Unplug Your PC During A Power Outage

Why Is Your Laptop Fan So Loud?

Researchers from Alibaba Group, Skywork AI, HKUST(GZ), HKUST, Zhejiang University, and UC Berkeley have proposed Meissonic, an innovative method to elevate non-autoregressive MIM text-to-image synthesis to a level comparable with state-of-the-art diffusion models like SDXL. Meissonic utilizes a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions to enhance MIM’s performance and efficiency. The model uses high-quality training data, micro-conditions informed by human preference scores, and feature compression layers to improve image fidelity and resolution. The Meissonic can produce 1024 × 1024 resolution images and often outperforms existing models in generating high-quality, high-resolution images.

Meissonic’s architecture integrates a CLIP text encoder, a vector-quantized (VQ) image encoder and decoder, and a multi-modal Transformer backbone for efficient high-performance text-to-image synthesis:

  • The VQ-VAE model converts raw image pixels into discrete semantic tokens using a learned codebook. 
  • A fine-tuned CLIP text encoder with a 1024 latent dimension is used for optimal performance. 
  • The multi-modal Transformer backbone utilizes sampling parameters and Rotary Position Embeddings for spatial information encoding. 
  • Feature compression layers are used to handle high-resolution generation efficiently.

The architecture also includes QK-Norm layers and implements gradient clipping to enhance training stability and reduce NaN Loss issues during distributed training.

Meissonic, optimized to 1 billion parameters, runs efficiently on 8GB VRAM, making inference and fine-tuning convenient. Qualitative comparisons show Meissonic’s image quality and text-image alignment capabilities. Human evaluations using K-Sort Arena and GPT-4 assessments indicate that Meissonic achieves performance comparable to DALL-E 2 and SDXL in human preference and text alignment, with improved efficiency. Meissonic is benchmarked against state-of-the-art models using the EMU-Edit dataset in image editing tasks, covering seven different operations. The model demonstrated versatility in both mask-guided and mask-free editing, achieving great performance without specific training on image editing data or instruction datasets.

In conclusion, researchers introduced Meissonic, an approach to elevate non-autoregressive MIM text-to-image synthesis. The model incorporates innovative elements such as a blended transformer architecture, advanced positional encoding, and adaptive masking rates to achieve superior performance in high-resolution image generation. Despite its compact 1B parameter size, Meissonic outperforms larger diffusion models while remaining accessible on consumer-grade GPUs. Moreover, Meissonic aligns with the emerging trend of offline text-to-image applications on mobile devices, exemplified by recent innovations from Google and Apple. It enhances the user experience and privacy in mobile imaging technology, empowering users with creative tools while ensuring data security.


Check out the Paper and Model. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 50k+ ML SubReddit.

[Upcoming Live Webinar- Oct 29, 2024] The Best Platform for Serving Fine-Tuned Models: Predibase Inference Engine (Promoted)


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Next Post
How a ‘putrid’ find in a museum cupboard could be the key to bringing the Tasmanian tiger back to life – The Guardian

How a ‘putrid’ find in a museum cupboard could be the key to bringing the Tasmanian tiger back to life - The Guardian

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Survivors return to flood-stricken communities in Nepal

Survivors return to flood-stricken communities in Nepal

September 19, 2026
Woman accused of photographing Lindsay Clancy jurors charged with witness intimidation

Woman accused of photographing Lindsay Clancy jurors charged with witness intimidation

September 19, 2026
LIVE: Steve Kornacki analyzes New Hampshire primary election results | Kornacki Cam | NBC News

LIVE: Steve Kornacki analyzes New Hampshire primary election results | Kornacki Cam | NBC News

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!