• bitcoinBitcoin(BTC)$79,288.000.45%
  • ethereumEthereum(ETH)$2,450.940.56%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$714.840.25%
  • rippleXRP(XRP)$1.411.10%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.31-0.27%
  • tronTRON(TRX)$0.3295450.22%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.93%
  • HyperliquidHyperliquid(HYPE)$84.952.36%
  • zcashZcash(ZEC)$979.3613.75%
  • dogecoinDogecoin(DOGE)$0.0850621.73%
  • RainRain(RAIN)$0.016604-0.25%
  • moneroMonero(XMR)$521.541.42%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$11.621.72%
  • whitebitWhiteBIT Coin(WBT)$72.921.00%
  • leo-tokenLEO Token(LEO)$9.30-0.12%
  • cardanoCardano(ADA)$0.2138441.54%
  • stellarStellar(XLM)$0.1805440.72%
  • bitcoin-cashBitcoin Cash(BCH)$251.060.32%
  • daiDai(DAI)$1.00-0.04%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.108506-1.00%
  • USD1USD1(USD1)$1.000.01%
  • uniswapUniswap(UNI)$6.313.24%
  • litecoinLitecoin(LTC)$50.33-0.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.362.19%
  • hedera-hashgraphHedera(HBAR)$0.077556-0.38%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.370.68%
  • shiba-inuShiba Inu(SHIB)$0.0000050.06%
  • suiSui(SUI)$0.75-1.88%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0564832.70%
  • tether-goldTether Gold(XAUT)$4,419.40-0.90%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • nearNEAR Protocol(NEAR)$1.962.37%
  • MemeCoreMemeCore(M)$1.103.51%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$107.74-0.05%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.03%
  • BittensorBittensor(TAO)$223.550.93%
  • aaveAave(AAVE)$131.130.37%
  • AsterAster(ASTER)$0.73-0.20%
  • pax-goldPAX Gold(PAXG)$4,424.48-1.00%
  • mantleMantle(MNT)$0.582.51%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0571751.73%
  • OndoOndo(ONDO)$0.3562830.22%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

The Artist Pal in Your Pocket: SnapFusion is an AI Approach That Brings the Power of Diffusion Models to Mobile Devices

June 24, 2023
in AI & Technology
Reading Time: 6 mins read
A A
The Artist Pal in Your Pocket: SnapFusion is an AI Approach That Brings the Power of Diffusion Models to Mobile Devices
ShareShareShareShareShare

Diffusion models. This is a term you heard about a lot if you have been paying attention to the advancements in the AI domain. They were the key that enabled the revolution in generative AI methods. We now have models that can generate photorealistic images using text prompts in a span of seconds. They have revolutionized content generation, image editing, super-resolution, video synthesis, and 3D asset generation.

Though this impressive performance does not come cheap. Diffusion models are extremely demanding in terms of computation requirements. That means you need really high-end GPUs to make full use of them. Yes, there are also attempts to make them run on your local computers; but even then, you need a high-end one.  On the other hand, using a cloud provider can be an alternative solution, but then you might risk your privacy in that case.

Then, there is also the on-the-go aspect we need to think about. For the majority of people, they spend more time on their phones than their computers. If you want to use diffusion models on your mobile device, well, good luck with that, as it will be too demanding for the limited hardware power of the device itself. 

🔥 Unleash the power of Live Proxies: Private, undetectable residential and mobile IPs.

Diffusion models are the next big thing, but we need to tackle their complexity before applying them in practical applications. There have been multiple attempts that have focused on speeding up inference on mobile devices, but they have not achieved seamless user experience or quantitatively evaluated generation quality. Well, it was the story until now because we have a new player on the field, and it’s named SnapFusion.

SnapFusion is the first text-to-image diffusion model that generates images on mobile devices in less than 2 seconds. It optimizes the UNet architecture and reduces the number of denoising steps to improve inference speed. Additionally, it uses an evolving training framework, introduces data distillation pipelines, and enhances the learning objective during step distillation.

Overview of SnapFusion. Source: https://arxiv.org/pdf/2306.00980.pdf

Before making any changes to the structure, the authors of SnapFusion first investigated the architecture redundancy of SD-v1.5 to obtain efficient neural networks. However, applying conventional pruning or architecture search techniques to SD was challenging due to the high training cost. Any changes in the architecture may result in degraded performance, requiring extensive fine-tuning with significant computational resources. So, that road was blocked, and they had to develop alternative solutions that can preserve the performance of the pre-trained UNet model while gradually improving its efficacy.

To increase inference speed, SnapFusion focuses on optimizing the UNet architecture, which is a bottleneck in the conditional diffusion model. Existing works primarily focus on post-training optimizations, but SnapFusion identifies architecture redundancies and proposes an evolving training framework that outperforms the original Stable Diffusion model while significantly improving speed. It also introduces a data distillation pipeline to compress and accelerate the image decoder.

SnapFusion includes a robust training phase, where stochastic forward propagation is applied to execute each cross-attention and ResNet block with a certain probability. This robust training augmentation ensures that the network is tolerant to architecture permutations, allowing for accurate assessment of each block and stable architectural evolution. 

The efficient image decoder is achieved through a distillation pipeline that uses synthetic data to train the decoder obtained via channel reduction. This compressed decoder has significantly fewer parameters and is faster than the one from SD-v1.5. The distillation process involves generating two images, one from the efficient decoder and the other from SD-v1.5, using text prompts to obtain the latent representation from the UNet of SD-v1.5.

The proposed step distillation approach includes a vanilla distillation loss objective, which aims to minimize the discrepancy between the student UNet’s prediction and the teacher UNet’s noisy latent representation. Additionally, a CFG-aware distillation loss objective is introduced to improve the CLIP score. CFG-guided predictions are used in both the teacher and student models, where the CFG scale is randomly sampled to provide a trade-off between FID and CLIP scores during training.

Sample images generated by SnapFusion. Source: https://arxiv.org/pdf/2306.00980.pdf

Thanks to the improved step distillation and network architecture development, SnapFusion can generate 512 × 512 images from text prompts on mobile devices in less than 2 seconds. The generated images exhibit quality similar to the state-of-the-art Stable Diffusion model.


Check Out The Paper and Project Page. Don’t forget to join our 25k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools From AI Tools Club

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

A Worthy Android Ereader, With Some Tradeoffs

Insurance Spent Years Talking About AI. This Year It Actually Used It – Unite.AI

Ekrem Çetinkaya received his B.Sc. in 2018, and M.Sc. in 2019 from Ozyegin University, Istanbul, Türkiye. He wrote his M.Sc. thesis about image denoising using deep convolutional networks. He received his Ph.D. degree in 2023 from the University of Klagenfurt, Austria, with his dissertation titled “Video Coding Enhancements for HTTP Adaptive Streaming Using Machine Learning.” His research interests include deep learning, computer vision, video encoding, and multimedia networking.


Credit: Source link

ShareTweetSendSharePin

Related Posts

A Worthy Android Ereader, With Some Tradeoffs
AI & Technology

A Worthy Android Ereader, With Some Tradeoffs

September 4, 2026
Insurance Spent Years Talking About AI. This Year It Actually Used It – Unite.AI
AI & Technology

Insurance Spent Years Talking About AI. This Year It Actually Used It – Unite.AI

September 4, 2026
Sam Altman Apologizes as GPT-6 Astra Staged Launch Denies Paid Access – Unite.AI
AI & Technology

Sam Altman Apologizes as GPT-6 Astra Staged Launch Denies Paid Access – Unite.AI

September 4, 2026
This Rugged Smartphone’s Camera Is A Removable Action Cam
AI & Technology

This Rugged Smartphone’s Camera Is A Removable Action Cam

September 4, 2026
Next Post
Synovus Financial: Interesting Dividend But It’s Not Enough (NYSE:SNV)

Synovus Financial: Interesting Dividend But It's Not Enough (NYSE:SNV)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
FCC chair Carr defends new ban on foreign-made humanoid robots

FCC chair Carr defends new ban on foreign-made humanoid robots

September 2, 2026
Israeli UN Ambassador: Israel ‘will apply more power’ in Gaza if Hamas does not disarm

Israeli UN Ambassador: Israel ‘will apply more power’ in Gaza if Hamas does not disarm

September 2, 2026
Kentucky Gov. Beshear demands ‘transparency and honesty’ on McConnell’s health

Kentucky Gov. Beshear demands ‘transparency and honesty’ on McConnell’s health

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!