• bitcoinBitcoin(BTC)$81,117.006.23%
  • ethereumEthereum(ETH)$2,619.657.14%
  • tetherTether(USDT)$1.000.04%
  • binancecoinBNB(BNB)$762.843.90%
  • rippleXRP(XRP)$1.408.00%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$112.9111.43%
  • tronTRON(TRX)$0.3383521.11%
  • zcashZcash(ZEC)$1,509.732.14%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.24%
  • HyperliquidHyperliquid(HYPE)$92.2110.00%
  • dogecoinDogecoin(DOGE)$0.0882338.42%
  • moneroMonero(XMR)$563.219.39%
  • whitebitWhiteBIT Coin(WBT)$83.225.81%
  • USDSUSDS(USDS)$1.000.03%
  • RainRain(RAIN)$0.0134355.00%
  • chainlinkChainlink(LINK)$12.288.32%
  • cardanoCardano(ADA)$0.22593012.56%
  • leo-tokenLEO Token(LEO)$8.91-0.05%
  • stellarStellar(XLM)$0.1925934.95%
  • uniswapUniswap(UNI)$8.7414.50%
  • bitcoin-cashBitcoin Cash(BCH)$262.6012.86%
  • nearNEAR Protocol(NEAR)$3.7322.02%
  • Ethena USDeEthena USDe(USDE)$1.000.05%
  • daiDai(DAI)$1.00-0.02%
  • litecoinLitecoin(LTC)$57.837.33%
  • CantonCanton(CC)$0.11164010.87%
  • USD1USD1(USD1)$1.000.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.383.24%
  • avalanche-2Avalanche(AVAX)$8.198.13%
  • hedera-hashgraphHedera(HBAR)$0.0791806.64%
  • suiSui(SUI)$0.8211.31%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000057.33%
  • MemeCoreMemeCore(M)$1.309.63%
  • crypto-com-chainCronos(CRO)$0.0595393.97%
  • BittensorBittensor(TAO)$248.838.47%
  • paypal-usdPayPal USD(PYUSD)$1.000.08%
  • tether-goldTether Gold(XAUT)$4,376.840.78%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.503.71%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.14%
  • aaveAave(AAVE)$137.847.44%
  • AsterAster(ASTER)$0.774.70%
  • mantleMantle(MNT)$0.6310.20%
  • Pump.funPump.fun(PUMP)$0.0043076.95%
  • OndoOndo(ONDO)$0.3960476.63%
  • polkadotPolkadot(DOT)$1.135.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MoMA: An Open-Vocabulary and Training Free Personalized Image Model that Boasts Flexible Zero-Shot Capabilities

April 12, 2024
in AI & Technology
Reading Time: 4 mins read
A A
MoMA: An Open-Vocabulary and Training Free Personalized Image Model that Boasts Flexible Zero-Shot Capabilities
ShareShareShareShareShare

Modern image-generating tools have come a long way thanks to large-scale text-to-image diffusion models like GLIDE, DALL-E 2, Imagen, Stable Diffusion, and eDiff-I. Thanks to these models, users can create realistic pictures using a variety of textual cues. Kandinsky and Stable Unclip take images as inputs to generate variations that retain the visual components of the reference. The emergence of image-conditioned generation works like Kandinsky, and Stable Unclip is a response to the fact that textual descriptions, while effective, frequently fail to convey detailed visual features.

Image personalization or subject-driven generation is the next logical step in this area. Early attempts in this field include using learnable text tokens to represent target concepts and converting input photos to text. Nevertheless, the substantial resources needed for instance-specific tuning and model storage severely restrict the practicality of these approaches despite their accuracy. To overcome these constraints, tuning-free methods have become more popular. Despite their efficacy in modifying textures, these methods frequently produce tuning-free detail defects and necessitate further tuning to achieve ideal results with target objects.

A recent study by ByteDance and Rutgers University presents a new model called MoMA for picture personalization that does not require tweaking and uses an open vocabulary. It overcomes these issues by effectively integrating logical textual prompts, achieving excellent detail fidelity, and resembling object identities. MoMA for text-to-image diffusion model rapid picture customization.

This approach consists of three parts:

  1. First, the researchers use a generative multimodal decoder to retrieve the reference picture’s features. Then, they alter them according to the target prompt to get the contextualized image feature. 
  2. Meanwhile, they use the original UNet’s self-attention layers to extract the object image feature by replacing the background of the original image with white color and leaving only the object’s pixels. 
  3. Lastly, they used the UNet diffusion model with the object-cross-attention layers and the contextualized picture attributes to generate new images. The layers were trained specifically for this purpose.

The team used the OpenImage-V7 dataset to build a dataset of 282K image/caption/image-mask triplets for model training. After generating image captions using BLIP-2 OPT6.7B, any subjects pertaining to humans, color, form, and texture keywords were eliminated.

The experimental results speak volumes about the MoMA model’s superiority. By harnessing the power of Multimodal Large Language Models (MLLMs), the model seamlessly combines the visual characteristics of the target object with text prompts, enabling changes to both the backdrop context and object texture. The suggested self-attention shortcut significantly enhances detail quality while imposing a minimal computational burden. The model’s expanded applicability is a testament to its potential, as it can be directly integrated with community models that have been fine-tuned using the same basic model, opening up new possibilities in the field of image generation and machine learning. 


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
AI & Technology

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

September 18, 2026
Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright
AI & Technology

Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright

September 18, 2026
Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI
AI & Technology

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI

September 18, 2026
The New Resident Evil Movie Captures The Survival Horror Magic Of The Games
AI & Technology

The New Resident Evil Movie Captures The Survival Horror Magic Of The Games

September 18, 2026
Next Post
Meet the Press NOW — March 18

Meet the Press NOW — March 18

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

September 18, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
An Incremental Price Hike For Incremental Updates

An Incremental Price Hike For Incremental Updates

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!