• bitcoinBitcoin(BTC)$81,181.005.21%
  • ethereumEthereum(ETH)$2,497.034.73%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$724.195.48%
  • rippleXRP(XRP)$1.457.86%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.934.08%
  • tronTRON(TRX)$0.3306051.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.79%
  • HyperliquidHyperliquid(HYPE)$86.936.74%
  • zcashZcash(ZEC)$947.8616.48%
  • dogecoinDogecoin(DOGE)$0.0876007.75%
  • RainRain(RAIN)$0.0171342.44%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$520.203.64%
  • chainlinkChainlink(LINK)$11.796.39%
  • whitebitWhiteBIT Coin(WBT)$73.964.61%
  • leo-tokenLEO Token(LEO)$9.351.17%
  • cardanoCardano(ADA)$0.22070810.26%
  • stellarStellar(XLM)$0.1838105.22%
  • bitcoin-cashBitcoin Cash(BCH)$257.015.54%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.1126812.09%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$51.253.37%
  • uniswapUniswap(UNI)$6.338.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.373.13%
  • hedera-hashgraphHedera(HBAR)$0.0789876.18%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.504.67%
  • suiSui(SUI)$0.784.95%
  • shiba-inuShiba Inu(SHIB)$0.0000054.33%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0577896.62%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,465.541.92%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.964.33%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.04-1.69%
  • okbOKB(OKB)$109.014.43%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.44%
  • BittensorBittensor(TAO)$225.314.15%
  • aaveAave(AAVE)$132.674.61%
  • AsterAster(ASTER)$0.72-0.83%
  • pax-goldPAX Gold(PAXG)$4,474.321.92%
  • mantleMantle(MNT)$0.571.71%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572942.40%
  • OndoOndo(ONDO)$0.3624394.93%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Revolutionizing Scene Reconstruction with Break-A-Scene: The Future of AI-Powered Object Extraction and Remixing

June 5, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Revolutionizing Scene Reconstruction with Break-A-Scene: The Future of AI-Powered Object Extraction and Remixing
ShareShareShareShareShare

Humans naturally possess the ability to break down complicated scenes into component elements and imagine them in various scenarios. One might easily picture the same creature in multiple attitudes and locales or imagine the same bowl in a new environment, given a snapshot of a ceramic artwork showing a creature reclining on a bowl. Today’s generative models, however, need help with tasks of this nature. Recent research suggests personalizing large-scale text-to-image models by optimizing freshly added specialized text embeddings or fine-tuning the model weights, given many pictures of a single idea, to enable synthesizing instances of this concept in unique situations.

In this study, researchers from the Hebrew University of Jerusalem, Google Research, Reichman University and Tel Aviv University present a novel scenario for textual scene decomposition: given a single image of a scene that might include several concepts of various types, their objective is to separate out a specific text token for each idea. This permits the creation of innovative pictures from verbal prompts that highlight certain concepts or combinations of many themes. The ideas they want to learn or extract from the customization activity are only sometimes apparent, which makes it potentially unclear. Previous works have dealt with this ambiguity by focusing on a single topic at a time and using a variety of photographs to show the notion in various settings. However, alternative methods are required to resolve the problem when transitioning to a single-picture situation. 

They specifically suggest adding a series of masks to the input image to add further information about the concepts they want to extract. These masks may be free-form ones that the user supplies or ones produced by an automated segmentation approach (such as). Adapting the two primary techniques, TI and DB, to this environment indicate a reconstruction-editability tradeoff. Whereas TI fails to rebuild the ideas in a new context properly, DB needs more context control due to overfitting. In this study, the authors suggest a unique customization pipeline that successfully strikes a compromise between maintaining learned concept identity and preventing overfitting. 

🚀 JOIN the fastest ML Subreddit Community

Figure 1 provides an overview of our methodology, which has four main parts: (1) We use a union-sampling approach, in which a new subset of the tokens is sampled every time, to train the model to handle various combinations of created ideas. Additionally, (2) in order to prevent overfitting, we employ a two-phase training regime, starting with the optimisation of just the recently inserted tokens with a high learning rate and continuing with the model weights in the second phase with a reduced learning rate. The desired ideas are reconstructed by use of a (3) disguised diffusion loss. Fourth, we employ a unique cross-attention loss to promote disentanglement between the learned ideas.

Their pipeline contains two steps, which are shown in Figure 1. To rebuild the input image, they first identify a group of special text characters (called handles), freeze the model weights, and then optimize the handles. They continue to refine the handles while switching over to fine-tuning the model weights in the second phase. Their method strongly emphasizes disentangling concept extraction or ensuring that each handle is connected to just one target concept. They also understand that the customization procedure cannot be performed independently for each idea to develop graphics showcasing combinations of notions. In response to this discovery, we offer union sampling, a training approach that meets this need and improves the creation of idea combinations. 

They do this by utilizing the masked diffusion loss, a modified variation of the standard diffusion loss. The model is not penalized if a handle is linked to more than one concept because of this loss, which guarantees that each custom handle may deliver its intended idea. Their main finding is that they may punish such entanglement by additionally imposing a loss on the cross-attention maps, which are known to correlate with the scene layout. Due to the additional loss, each handle will concentrate solely on the areas covered by its target concept. They offer several automatic measurements for the task to compare their methodology to the benchmarks. 

They have made the following contributions, in order: (1) they introduce the novel task of textual scene decomposition; (2) they propose a novel method for this situation that strikes a balance between concept fidelity and scene editability by learning a set of disentangled concept handles; and (3) they suggest several automatic evaluation metrics and use them, along with a user study, to demonstrate the effectiveness of their approach. They also conduct user research, which shows that human assessors also like their methodology. In their last part, they suggest several applications for their technique.


Check Out The Paper and Project Page. Don’t forget to join our 23k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Nvidia Deepens Chip Ties With $3.5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026

Anthropic IPO Could Open AI Listings Floodgates: Madrona

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

Nvidia Deepens Chip Ties With .5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026
AI & Technology

Nvidia Deepens Chip Ties With $3.5 Billion MediaTek Bet | Bloomberg Tech 8/31/2026

September 3, 2026
Anthropic IPO Could Open AI Listings Floodgates: Madrona
AI & Technology

Anthropic IPO Could Open AI Listings Floodgates: Madrona

September 3, 2026
Why the US Military Is Betting on Portable Nuclear
AI & Technology

Why the US Military Is Betting on Portable Nuclear

September 3, 2026
AI Boom Fuels a New Tech Debt Binge
AI & Technology

AI Boom Fuels a New Tech Debt Binge

September 3, 2026
Next Post
Gold Off Its Two-year Highs Ahead of Nonfarm Payrolls Data

Gold Off Its Two-year Highs Ahead of Nonfarm Payrolls Data

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
SCOTUS should weigh in on prediction markets, says New Jersey AG who wants to shut down sites like Kalshi, Polymarket

SCOTUS should weigh in on prediction markets, says New Jersey AG who wants to shut down sites like Kalshi, Polymarket

September 2, 2026
Nvidia May Be Close to  Billion Deal for Hugging Face

Nvidia May Be Close to $14 Billion Deal for Hugging Face

September 3, 2026
Police release Nancy Guthrie ransom notes

Police release Nancy Guthrie ransom notes

August 31, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!