• bitcoinBitcoin(BTC)$84,028.00-0.69%
  • ethereumEthereum(ETH)$2,690.25-0.07%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$777.10-0.06%
  • rippleXRP(XRP)$1.571.00%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$122.264.12%
  • tronTRON(TRX)$0.337881-0.65%
  • zcashZcash(ZEC)$1,554.200.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.54%
  • HyperliquidHyperliquid(HYPE)$92.330.55%
  • dogecoinDogecoin(DOGE)$0.0991333.14%
  • moneroMonero(XMR)$558.97-0.91%
  • chainlinkChainlink(LINK)$13.974.98%
  • whitebitWhiteBIT Coin(WBT)$83.85-0.58%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2583593.59%
  • leo-tokenLEO Token(LEO)$8.81-1.56%
  • RainRain(RAIN)$0.011248-6.53%
  • stellarStellar(XLM)$0.219940-1.23%
  • bitcoin-cashBitcoin Cash(BCH)$341.371.25%
  • nearNEAR Protocol(NEAR)$4.957.54%
  • uniswapUniswap(UNI)$9.665.72%
  • litecoinLitecoin(LTC)$72.231.33%
  • CantonCanton(CC)$0.13130615.42%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • suiSui(SUI)$1.1916.07%
  • avalanche-2Avalanche(AVAX)$10.633.81%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.05%
  • hedera-hashgraphHedera(HBAR)$0.0954582.14%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.453.46%
  • BitwayBitway(BTW)$1.3439.98%
  • BittensorBittensor(TAO)$316.246.38%
  • shiba-inuShiba Inu(SHIB)$0.0000063.03%
  • crypto-com-chainCronos(CRO)$0.0658164.29%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.21-1.25%
  • EthenaEthena(ENA)$0.27266019.82%
  • OndoOndo(ONDO)$0.556.39%
  • tether-goldTether Gold(XAUT)$4,283.520.44%
  • okbOKB(OKB)$121.231.09%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • aaveAave(AAVE)$154.485.29%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • mantleMantle(MNT)$0.67-1.20%
  • polkadotPolkadot(DOT)$1.215.02%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

UC Berkeley And MIT Researchers Propose A Policy Gradient Algorithm Called Denoising Diffusion Policy Optimization (DDPO) That Can Optimize A Diffusion Model For Downstream Tasks Using Only A Black-Box Reward Function

July 9, 2023
in AI & Technology
Reading Time: 4 mins read
A A
UC Berkeley And MIT Researchers Propose A Policy Gradient Algorithm Called Denoising Diffusion Policy Optimization (DDPO) That Can Optimize A Diffusion Model For Downstream Tasks Using Only A Black-Box Reward Function
ShareShareShareShareShare

Researchers have made notable strides in training diffusion models using reinforcement learning (RL) to enhance prompt-image alignment and optimize various objectives. Introducing denoising diffusion policy optimization (DDPO), which treats denoising diffusion as a multi-step decision-making problem, enables fine-tuning Stable Diffusion on challenging downstream objectives.

By directly training diffusion models on RL-based objectives, the researchers demonstrate significant improvements in prompt-image alignment and optimizing objectives that are difficult to express through traditional prompting methods. DDPO presents a class of policy gradient algorithms designed for this purpose. To improve prompt-image alignment, the research team incorporates feedback from a large vision-language model known as LLaVA. By leveraging RL training, they achieved remarkable progress in aligning prompts with generated images. Notably, the models shift towards a more cartoon-like style, potentially influenced by the prevalence of such representations in the pretraining data.

The results obtained using DDPO for various reward functions are promising. Evaluations on objectives such as compressibility, incompressibility, and aesthetic quality show notable enhancements compared to the base model. The researchers also highlight the generalization capabilities of the RL-trained models, which extend to unseen animals, everyday objects, and novel combinations of activities and objects. While RL training brings substantial benefits, the researchers note the potential challenge of over-optimization. Fine-tuning learned reward functions can lead to models exploiting the rewards non-usefully, often destroying meaningful image content.

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!

Additionally, the researchers observe a susceptibility of the LLaVA model to typographic attacks. RL-trained models can loosely generate text resembling the correct number of animals, fooling LLaVA in prompt-based alignment scenarios.

In summary, introducing DDPO and using RL training for diffusion models represent significant progress in improving prompt-image alignment and optimizing diverse objectives. The results showcase advancements in compressibility, incompressibility, and aesthetic quality. However, challenges such as reward over-optimization and vulnerabilities in prompt-based alignment methods warrant further investigation. These findings open up new opportunities for research and development in diffusion models, particularly in image generation and completion tasks.


Check out the Paper, Project, and GitHub Link. Don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How To Stop Meta Training Its AI Models On Your Smart Glasses’ Visual Data

New Mexico Jury Rules Meta Misled State Residents About Data Privacy

Niharika is a Technical consulting intern at Marktechpost. She is a third year undergraduate, currently pursuing her B.Tech from Indian Institute of Technology(IIT), Kharagpur. She is a highly enthusiastic individual with a keen interest in Machine learning, Data science and AI and an avid reader of the latest developments in these fields.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Stop Meta Training Its AI Models On Your Smart Glasses’ Visual Data
AI & Technology

How To Stop Meta Training Its AI Models On Your Smart Glasses’ Visual Data

September 25, 2026
New Mexico Jury Rules Meta Misled State Residents About Data Privacy
AI & Technology

New Mexico Jury Rules Meta Misled State Residents About Data Privacy

September 25, 2026
Apple’s HomePod Mini 2 Will Reportedly Come In New Colors, But Feature A Similar Design
AI & Technology

Apple’s HomePod Mini 2 Will Reportedly Come In New Colors, But Feature A Similar Design

September 25, 2026
Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB
AI & Technology

Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

September 25, 2026
Next Post
PVH Beats Earnings Estimates, but Faced Significant Foreign Exchange Headwinds

PVH Beats Earnings Estimates, but Faced Significant Foreign Exchange Headwinds

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Current with Christine Romans – Sept. 2 | NBC News NOW

Current with Christine Romans – Sept. 2 | NBC News NOW

September 19, 2026
Bear steals a gear bag from a fire station in Colorado

Bear steals a gear bag from a fire station in Colorado

September 23, 2026
Steve Kornacki previews generational change fights in Massachusetts primary

Steve Kornacki previews generational change fights in Massachusetts primary

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!