• bitcoinBitcoin(BTC)$79,732.000.27%
  • ethereumEthereum(ETH)$2,458.930.12%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$765.826.75%
  • rippleXRP(XRP)$1.410.78%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$102.741.13%
  • tronTRON(TRX)$0.3337721.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.30%
  • HyperliquidHyperliquid(HYPE)$85.370.86%
  • zcashZcash(ZEC)$1,016.593.32%
  • dogecoinDogecoin(DOGE)$0.0877113.66%
  • RainRain(RAIN)$0.016425-1.12%
  • moneroMonero(XMR)$542.983.64%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$11.932.16%
  • whitebitWhiteBIT Coin(WBT)$73.250.21%
  • leo-tokenLEO Token(LEO)$9.27-0.22%
  • cardanoCardano(ADA)$0.2171611.55%
  • stellarStellar(XLM)$0.1840122.67%
  • bitcoin-cashBitcoin Cash(BCH)$250.03-0.48%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.1095441.24%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.777.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.424.32%
  • uniswapUniswap(UNI)$6.350.85%
  • hedera-hashgraphHedera(HBAR)$0.0804903.95%
  • suiSui(SUI)$0.806.32%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.532.19%
  • shiba-inuShiba Inu(SHIB)$0.0000054.80%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.1812.16%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056293-0.04%
  • tether-goldTether Gold(XAUT)$4,425.08-0.23%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.120.30%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.425.13%
  • BittensorBittensor(TAO)$235.655.27%
  • AsterAster(ASTER)$0.8211.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • aaveAave(AAVE)$131.620.11%
  • pax-goldPAX Gold(PAXG)$4,432.45-0.20%
  • mantleMantle(MNT)$0.580.54%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056959-0.34%
  • OndoOndo(ONDO)$0.3692794.36%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

UC Berkeley And MIT Researchers Propose A Policy Gradient Algorithm Called Denoising Diffusion Policy Optimization (DDPO) That Can Optimize A Diffusion Model For Downstream Tasks Using Only A Black-Box Reward Function

July 9, 2023
in AI & Technology
Reading Time: 4 mins read
A A
UC Berkeley And MIT Researchers Propose A Policy Gradient Algorithm Called Denoising Diffusion Policy Optimization (DDPO) That Can Optimize A Diffusion Model For Downstream Tasks Using Only A Black-Box Reward Function
ShareShareShareShareShare

Researchers have made notable strides in training diffusion models using reinforcement learning (RL) to enhance prompt-image alignment and optimize various objectives. Introducing denoising diffusion policy optimization (DDPO), which treats denoising diffusion as a multi-step decision-making problem, enables fine-tuning Stable Diffusion on challenging downstream objectives.

By directly training diffusion models on RL-based objectives, the researchers demonstrate significant improvements in prompt-image alignment and optimizing objectives that are difficult to express through traditional prompting methods. DDPO presents a class of policy gradient algorithms designed for this purpose. To improve prompt-image alignment, the research team incorporates feedback from a large vision-language model known as LLaVA. By leveraging RL training, they achieved remarkable progress in aligning prompts with generated images. Notably, the models shift towards a more cartoon-like style, potentially influenced by the prevalence of such representations in the pretraining data.

The results obtained using DDPO for various reward functions are promising. Evaluations on objectives such as compressibility, incompressibility, and aesthetic quality show notable enhancements compared to the base model. The researchers also highlight the generalization capabilities of the RL-trained models, which extend to unseen animals, everyday objects, and novel combinations of activities and objects. While RL training brings substantial benefits, the researchers note the potential challenge of over-optimization. Fine-tuning learned reward functions can lead to models exploiting the rewards non-usefully, often destroying meaningful image content.

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!

Additionally, the researchers observe a susceptibility of the LLaVA model to typographic attacks. RL-trained models can loosely generate text resembling the correct number of animals, fooling LLaVA in prompt-based alignment scenarios.

In summary, introducing DDPO and using RL training for diffusion models represent significant progress in improving prompt-image alignment and optimizing diverse objectives. The results showcase advancements in compressibility, incompressibility, and aesthetic quality. However, challenges such as reward over-optimization and vulnerabilities in prompt-based alignment methods warrant further investigation. These findings open up new opportunities for research and development in diffusion models, particularly in image generation and completion tasks.


Check out the Paper, Project, and GitHub Link. Don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How To Check Your MacBook’s Hard Drive Health

Remote Work As A Worm, Colorful Platformers And Other New Indie Games Worth Checking Out

Niharika is a Technical consulting intern at Marktechpost. She is a third year undergraduate, currently pursuing her B.Tech from Indian Institute of Technology(IIT), Kharagpur. She is a highly enthusiastic individual with a keen interest in Machine learning, Data science and AI and an avid reader of the latest developments in these fields.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Check Your MacBook’s Hard Drive Health
AI & Technology

How To Check Your MacBook’s Hard Drive Health

September 5, 2026
Remote Work As A Worm, Colorful Platformers And Other New Indie Games Worth Checking Out
AI & Technology

Remote Work As A Worm, Colorful Platformers And Other New Indie Games Worth Checking Out

September 5, 2026
OpenAI Plans Misalignment Incident Reporting Framework After Wiki Incident – Unite.AI
AI & Technology

OpenAI Plans Misalignment Incident Reporting Framework After Wiki Incident – Unite.AI

September 5, 2026
Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus
AI & Technology

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

September 5, 2026
Next Post
PVH Beats Earnings Estimates, but Faced Significant Foreign Exchange Headwinds

PVH Beats Earnings Estimates, but Faced Significant Foreign Exchange Headwinds

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Bryan Kohberger seeks to withdraw guilty plea

Bryan Kohberger seeks to withdraw guilty plea

September 4, 2026
Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking

Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking

September 4, 2026
Nvidia Nears  Billion Hugging Face Deal | Bloomberg Tech 09/02/2026

Nvidia Nears $14 Billion Hugging Face Deal | Bloomberg Tech 09/02/2026

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!