• bitcoinBitcoin(BTC)$77,787.000.67%
  • ethereumEthereum(ETH)$2,516.31-0.18%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$724.18-0.49%
  • rippleXRP(XRP)$1.370.56%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.19-0.76%
  • tronTRON(TRX)$0.338825-0.36%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,109.71-2.11%
  • HyperliquidHyperliquid(HYPE)$79.780.62%
  • dogecoinDogecoin(DOGE)$0.084078-1.02%
  • RainRain(RAIN)$0.015228-3.29%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$522.86-3.54%
  • whitebitWhiteBIT Coin(WBT)$80.610.53%
  • chainlinkChainlink(LINK)$11.41-0.90%
  • leo-tokenLEO Token(LEO)$9.03-0.32%
  • cardanoCardano(ADA)$0.2079540.20%
  • stellarStellar(XLM)$0.1812290.61%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$224.22-0.44%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.481.49%
  • uniswapUniswap(UNI)$6.430.91%
  • CantonCanton(CC)$0.096534-1.08%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.55%
  • hedera-hashgraphHedera(HBAR)$0.0761391.47%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.430.08%
  • nearNEAR Protocol(NEAR)$2.381.30%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.06%
  • suiSui(SUI)$0.72-0.98%
  • crypto-com-chainCronos(CRO)$0.058167-2.83%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,335.97-0.33%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-4.21%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.13-0.34%
  • BittensorBittensor(TAO)$235.380.28%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.13%
  • aaveAave(AAVE)$126.25-0.93%
  • AsterAster(ASTER)$0.700.05%
  • pax-goldPAX Gold(PAXG)$4,338.80-0.39%
  • BitwayBitway(BTW)$0.6824.36%
  • mantleMantle(MNT)$0.560.35%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0569810.04%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Pinterest Researchers Present an Effective Scalable Algorithm to Improve Diffusion Models Using Reinforcement Learning (RL)

February 12, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Pinterest Researchers Present an Effective Scalable Algorithm to Improve Diffusion Models Using Reinforcement Learning (RL)
ShareShareShareShareShare

Diffusion models are a set of generative models that work by adding noise to the training data and then learn to recover the same by reversing the noising process. This process allows these models to achieve state-of-the-art image quality, making them one of the most significant developments in Machine Learning (ML) in the past few years. Their performance, however, is greatly determined by the distribution of the training data (mainly web-scale text-image pairs), which leads to issues like human aesthetic mismatch, biases, and stereotypes.

Previous works focus on using curated datasets or intervening in the sampling process to address the abovementioned issues and achieve controllability. However, these methods affect the sampling time of the model without improving its inherent capabilities. In this work, researchers from Pinterest have proposed a reinforcement learning (RL) framework for fine-tuning diffusion models to achieve results that are more aligned with human preferences.

The proposed framework enables training over millions of prompts across diverse tasks. Moreover, to ensure that the model generates diverse outputs, the researchers used a distribution-based reward function for reinforcement learning fine-tuning. Additionally, the researchers also performed multi-task joint training so that the model is better equipped to deal with a diverse set of objectives simultaneously.

For evaluation, the authors considered three separate reward functions – image composition, human preference, and diversity and fairness. They used the ImageReward model to calculate the human preference score, which was then used as the reward during the model’s training. They also compared their framework with various baseline models such as ReFL, RAFT, DRaFT, etc.

  • They found that their method is generalizable to all the rewards and got the best rank in terms of human preference. They hypothesized that the ReFL model is influenced by the reward hacking problem (the model over-optimizes a single metric at the cost of overall performance). In contrast, their method is much more robust to these effects.
  • The results show that the SDv2 model is biased towards light skin tone for images of dentists and judges, whereas their method has a much more balanced distribution.
  • The proposed framework is also able to tackle the problem of compositionality in diffusion models, i.e., generating different compositions of objects in a scene, and performs much better than the SDv2 model.
  • Lastly, in terms of multi-reward joint optimization, the model outperforms the base models on all three tasks.

In conclusion, to address the issues with the existing diffusion models, the authors of this research paper have introduced a scalable RL training framework that fine-tunes diffusion models to achieve better results. The method performed significantly better than existing models and demonstrated its superiority in generality, robustness, and the ability to generate diverse images. With this work, the authors aim to inspire future research on this topic to further enhance diffusion models’ capabilities and mitigate significant issues like bias and fairness.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Anthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?

Which Is Better For Charging Your MacBook?

I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?
AI & Technology

Anthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?

September 14, 2026
Which Is Better For Charging Your MacBook?
AI & Technology

Which Is Better For Charging Your MacBook?

September 14, 2026
At What Length Do Ethernet Cables Drop To Lower Speeds?
AI & Technology

At What Length Do Ethernet Cables Drop To Lower Speeds?

September 14, 2026
Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI
AI & Technology

Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI

September 13, 2026
Next Post
Israel says two hostages rescued from Gaza in special operation, 128 days after their capture

Israel says two hostages rescued from Gaza in special operation, 128 days after their capture

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
First-ever adult T-Rex footprints discovered in North Dakota

First-ever adult T-Rex footprints discovered in North Dakota

September 13, 2026
Anthropic Caught Scientists Using Claude To Further Biological Weapon Research

Anthropic Caught Scientists Using Claude To Further Biological Weapon Research

September 10, 2026
9/11 victim’s son calls on others to not ‘spread hate’

9/11 victim’s son calls on others to not ‘spread hate’

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!