• bitcoinBitcoin(BTC)$76,905.00-1.07%
  • ethereumEthereum(ETH)$2,475.66-1.27%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$718.32-0.45%
  • rippleXRP(XRP)$1.400.23%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$100.78-0.55%
  • tronTRON(TRX)$0.338672-0.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.66%
  • zcashZcash(ZEC)$1,130.49-0.15%
  • HyperliquidHyperliquid(HYPE)$78.97-1.33%
  • dogecoinDogecoin(DOGE)$0.082462-1.88%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$514.810.48%
  • whitebitWhiteBIT Coin(WBT)$79.53-1.19%
  • RainRain(RAIN)$0.013206-12.23%
  • chainlinkChainlink(LINK)$11.37-0.07%
  • leo-tokenLEO Token(LEO)$8.980.21%
  • cardanoCardano(ADA)$0.204320-2.53%
  • stellarStellar(XLM)$0.1937502.47%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$221.81-0.51%
  • USD1USD1(USD1)$1.00-0.01%
  • uniswapUniswap(UNI)$6.706.11%
  • litecoinLitecoin(LTC)$52.34-2.18%
  • CantonCanton(CC)$0.095052-1.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-0.47%
  • hedera-hashgraphHedera(HBAR)$0.0770850.53%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.470.20%
  • nearNEAR Protocol(NEAR)$2.37-1.94%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.74%
  • suiSui(SUI)$0.71-2.31%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057366-3.07%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,285.14-0.11%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$224.63-4.30%
  • MemeCoreMemeCore(M)$1.11-1.18%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$112.53-1.13%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.34%
  • BitwayBitway(BTW)$0.7311.33%
  • aaveAave(AAVE)$127.140.64%
  • AsterAster(ASTER)$0.69-1.72%
  • pax-goldPAX Gold(PAXG)$4,286.57-0.17%
  • mantleMantle(MNT)$0.56-1.76%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057121-0.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft AI Introduces Direct Nash Optimization (DNO): A Scalable Machine Learning Algorithm that Combines the Simplicity and Stability of Contrastive Learning with the Theoretical Generality of Optimizing General Preferences

April 10, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft AI Introduces Direct Nash Optimization (DNO): A Scalable Machine Learning Algorithm that Combines the Simplicity and Stability of Contrastive Learning with the Theoretical Generality of Optimizing General Preferences
ShareShareShareShareShare

The evolution of artificial intelligence through the development of Large Language Models (LLMs) has marked a significant milestone in the quest to mirror human-like abilities in generating text, reasoning, and decision-making. However, aligning these models with human ethics and values has remained complex. Traditional methods, such as Reinforcement Learning from Human Feedback (RLHF), have made strides in integrating human preferences by fine-tuning LLMs post-training. These methods, however, often rely on simplifying the multifaceted nature of human preferences into scalar rewards, a process that may not capture the entirety of human values and ethical considerations.

Researchers from Microsoft Research have introduced an approach known as Direct Nash Optimization (DNO), a novel strategy aimed at refining LLMs by focusing on general preferences rather than solely on reward maximization. The method emerges as a response to the limitations of traditional RLHF techniques, which, despite their advances, struggle to fully embody complex human preferences within the full training of LLMs. DNO introduces a paradigm shift by employing a batched on-policy algorithm alongside a regression-based learning objective.

DNO is rooted in the observation that existing methods might not fully harness the potential of LLMs to understand and generate content that aligns with nuanced human values. DNO offers a comprehensive framework for post-training LLMs by directly optimising general preferences. This approach is characterized by its simplicity and scalability, attributed to the method’s innovative use of batched on-policy updates and regression-based objectives. These features allow DNO to provide a more refined alignment of LLMs with human values, as demonstrated in extensive empirical evaluations.

One of DNO’s standout achievements is its implementation with the 7B parameter Orca-2.5 model, which showed an unprecedented 33% win rate against GPT-4-Turbo in AlpacaEval 2.0. This represents a significant leap from the model’s initial 7% win rate, showcasing an absolute gain of 26% through the application of DNO. This remarkable performance positions DNO as a leading method for post-training LLMs. It highlights its potential to surpass traditional models and methodologies in aligning LLMs more closely with human preferences and ethical standards.

Research Snapshot

In conclusion, the DNO method emerges as a pivotal advancement in refining LLMs, addressing the significant challenge of aligning these models with human ethical standards and complex preferences. By shifting focus from traditional reward maximization to optimizing general preferences, DNO overcomes the limitations of previous RLHF techniques and sets a new benchmark for post-training LLMs. The remarkable success demonstrated by the Orca-2.5 model’s impressive performance gain in AlpacaEval 2.0 underscores its potential to revolutionize the field.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus
AI & Technology

Apple TV Cleaned Up At The Emmys With Eight Wins For Widow’s Bay And Pluribus

September 15, 2026
Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI
AI & Technology

Elsevier Integrates LG AI Research’s Chemistry Vision Model Into Reaxys – Unite.AI

September 15, 2026
Double The Range And Smarter Safety, Too
AI & Technology

Double The Range And Smarter Safety, Too

September 15, 2026
Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second
AI & Technology

Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second

September 15, 2026
Next Post
McDonald’s  ‘deal’ goes viral, sparks debate over California’s minimum wage increase

McDonald's $25 'deal' goes viral, sparks debate over California's minimum wage increase

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
You Can Use Gemini To Help You Organize Your Files On Google Drive

You Can Use Gemini To Help You Organize Your Files On Google Drive

September 14, 2026
More AI researchers warn of AI’s threat to humanity

More AI researchers warn of AI’s threat to humanity

September 13, 2026
Vance, Trump issue plea for Americans to vote Republican on last night of convention – NPR

Vance, Trump issue plea for Americans to vote Republican on last night of convention – NPR

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!