• bitcoinBitcoin(BTC)$77,250.001.22%
  • ethereumEthereum(ETH)$2,475.081.80%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$749.953.54%
  • rippleXRP(XRP)$1.321.72%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$104.684.91%
  • tronTRON(TRX)$0.3360860.15%
  • zcashZcash(ZEC)$1,513.9910.77%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.15%
  • HyperliquidHyperliquid(HYPE)$86.739.43%
  • dogecoinDogecoin(DOGE)$0.0842504.00%
  • moneroMonero(XMR)$516.713.52%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$79.561.40%
  • RainRain(RAIN)$0.012683-1.55%
  • chainlinkChainlink(LINK)$11.765.96%
  • leo-tokenLEO Token(LEO)$8.89-0.57%
  • cardanoCardano(ADA)$0.2150579.86%
  • stellarStellar(XLM)$0.1883473.14%
  • uniswapUniswap(UNI)$8.6628.15%
  • bitcoin-cashBitcoin Cash(BCH)$246.8811.95%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.4829.92%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.10805611.42%
  • litecoinLitecoin(LTC)$54.614.92%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.353.22%
  • avalanche-2Avalanche(AVAX)$7.945.17%
  • hedera-hashgraphHedera(HBAR)$0.0771854.19%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.788.82%
  • shiba-inuShiba Inu(SHIB)$0.0000057.54%
  • crypto-com-chainCronos(CRO)$0.0588900.71%
  • MemeCoreMemeCore(M)$1.2713.73%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BittensorBittensor(TAO)$239.386.72%
  • tether-goldTether Gold(XAUT)$4,357.491.52%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$113.952.39%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.04%
  • aaveAave(AAVE)$134.5110.97%
  • AsterAster(ASTER)$0.754.14%
  • mantleMantle(MNT)$0.584.08%
  • Pump.funPump.fun(PUMP)$0.0041077.63%
  • polkadotPolkadot(DOT)$1.1210.33%
  • OndoOndo(ONDO)$0.38900611.26%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Demonstration ITerated Task Optimization (DITTO): A Novel AI Method that Aligns Language Model Outputs Directly with User’s Demonstrated Behaviors

June 7, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Demonstration ITerated Task Optimization (DITTO): A Novel AI Method that Aligns Language Model Outputs Directly with User’s Demonstrated Behaviors
ShareShareShareShareShare

Language models (LMs) are designed to reflect a broad range of voices, leading to outputs that don’t perfectly match any single perspective. To avoid generic responses, one can use LLMs through supervised fine-tuning (SFT) or reinforcement learning with human feedback (RLHF). However, these methods need huge datasets, making them impractical for new and specific tasks. Moreover, there is often a mismatch between the universal style trained into an LLM through instruction and preference tuning needed for specific applications. This mismatch results in LLM outputs feeling generic and lacking a distinctive voice.

Several methods have been developed to address these challenges. One of the approaches involves LLMs and Preference Finetuning in which LLMs are trained on huge datasets to perform well with careful prompting. However, designing prompts can be difficult and sensitive to variations, so it is often necessary to finetune these models on large datasets and use RLHF. Another strategy is self-improvement, where iterative sampling is used to enhance LLMs. For example, methods like STaR are supervised by verifying the correctness of its outputs. Lastly, Online Imitation Learning can improve a policy beyond the demonstrator’s performance. However, these approaches need to learn a reward function and are not applicable to LLMs.

Researchers from Standford University have introduced Demonstration ITerated Task Optimization (DITTO), a method that aligns language model outputs directly with the user’s demonstrated behaviors. It is derived using ideas from online imitation learning and can generate online comparison data at a low cost. To generate these data, DITTO prioritizes users’ demonstrations over output from the LLM and its intermediate checkpoints. Moreover, the win rates of this method outperform few-shot prompting, supervised fine-tuning, and other self-play methods by an average of 19% points. Also, it provides a novel way to effectively customize LLMs using direct feedback from demonstrations.

DITTO is capable of learning fine-grained style and task alignment across domains like news articles, emails, and blog posts. It is an iterative process that contains three components: (a) On the set of expert demonstrations, supervised fine-tuning is executed for a limited number of gradient steps; (b) a New dataset is constructed during the training process by sampling completions for each demonstration and adding it to the ranking over policies,  and (c) RLHF is used for updating the policy, particularly using batches sampled through the previously mentioned process. 

The results of DITTO is evaluated with GPT-4 eval and averaged across all authors, where it outperforms all baselines with an average win rate of 77.09% across CMCC (71.67%) and CCAT50 (82.50%). It provides an average increase of 11.7% win rate as compared to SFT which serves as a strong baseline (56.78% on CMCC, 73.89% on CCAT). Further, in user study results, DITTO outperforms baseline methods with DITTO (72.1% win-rate) > SFT (60.1%) > few-shot (48.1%) > self-prompt (44.2%) > zero-shot (25.0%). Also, self-promoting performs a little worse than giving examples in a few-shot prompt and underperforms DITTO.

In conclusion, researchers from Standford University have introduced Demonstration ITerated Task Optimization (DITTO), a method that aligns language model outputs directly with the user’s demonstrated behaviors and generates online comparison data from demonstrations. In this paper, researchers highlighted the importance of using demonstrations as feedback and proved that even a small number of demonstrated behaviors can provide a strong signal of an individual’s specific preferences. However, other model sizes are not tested by researchers because of computational cost, and additional analysis is needed by the types of preference data needed. So, there is a need for future work in this domain. 


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 43k+ ML SubReddit | Also, check out our AI Events Platform


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Google’s Revamped CC Is An AI Agent For Families And Groups

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Next Post
Following GameStop action during Roaring Kitty livestream

Following GameStop action during Roaring Kitty livestream

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Bill de Blasio says it was ‘painful’ to see Fetterman ‘switch sides’ at GOP convention

Bill de Blasio says it was ‘painful’ to see Fetterman ‘switch sides’ at GOP convention

September 14, 2026
GOP Sen. John Kennedy says the U.S. must ‘choke [Putin] to death’ as Trump admin meets with him

GOP Sen. John Kennedy says the U.S. must ‘choke [Putin] to death’ as Trump admin meets with him

September 16, 2026
Protalix BioTherapeutics, Inc. (PLX) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

Protalix BioTherapeutics, Inc. (PLX) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!