• bitcoinBitcoin(BTC)$76,837.00-0.66%
  • ethereumEthereum(ETH)$2,478.23-2.57%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$717.74-2.52%
  • rippleXRP(XRP)$1.34-1.91%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.07-1.85%
  • tronTRON(TRX)$0.3407660.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,093.06-5.07%
  • HyperliquidHyperliquid(HYPE)$78.04-2.65%
  • dogecoinDogecoin(DOGE)$0.083397-1.83%
  • RainRain(RAIN)$0.0152621.12%
  • moneroMonero(XMR)$530.890.17%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$79.64-0.99%
  • chainlinkChainlink(LINK)$11.30-2.19%
  • leo-tokenLEO Token(LEO)$9.06-0.66%
  • cardanoCardano(ADA)$0.206558-0.85%
  • stellarStellar(XLM)$0.179506-0.84%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$224.92-2.49%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.69-0.30%
  • uniswapUniswap(UNI)$6.27-1.76%
  • CantonCanton(CC)$0.095175-2.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.59%
  • hedera-hashgraphHedera(HBAR)$0.0758791.86%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.37-0.85%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.57%
  • nearNEAR Protocol(NEAR)$2.30-2.34%
  • suiSui(SUI)$0.71-1.79%
  • crypto-com-chainCronos(CRO)$0.0583580.13%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.15-0.05%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-2.77%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.50-1.44%
  • BittensorBittensor(TAO)$236.220.71%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.17%
  • aaveAave(AAVE)$126.970.84%
  • AsterAster(ASTER)$0.701.88%
  • pax-goldPAX Gold(PAXG)$4,349.84-0.10%
  • mantleMantle(MNT)$0.57-1.04%
  • BitwayBitway(BTW)$0.6823.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572770.57%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers at Stanford Unveil C3PO: A Novel Machine Learning Approach for Context-Sensitive Customization of Large Language Models

February 27, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Researchers at Stanford Unveil C3PO: A Novel Machine Learning Approach for Context-Sensitive Customization of Large Language Models
ShareShareShareShareShare

In the evolving landscape of artificial intelligence, language models transform interaction and information processing. However, aligning these models with specific user feedback while avoiding unintended overgeneralization poses a challenge. Traditional approaches often need to discern the applicability of feedback, leading to models extending rules beyond intended contexts. This issue highlights the need for advanced methods to ensure language models can adapt precisely to user preferences without compromising their utility in diverse applications.

Existing works have explored improving language or dialogue systems through various types of feedback, including learned or heuristic rewards, preferences or rankings, and natural language feedback. Natural language feedback has enhanced performance in code generation, dialogue, and summarization tasks. Some studies have focused on leveraging natural language feedback to refine general model behaviors rather than improving a single model output. Related research areas include constitutional AI, context distillation, model editing, and debiasing LLMs.

Researchers from Cornell University have introduced a novel method, Contextualized Critiques with Constrained Preference Optimization (C3PO), to refine models’ response behavior. The C3PO method strategically fine-tunes language models to apply feedback where relevant while averting overgeneralization meticulously. It achieves this by utilizing Direct Preference Optimization (DPO) for data deemed in-scope and Supervised Fine-Tuning (SFT) losses for out-of-scope and near-scope data, ensuring the model’s performance remains robust across various contexts. 

The generation of datasets Dnear-scope and Dout-of-scope, filled with prompts and completions from the initial model, maintains the model’s integrity for inputs unrelated to the feedback. Incorporating a sophisticated combined loss function, LC3PO, the approach not only embraces feedback for pertinent prompts but also actively prevents the model’s performance from deteriorating on irrelevant prompts. This is further enhanced by C3PO’s creation of synthetic two-policy preference data, which enables learning of the optimal policy under the Bradley-Terry preference model framework. This optimal policy delicately balances the model’s original capabilities with the new feedback, penalizing responses that deviate from the input, thus refining the model’s responses precisely, feedback-aligned.

The experiments rigorously evaluate C3PO’s ability to incorporate verbal feedback without overgeneralizing, comparing it against traditional methods and exploring its proficiency in assimilating multiple feedbacks. Utilizing a feedback dataset of 100 entries, both authored and GPT-4 generated, C3PO demonstrates superior performance by effectively adhering to in-scope prompts while minimizing overgeneralization, a notable improvement over modified In-Context and SCD methods. Mixing Learned Low-Rank Adjustment (LoRA) parameters underscores C3PO’s efficient feedback integration, supported by a strategic constraint formulation that outperforms full knowledge distillation.

In conclusion, the development of C3PO marks a significant stride towards more adaptable and user-centric language models. By addressing the challenge of overgeneralization, this method paves the way for more personalized and efficient AI tools tailored to meet the diverse needs of users without sacrificing broader applicability. The implications of this research extend beyond technical achievements, heralding a future where AI can seamlessly adapt to individual preferences, enhancing both its utility and accessibility.


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 38k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel

You may also like our FREE AI Courses….


YOU MAY ALSO LIKE

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Get Your Cut Of PlayStation’s .85 Million Settlement
AI & Technology

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

September 13, 2026
AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
AI & Technology

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Next Post
THIS is Crazy!!! Juggernaut XL Lightning in only 4 Steps – Automatic 1111

THIS is Crazy!!! Juggernaut XL Lightning in only 4 Steps - Automatic 1111

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Harvard fellow ‘deeply concerned’ over efforts to build bioweapons after Anthropic report

Harvard fellow ‘deeply concerned’ over efforts to build bioweapons after Anthropic report

September 12, 2026
Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours

Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours

September 6, 2026
Plane carrying Zelensky threatened by drone in Moldova, Norway’s premier says – The Washington Post

Plane carrying Zelensky threatened by drone in Moldova, Norway’s premier says – The Washington Post

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!