• bitcoinBitcoin(BTC)$83,838.00-0.62%
  • ethereumEthereum(ETH)$2,681.24-0.24%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$772.62-0.81%
  • rippleXRP(XRP)$1.550.50%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$121.003.41%
  • tronTRON(TRX)$0.337822-0.75%
  • zcashZcash(ZEC)$1,528.04-1.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.85%
  • HyperliquidHyperliquid(HYPE)$91.71-1.45%
  • dogecoinDogecoin(DOGE)$0.0977381.71%
  • moneroMonero(XMR)$552.06-3.27%
  • chainlinkChainlink(LINK)$13.743.51%
  • whitebitWhiteBIT Coin(WBT)$83.67-0.93%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2531451.50%
  • RainRain(RAIN)$0.011950-0.81%
  • leo-tokenLEO Token(LEO)$8.84-0.94%
  • stellarStellar(XLM)$0.2178301.21%
  • bitcoin-cashBitcoin Cash(BCH)$338.85-0.64%
  • nearNEAR Protocol(NEAR)$4.905.59%
  • uniswapUniswap(UNI)$9.442.75%
  • litecoinLitecoin(LTC)$70.78-1.37%
  • CantonCanton(CC)$0.12791112.84%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.1614.37%
  • avalanche-2Avalanche(AVAX)$10.46-1.31%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.03%
  • hedera-hashgraphHedera(HBAR)$0.0941050.05%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.62%
  • BitwayBitway(BTW)$1.3136.17%
  • BittensorBittensor(TAO)$310.404.33%
  • shiba-inuShiba Inu(SHIB)$0.0000060.15%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0649592.72%
  • MemeCoreMemeCore(M)$1.19-2.61%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,286.280.64%
  • EthenaEthena(ENA)$0.26051716.63%
  • OndoOndo(ONDO)$0.542.03%
  • okbOKB(OKB)$120.180.29%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • aaveAave(AAVE)$150.813.92%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • mantleMantle(MNT)$0.66-3.01%
  • polkadotPolkadot(DOT)$1.191.58%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MUSE: A Comprehensive AI Framework for Evaluating Machine Unlearning in Language Models

July 20, 2024
in AI & Technology
Reading Time: 6 mins read
A A
MUSE: A Comprehensive AI Framework for Evaluating Machine Unlearning in Language Models
ShareShareShareShareShare

Language models (LMs) face significant challenges related to privacy and copyright concerns due to their training on vast amounts of text data. The inadvertent inclusion of private and copyrighted content in training datasets has led to legal and ethical issues, including copyright lawsuits and compliance requirements with regulations like GDPR. Data owners increasingly demand the removal of their data from trained models, highlighting the need for effective machine unlearning techniques. These developments have spurred research into methods that can transform existing trained models to behave as if they had never been exposed to certain data, while maintaining overall performance and efficiency.

Researchers have made various attempts to address the challenges of machine unlearning in language models. Exact unlearning methods, which aim to make the unlearned model identical to a model retrained without the forgotten data, have been developed for simple models like SVMs and naive Bayes classifiers. However, these approaches are computationally infeasible for modern large language models.

YOU MAY ALSO LIKE

New Mexico Jury Rules Meta Misled State Residents About Data Privacy

Apple’s HomePod Mini 2 Will Reportedly Come In New Colors, But Feature A Similar Design

Approximate unlearning methods have emerged as more practical alternatives. These include parameter optimization techniques like Gradient Ascent, localization-informed unlearning that targets specific model units, and in-context unlearning that modifies model outputs using external knowledge. Researchers have also explored applying unlearning to specific downstream tasks and for eliminating harmful behaviors in language models.

Evaluation methods for machine unlearning in language models have primarily focused on specific tasks like question answering or sentence completion. Metrics such as familiarity scores and comparisons with retrained models have been used to assess unlearning effectiveness. However, existing evaluations often lack comprehensiveness and fail to adequately address real-world deployment considerations like scalability and sequential unlearning requests.

Researchers from the University of Washington, Princeton University, the University of Southern California, the University of Chicago, and Google Research introduce MUSE (Machine Unlearning Six-Way Evaluation), a comprehensive framework designed to assess the effectiveness of machine unlearning algorithms for language models. This systematic approach evaluates six critical properties that address both data owners’ and model deployers’ requirements for practical unlearning. MUSE examines the ability of unlearning algorithms to remove verbatim memorization, knowledge memorization, and privacy leakage while also assessing their capacity to preserve utility, scale effectively, and sustain performance across multiple unlearning requests. By applying this framework to evaluate eight representative machine unlearning algorithms on datasets focused on unlearning Harry Potter books and news articles, MUSE provides a holistic view of the current state and limitations of unlearning techniques in real-world scenarios.

MUSE proposes a comprehensive set of evaluation metrics that address both data owner and model deployer expectations for machine unlearning in language models. The framework consists of six key criteria:

Data Owner Expectations:

1. No verbatim memorization: Measured by prompting the model with the beginning of a sequence from the forget set and comparing the model’s continuation with the true continuation using ROUGE-L F1 score.

2. No knowledge memorization: Assessed by testing the model’s ability to answer questions derived from the forget set, using ROUGE scores to compare model-generated answers with true answers.

3. No privacy leakage: Evaluated using a membership inference attack (MIA) method to detect if the model retains information indicating that the forget set was part of the training data.

Model Deployer Expectations:

4. Utility preservation: Measured by evaluating the model’s performance on the retain set using the knowledge memorization metric.

5. Scalability: Assessed by examining the model’s performance on forget sets of varying sizes.

6. Sustainability: Analyzed by tracking the model’s performance over sequential unlearning requests.

MUSE evaluates these metrics on two representative datasets: NEWS (BBC news articles) and BOOKS (Harry Potter series), providing a realistic testbed for assessing unlearning algorithms in practical scenarios.

The MUSE framework’s evaluation of eight unlearning methods revealed significant challenges in machine unlearning for language models. While most methods effectively removed verbatim and knowledge memorization, they struggled with privacy leakage, often under- or over-unlearning. All methods significantly degraded model utility, with some rendering models unusable. Scalability issues emerged as forget set sizes increased, and sustainability proved problematic with sequential unlearning requests, leading to progressive performance degradation. These findings underscore the substantial trade-offs and limitations in current unlearning techniques, highlighting the pressing need for more effective and balanced approaches to meet both data owner and deployer expectations.

This research introduces MUSE, a comprehensive machine unlearning evaluation benchmark, assesses six key properties crucial for both data owners and model deployers. The evaluation reveals that while current unlearning methods effectively prevent content memorization, they do so at a substantial cost to model utility on retained data. Also, these methods often result in significant privacy leakage and struggle with scalability and sustainability when handling large-scale content removal or successive unlearning requests. These findings underscore the limitations of existing approaches and emphasize the urgent need for developing more robust and balanced machine unlearning techniques that can better address the complex requirements of real-world applications.


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

New Mexico Jury Rules Meta Misled State Residents About Data Privacy
AI & Technology

New Mexico Jury Rules Meta Misled State Residents About Data Privacy

September 25, 2026
Apple’s HomePod Mini 2 Will Reportedly Come In New Colors, But Feature A Similar Design
AI & Technology

Apple’s HomePod Mini 2 Will Reportedly Come In New Colors, But Feature A Similar Design

September 25, 2026
Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB
AI & Technology

Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

September 25, 2026
Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
AI & Technology

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

September 25, 2026
Next Post
Musk Is Moving SpaceX Headquarters to Texas

Musk Is Moving SpaceX Headquarters to Texas

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

September 25, 2026
Lindsay Clancy jurors deadlocked, judge asks them to continue deliberations

Lindsay Clancy jurors deadlocked, judge asks them to continue deliberations

September 20, 2026
Anthropic Investor Franklin: AI Safety Concerns Won’t Slow Spending

Anthropic Investor Franklin: AI Safety Concerns Won’t Slow Spending

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!