• bitcoinBitcoin(BTC)$76,832.00-1.60%
  • ethereumEthereum(ETH)$2,445.50-0.44%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$711.73-1.21%
  • rippleXRP(XRP)$1.34-3.36%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$99.28-2.01%
  • tronTRON(TRX)$0.3403840.51%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.041.04%
  • zcashZcash(ZEC)$1,098.35-11.40%
  • HyperliquidHyperliquid(HYPE)$79.49-4.62%
  • dogecoinDogecoin(DOGE)$0.083448-2.68%
  • RainRain(RAIN)$0.015751-0.44%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$509.37-0.22%
  • whitebitWhiteBIT Coin(WBT)$79.49-1.33%
  • chainlinkChainlink(LINK)$11.52-1.87%
  • leo-tokenLEO Token(LEO)$9.190.06%
  • cardanoCardano(ADA)$0.205820-2.53%
  • stellarStellar(XLM)$0.175744-2.70%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$224.28-10.18%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.32-1.25%
  • CantonCanton(CC)$0.098961-4.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.32%
  • uniswapUniswap(UNI)$5.99-2.67%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.075213-1.42%
  • nearNEAR Protocol(NEAR)$2.490.61%
  • avalanche-2Avalanche(AVAX)$7.50-2.93%
  • suiSui(SUI)$0.73-5.05%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.05%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.055950-4.15%
  • tether-goldTether Gold(XAUT)$4,321.79-1.70%
  • MemeCoreMemeCore(M)$1.16-3.90%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$110.47-1.56%
  • BittensorBittensor(TAO)$238.40-5.77%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.09%
  • AsterAster(ASTER)$0.70-4.10%
  • polkadotPolkadot(DOT)$1.11-0.17%
  • mantleMantle(MNT)$0.57-5.06%
  • aaveAave(AAVE)$121.87-2.24%
  • pax-goldPAX Gold(PAXG)$4,323.90-1.73%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0563140.43%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Introduces the ‘ForgetFilter’: A Machine Learning Algorithm that Filters Unsafe Data based on How Strong the Model’s Forgetting Signal is for that Data

December 25, 2023
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper Introduces the ‘ForgetFilter’: A Machine Learning Algorithm that Filters Unsafe Data based on How Strong the Model’s Forgetting Signal is for that Data
ShareShareShareShareShare

A pressing concern has surfaced in large language models (LLMs), drawing attention to the safety implications of downstream customized finetuning. As LLMs become increasingly sophisticated, their potential to inadvertently generate biased, toxic, or harmful outputs poses a substantial challenge. This paper (from a team of researchers from the University of Massachusetts Amherst, Columbia University, Google, Stanford University, and New York University) is a significant contribution to the ongoing discourse surrounding LLM safety, as it meticulously explores the intricate dynamics of these models during the finetuning process.

The predominant approach to aligning LLMs with human preferences within the current milieu involves finetuning. This can be achieved through reinforcement learning from human feedback (RLHF) or traditional supervised learning. The paper introduces a groundbreaking alternative named ForgetFilter, designed to grapple with the inherent complexities of safety finetuning. ForgetFilter represents a paradigm shift by delving into the nuanced behaviors of LLMs, particularly focusing on semantic-level differences and conflicts during the finetuning phase.

ForgetFilter operates by dissecting the forgetting process intrinsic to safety finetuning. Its novel approach involves strategically filtering unsafe examples from noisy downstream data, mitigating the risks associated with biased or harmful model outputs. The paper outlines the key parameters governing ForgetFilter’s effectiveness, offering valuable insights. Notably, the method demonstrates an interesting insensitivity of classification performance to the number of training steps on safe examples. The experimentations reveal that opting for a relatively smaller number of training steps enhances model efficiency and optimizes computational resources.

https://arxiv.org/abs/2312.12736

A critical aspect of ForgetFilter’s success is carefully selecting a threshold for forgetting rates (ϕ). The research underscores that a small ϕ value is effective across diverse scenarios while acknowledging the need for automated approaches to identify an optimal ϕ, especially in scenarios with varying percentages of unsafe examples. Furthermore, the research team delves into the influence of the size of safe examples during safety finetuning on ForgetFilter’s filtering performance. The intriguing finding that reducing the number of safe examples has minimal effects on classification outcomes raises essential considerations for resource-efficient model deployment.

The research extends its investigation into the long-term safety of LLMs, particularly in an “interleaved training” setup involving continuous downstream finetuning followed by safety alignment. This exploration underscores the limitations of safety finetuning in eradicating unsafe knowledge from the model, emphasizing the proactive filtering of unsafe examples as a crucial component for ensuring sustained long-term safety.

Additionally, the research team acknowledges the ethical dimensions of their work. They recognize the potential societal impact of biased or harmful outputs generated by LLMs and stress the importance of mitigating such risks through advanced safety measures. This ethical consciousness adds depth to the paper’s contributions, aligning it with broader discussions on responsible AI development and deployment.

In conclusion, the paper significantly addresses the multifaceted safety challenges in LLMs. ForgetFilter emerges as a promising solution with its nuanced understanding of forgetting behaviors and semantic-level filtering. The study introduces a novel method and prompts future investigations into the factors influencing LLM forgetting behaviors. ForgetFilter signifies a critical step toward the responsible development and deployment of large language models by balancing model utility and safety. As the AI community grapples with these challenges, ForgetFilter offers a valuable contribution to the ongoing dialogue on AI ethics and safety.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

🚨 New Research Alert!
People have found safety training of LLMs can be easily undone through finetuning. How can we ensure safety in customized LLM finetuning while making finetuning still useful? Check out our latest work led by Jiachen Zhao! @jcz12856876
🔍 Our study reveals:… pic.twitter.com/buNu9td5xk

— Mengye Ren (@mengyer) December 21, 2023


YOU MAY ALSO LIKE

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

Madhur Garg is a consulting intern at MarktechPost. He is currently pursuing his B.Tech in Civil and Environmental Engineering from the Indian Institute of Technology (IIT), Patna. He shares a strong passion for Machine Learning and enjoys exploring the latest advancements in technologies and their practical applications. With a keen interest in artificial intelligence and its diverse applications, Madhur is determined to contribute to the field of Data Science and leverage its potential impact in various industries.


🚀 Boost your LinkedIn presence with Taplio: AI-driven content creation, easy scheduling, in-depth analytics, and networking with top creators – Try it free now!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.
AI & Technology

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

September 10, 2026
IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses
AI & Technology

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

September 10, 2026
OpenAI Launches ChatGPT for Financial Services With Built-In Data – Unite.AI
AI & Technology

OpenAI Launches ChatGPT for Financial Services With Built-In Data – Unite.AI

September 10, 2026
Cognition Adds Dioxus Team to Advance Devin Coding Agent – Unite.AI
AI & Technology

Cognition Adds Dioxus Team to Advance Devin Coding Agent – Unite.AI

September 10, 2026
Next Post
Ukraine attacks shipyard in Russia-controlled Crimea with cruise missiles, Moscow says

Ukraine attacks shipyard in Russia-controlled Crimea with cruise missiles, Moscow says

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Proteomic Aging Clocks Track Biological Age Reversal in Rentosertib Trial – Unite.AI

Proteomic Aging Clocks Track Biological Age Reversal in Rentosertib Trial – Unite.AI

September 7, 2026
Acadia Retail Trust has big presence on Bleecker Street

Acadia Retail Trust has big presence on Bleecker Street

September 7, 2026
A major earthquake strikes southern Japan

A major earthquake strikes southern Japan

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!