• bitcoinBitcoin(BTC)$86,676.007.10%
  • ethereumEthereum(ETH)$2,771.945.54%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$804.845.16%
  • rippleXRP(XRP)$1.5510.23%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$119.418.61%
  • tronTRON(TRX)$0.3442150.46%
  • zcashZcash(ZEC)$1,458.89-1.72%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • HyperliquidHyperliquid(HYPE)$93.360.70%
  • dogecoinDogecoin(DOGE)$0.09956814.28%
  • moneroMonero(XMR)$596.545.86%
  • whitebitWhiteBIT Coin(WBT)$87.195.52%
  • RainRain(RAIN)$0.013975-0.85%
  • chainlinkChainlink(LINK)$13.165.69%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2450948.05%
  • leo-tokenLEO Token(LEO)$8.940.17%
  • stellarStellar(XLM)$0.21700911.00%
  • uniswapUniswap(UNI)$8.810.40%
  • bitcoin-cashBitcoin Cash(BCH)$268.747.48%
  • nearNEAR Protocol(NEAR)$4.11-1.92%
  • avalanche-2Avalanche(AVAX)$11.310.11%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$62.226.17%
  • daiDai(DAI)$1.00-0.03%
  • CantonCanton(CC)$0.1159957.72%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.0214.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.455.35%
  • hedera-hashgraphHedera(HBAR)$0.0920706.44%
  • shiba-inuShiba Inu(SHIB)$0.0000068.98%
  • BittensorBittensor(TAO)$308.1418.12%
  • MemeCoreMemeCore(M)$1.47-0.33%
  • crypto-com-chainCronos(CRO)$0.06560410.34%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,340.14-0.69%
  • okbOKB(OKB)$123.294.55%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BitwayBitway(BTW)$0.8715.34%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.12%
  • aaveAave(AAVE)$144.285.38%
  • OndoOndo(ONDO)$0.4515665.85%
  • mantleMantle(MNT)$0.657.42%
  • EthenaEthena(ENA)$0.208970-3.86%
  • polkadotPolkadot(DOT)$1.193.25%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing AI Safety and Reliability through Short-Circuiting Techniques

June 11, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Enhancing AI Safety and Reliability through Short-Circuiting Techniques
ShareShareShareShareShare

The vulnerability of AI systems, particularly large language models (LLMs) and multimodal models, to adversarial attacks can lead to harmful outputs. These models are designed to assist and provide helpful responses, but adversaries can manipulate them to produce undesirable or even dangerous outputs. The attacks exploit inherent weaknesses in the models, raising concerns about their safety and reliability. Existing defenses, such as refusal training and adversarial training, have significant limitations, often compromising model performance without effectively preventing harmful outputs.

Current methods to improve AI model alignment and robustness include refusal training and adversarial training. Refusal training teaches models to reject harmful prompts, but sophisticated adversarial attacks often bypass these safeguards. Adversarial training involves exposing models to adversarial examples during training to improve robustness, but this method tends to fail against new, unseen attacks and can degrade the model’s performance.

To address these shortcomings, a team of researchers from Black Swan AI, Carnegie Mellon University, and Center for AI Safety proposes a novel method that involves short-circuiting. Inspired by representation engineering, this approach directly manipulates the internal representations responsible for generating harmful outputs. Instead of focusing on specific attacks or outputs, short-circuiting interrupts the harmful generation process by rerouting the model’s internal states to neutral or refusal states. This method is designed to be attack-agnostic and does not require additional training or fine-tuning, making it more efficient and broadly applicable.

The core of the short-circuiting method is a technique called Representation Rerouting (RR). This technique intervenes in the model’s internal processes, particularly the representations that contribute to harmful outputs. By modifying these internal representations, the method prevents the model from completing harmful actions, even under strong adversarial pressure.

Experimentally, RR was applied to a refusal-trained Llama-3-8B-Instruct model. The results showed a significant reduction in the success rate of adversarial attacks across various benchmarks without sacrificing performance on standard tasks. For instance, the short-circuited model demonstrated lower attack success rates on HarmBench prompts while maintaining high scores on capability benchmarks like MT Bench and MMLU. Additionally, the method proved effective in multimodal settings, improving robustness against image-based attacks and ensuring the model’s harmlessness without impacting its utility.

The short-circuiting method operates by using datasets and loss functions tailored to the task. The training data is divided into two sets: the Short Circuit Set and the Retain Set. The Short Circuit Set contains data that triggers harmful outputs, and the Retain Set includes data that represents safe or desired outputs. The loss functions are designed to adjust the model’s representations to redirect harmful processes to incoherent or refusal states, effectively short-circuiting the harmful outputs.

The problem of AI systems producing harmful outputs due to adversarial attacks is a significant concern. Existing methods like refusal training and adversarial training have limitations that the proposed short-circuiting method aims to overcome. By directly manipulating internal representations, short-circuiting offers a robust, attack-agnostic solution that maintains model performance while significantly enhancing safety and reliability. This approach represents a promising advancement in the development of safer AI systems.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit

No LLM is secure! A year ago, we unveiled the first of many automated jailbreak capable of cracking all major LLMs. 🚨

But there is hope?!

We introduce Short Circuiting: the first alignment technique that is adversarially robust. 🧵

📄 Paper: https://t.co/hY7koqrLyl pic.twitter.com/GKWLLv0fox

— Andy Zou (@andyzou_jiaming) June 8, 2024


YOU MAY ALSO LIKE

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

Shreya Maji is a consulting intern at MarktechPost. She is pursued her B.Tech at the Indian Institute of Technology (IIT), Bhubaneswar. An AI enthusiast, she enjoys staying updated on the latest advancements. Shreya is particularly interested in the real-life applications of cutting-edge technology, especially in the field of data science.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Here’s Why Apple’s Mac Studio Has Become So Expensive
AI & Technology

Here’s Why Apple’s Mac Studio Has Become So Expensive

September 21, 2026
Tesla Will Soon Roll Out FSD Supervised In The Czech Republic
AI & Technology

Tesla Will Soon Roll Out FSD Supervised In The Czech Republic

September 21, 2026
Next Post
Why are more polar bears roaming a tiny Canadian town?

Why are more polar bears roaming a tiny Canadian town?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
An Incremental Price Hike For Incremental Updates

An Incremental Price Hike For Incremental Updates

September 18, 2026
How To Use Meta Display Glasses While Driving With The Audio Only Feature

How To Use Meta Display Glasses While Driving With The Audio Only Feature

September 15, 2026
Lindsay Clancy’s defense attorney requests Trump pardon: Is it legally possible?

Lindsay Clancy’s defense attorney requests Trump pardon: Is it legally possible?

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!