• bitcoinBitcoin(BTC)$80,401.00-0.97%
  • ethereumEthereum(ETH)$2,580.45-2.00%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$748.37-1.77%
  • rippleXRP(XRP)$1.38-2.29%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$108.56-3.00%
  • tronTRON(TRX)$0.3424481.38%
  • zcashZcash(ZEC)$1,447.88-7.56%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.31%
  • HyperliquidHyperliquid(HYPE)$91.16-2.07%
  • dogecoinDogecoin(DOGE)$0.085095-2.30%
  • moneroMonero(XMR)$523.53-8.63%
  • whitebitWhiteBIT Coin(WBT)$81.82-1.75%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013394-4.24%
  • chainlinkChainlink(LINK)$12.02-2.88%
  • cardanoCardano(ADA)$0.220568-1.11%
  • leo-tokenLEO Token(LEO)$8.900.16%
  • stellarStellar(XLM)$0.190600-1.17%
  • uniswapUniswap(UNI)$8.75-4.72%
  • bitcoin-cashBitcoin Cash(BCH)$247.150.14%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.01%
  • nearNEAR Protocol(NEAR)$3.52-4.53%
  • litecoinLitecoin(LTC)$56.95-0.30%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6812.19%
  • CantonCanton(CC)$0.104150-5.17%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.50%
  • hedera-hashgraphHedera(HBAR)$0.0818113.81%
  • suiSui(SUI)$0.821.01%
  • MemeCoreMemeCore(M)$1.4713.48%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.46%
  • crypto-com-chainCronos(CRO)$0.058362-1.31%
  • BittensorBittensor(TAO)$253.04-0.34%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,370.18-0.02%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.57-0.91%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.19%
  • aaveAave(AAVE)$137.16-4.47%
  • AsterAster(ASTER)$0.74-1.68%
  • EthenaEthena(ENA)$0.1989029.24%
  • OndoOndo(ONDO)$0.4100863.17%
  • mantleMantle(MNT)$0.59-2.10%
  • pax-goldPAX Gold(PAXG)$4,360.58-0.06%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Orthogonal Paths: Simplifying Jailbreaks in Language Models

June 23, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Orthogonal Paths: Simplifying Jailbreaks in Language Models
ShareShareShareShareShare

Ensuring the safety and ethical behavior of large language models (LLMs) in responding to user queries is of paramount importance. Problems arise from the fact that LLMs are designed to generate text based on user input, which can sometimes lead to harmful or offensive content. This paper investigates the mechanisms by which LLMs refuse to generate certain types of content and develops methods to improve their refusal capabilities.

Currently, LLMs use various methods to refuse user requests, such as inserting refusal phrases or using specific templates. However, these methods are often ineffective and can be bypassed by users who attempt to manipulate the models. The proposed solution by the researchers from ETH Zürich, Anthropic, MIT and others involve a novel approach called “weight orthogonalization,” which ablates the refusal direction in the model’s weights. This method is designed to make the refusal more robust and difficult to bypass.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

How To Record Audio On Your iPhone

Weight orthogonalization technique is simpler and more efficient than existing methods as it does not require gradient-based optimization or a dataset of harmful completions. The weight orthogonalization method involves adjusting the weights in the model so that the direction associated with refusals is orthogonalized, effectively preventing the model from following refusal directives while maintaining its original capabilities. It is based on the concept of directional ablation, an inference-time intervention where the component corresponding to the refusal direction is zeroed out in the model’s residual stream activations. In this approach, the researchers modify the weights directly to achieve the same effect. 

By orthogonalizing matrices like the embedding matrix, positional embedding matrix, attention-out matrices, and MLP out matrices, the model is prevented from writing to the refusal direction in the first place. This modification ensures the model retains its original capabilities while no longer adhering to the refusal mechanism. 

Performance evaluations of this method, conducted using the HARMBENCH test set, show promising results. The attack success rate (ASR) of the orthogonalized models indicates that this method is on par with prompt-specific jailbreak techniques, like GCG, which optimize jailbreaks for individual prompts. The weight orthogonalization method demonstrates high ASR across various models, including the LLAMA-2 and QWEN families, even when the system prompts are designed to enforce safety and ethical guidelines.

While the proposed method significantly simplifies the process of jailbreaking LLMs, it also raises important ethical considerations. The researchers acknowledge that this method marginally lowers the barrier for jailbreaking open-source model weights, potentially enabling misuse. However, they argue that it does not substantially alter the risk profile of open-sourcing models. The work underscores the fragility of current safety mechanisms and calls for a scientific consensus on the limitations of these techniques to inform future policy decisions and research efforts.

This research highlights a critical vulnerability in the safety mechanisms of LLMs and introduces an efficient method to exploit this weakness. The researchers demonstrate a simple yet powerful technique to bypass refusal mechanisms by orthogonalizing the refusal direction in the model’s weights. This work not only advances the understanding of LLM vulnerabilities but also emphasizes the need for robust and effective safety measures to prevent misuse.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


Shreya Maji is a consulting intern at MarktechPost. She is pursued her B.Tech at the Indian Institute of Technology (IIT), Bhubaneswar. An AI enthusiast, she enjoys staying updated on the latest advancements. Shreya is particularly interested in the real-life applications of cutting-edge technology, especially in the field of data science.

[Announcing Gretel Navigator] Create, edit, and augment tabular data with the first compound AI system trusted by EY, Databricks, Google, and Microsoft


Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
How To Record Audio On Your iPhone
AI & Technology

How To Record Audio On Your iPhone

September 20, 2026
What Is The Difference Between Apple CarPlay And CarPlay Ultra?
AI & Technology

What Is The Difference Between Apple CarPlay And CarPlay Ultra?

September 19, 2026
The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers
AI & Technology

The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers

September 19, 2026
Next Post
Dozens killed in Israeli strikes across northern Gaza amid continued West Bank violence

Dozens killed in Israeli strikes across northern Gaza amid continued West Bank violence

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Kylian Mbappe leaves Nike for Roger Federer-backed On

Kylian Mbappe leaves Nike for Roger Federer-backed On

September 18, 2026
Meet the Press NOW — September 4

Meet the Press NOW — September 4

September 17, 2026
Anne Thompson recalls reporting near ground zero on 9/11

Anne Thompson recalls reporting near ground zero on 9/11

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!