• bitcoinBitcoin(BTC)$77,235.000.06%
  • ethereumEthereum(ETH)$2,524.260.49%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$727.180.04%
  • rippleXRP(XRP)$1.370.77%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.78-0.63%
  • tronTRON(TRX)$0.3398660.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.08%
  • zcashZcash(ZEC)$1,124.41-3.29%
  • HyperliquidHyperliquid(HYPE)$79.800.48%
  • dogecoinDogecoin(DOGE)$0.0847940.59%
  • RainRain(RAIN)$0.0157522.27%
  • moneroMonero(XMR)$539.243.69%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.260.17%
  • chainlinkChainlink(LINK)$11.50-0.26%
  • leo-tokenLEO Token(LEO)$9.14-0.15%
  • cardanoCardano(ADA)$0.2072240.53%
  • stellarStellar(XLM)$0.1800321.06%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$226.29-0.64%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.600.57%
  • uniswapUniswap(UNI)$6.356.24%
  • CantonCanton(CC)$0.0980550.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.61%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.40-0.79%
  • hedera-hashgraphHedera(HBAR)$0.0746370.34%
  • shiba-inuShiba Inu(SHIB)$0.0000052.36%
  • nearNEAR Protocol(NEAR)$2.370.25%
  • suiSui(SUI)$0.72-0.02%
  • crypto-com-chainCronos(CRO)$0.0603357.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-0.49%
  • tether-goldTether Gold(XAUT)$4,350.010.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.451.13%
  • BittensorBittensor(TAO)$232.83-1.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • aaveAave(AAVE)$125.450.65%
  • pax-goldPAX Gold(PAXG)$4,354.62-0.01%
  • AsterAster(ASTER)$0.690.70%
  • mantleMantle(MNT)$0.56-3.58%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572113.34%
  • polkadotPolkadot(DOT)$1.02-1.99%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Why LLMs are vulnerable to the ‘butterfly effect’

January 23, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Why LLMs are vulnerable to the ‘butterfly effect’
ShareShareShareShareShare

Prompting is the way we get generative AI and large language models (LLMs) to talk to us. It is an art form in and of itself as we seek to get AI to provide us with ‘accurate’ answers. 

But what about variations? If we construct a prompt a certain way, will it change a model’s decision (and impact its accuracy)? 

YOU MAY ALSO LIKE

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

Diablo V Is Coming Out In Spring 2029

The answer: Yes, according to research from the University of Southern California Information Sciences Institute. 

Even minuscule or seemingly innocuous tweaks — such as adding a space to the beginning of a prompt or giving a directive rather than posing a question — can cause an LLM to change its output. More alarmingly, requesting responses in XML and applying commonly used jailbreaks can have “cataclysmic effects” on data labeled by models. 

Researchers compare this phenomenon to the butterfly effect in chaos theory, which purports that the minor perturbations caused by a butterfly flapping its wings could, several weeks later, cause a tornado in a distant land. 

In prompting, “each step requires a series of decisions from the person designing the prompt,” researchers write. However, “little attention has been paid to how sensitive LLMs are to variations in these decisions.”

Probing ChatGPT with four different prompt methods

The researchers — who were sponsored by the Defense Advanced Research Projects Agency (DARPA) — chose ChatGPT for their experiment and applied four different prompting variation methods. 

The first method asked the LLM for outputs in frequently used formats including Python List, ChatGPT’s JSON Checkbox, CSV, XML or YAML (or the researchers provided no specified format at all). 

The second method applied several minor variations to prompts. These included: 

  • Beginning with a single space. 
  • Ending with a single space. 
  • Starting with ‘Hello’ 
  • Beginning with ‘Hello!’
  • Starting with ‘Howdy!’
  • Ending with ‘Thank you.’
  • Rephrasing from a question to a command. For instance, ‘Which label is best?,’ followed by ‘Select the best label.’

The third method involved applying jailbreak techniques including: 

  • AIM, a top-rated jailbreak that instructs models to simulate a conversation between Niccolo Machiavelli and the character Always Intelligent and Machiavellian (AIM). The model in turn provides responses that are immoral, illegal and/or harmful. 
  • Dev Mode v2, which instructs the model to simulate a ChatGPT with Developer Mode enabled, thus allowing for unrestricted content generation (including that offensive or explicit). 
  • Evil Confidant, which instructs the model to adopt a malignant persona and provide “unhinged results without any remorse or ethics.”
  • Refusal Suppression, which demands prompts under specific linguistic constraints, such as avoiding certain words and constructs. 

The fourth method, meanwhile, involved ‘tipping’ the model — an idea taken from the viral notion that models will provide better prompts when offered money. In this scenario, researchers either added to the end of the prompt, “I won’t tip by the way,” or offered to tip in increments of $1, $10, $100 or $1,000. 

Accuracy drops, predictions change

The researchers ran experiments across 11 classification tasks — true-false and positive-negative question answering; premise-hypothesis relationships; humor and sarcasm detection; reading and math comprehension; grammar acceptability; binary and toxicity classification; and stance detection on controversial subjects. 

With each variation, they measured how often the LLM changed its prediction and what impact that had on its accuracy, then explored the similarity in prompt variations. 

For starters, researchers discovered that simply adding a specified output format yielded a minimum 10% prediction change. Even just utilizing ChatGPT’s JSON Checkbox feature via the ChatGPT API caused more prediction change compared to simply using the JSON specification.

Furthermore, formatting in YAML, XML or CSV led to a 3 to 6% loss in accuracy compared to Python List specification. CSV, for its part, displayed the lowest performance across all formats.

When it came to the perturbation method, meanwhile, rephrasing a statement had the most substantial impact. Also, just introducing a simple space at the beginning of the prompt led to more than 500 prediction changes. This also applies when adding common greetings or ending with a thank-you.

“While the impact of our perturbations is smaller than changing the entire output format, a significant number of predictions still undergo change,” researchers write. 

‘Inherent instability’ in jailbreaks

Similarly, the experiment revealed a “significant” performance drop when using certain jailbreaks. Most notably, AIM and Dev Mode V2 yielded invalid responses in about 90% of predictions. This, researchers noted, is primarily due to the model’s standard response of ‘I’m sorry, I cannot comply with that request.’

Meanwhile, Refusal Suppression and Evil Confidant usage resulted in more than 2,500 prediction changes. Evil Confidant (guided toward ‘unhinged’ responses) yielded low accuracy, while Refusal Suppression alone leads to a loss of more than 10% accuracy, “highlighting the inherent instability even in seemingly innocuous jailbreaks,” researchers emphasize.

Finally (at least for now), models don’t seem to be easily swayed by money, the study found.

“When it comes to influencing the model by specifying a tip versus specifying we will not tip, we noticed minimal performance changes,” researchers write. 

LLMs are young; there’s much more work to be done

But why do slight changes in prompts lead to such significant changes? Researchers are still puzzled. 

They questioned whether the instances that changed the most were ‘confusing’ the model — confusion referring to the Shannon entropy, which measures the uncertainty in random processes.

To measure this confusion, they focused on a subset of tasks that had individual human annotations, and then studied the correlation between confusion and the instance’s likelihood of having its answer changed. Through this analysis, they found that this was “not really” the case.

“The confusion of the instance provides some explanatory power for why the prediction changes,” researchers report, “but there are other factors at play.”

Clearly, there is still much more work to be done. The obvious “major next step” would be to generate LLMs that are resistant to changes and provide consistent answers, researchers note. This requires a deeper understanding of why responses change under minor tweaks and developing ways to better anticipate them. 

As researchers write: “This analysis becomes increasingly crucial as ChatGPT and other large language models are integrated into systems at scale.”

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait
AI & Technology

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

September 12, 2026
Diablo V Is Coming Out In Spring 2029
AI & Technology

Diablo V Is Coming Out In Spring 2029

September 12, 2026
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
AI & Technology

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

September 12, 2026
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI
AI & Technology

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

September 12, 2026
Next Post
Am I Being Too Serious About Paying Off Debt?

Am I Being Too Serious About Paying Off Debt?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Voya Infrastructure, Industrials And Materials Fund Q2 2026 Commentary

Voya Infrastructure, Industrials And Materials Fund Q2 2026 Commentary

September 6, 2026
CPI REPORT IS COMING: Do This Before Market Open!

CPI REPORT IS COMING: Do This Before Market Open!

September 11, 2026
Iran-backed rebels in Yemen tried using Anthropic’s Claude to build guided missiles, alarming report finds

Iran-backed rebels in Yemen tried using Anthropic’s Claude to build guided missiles, alarming report finds

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!