• bitcoinBitcoin(BTC)$76,741.00-0.76%
  • ethereumEthereum(ETH)$2,479.26-2.22%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$716.07-2.70%
  • rippleXRP(XRP)$1.34-2.17%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.72-2.25%
  • tronTRON(TRX)$0.3409780.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,095.61-4.71%
  • HyperliquidHyperliquid(HYPE)$77.40-3.61%
  • dogecoinDogecoin(DOGE)$0.083500-1.85%
  • RainRain(RAIN)$0.0153381.49%
  • moneroMonero(XMR)$538.110.88%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.59-0.99%
  • chainlinkChainlink(LINK)$11.29-2.28%
  • leo-tokenLEO Token(LEO)$9.06-0.60%
  • cardanoCardano(ADA)$0.204574-2.23%
  • stellarStellar(XLM)$0.178346-2.15%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.38-3.29%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.78-0.51%
  • uniswapUniswap(UNI)$6.27-1.50%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.89%
  • CantonCanton(CC)$0.094908-4.09%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0752190.72%
  • avalanche-2Avalanche(AVAX)$7.32-1.69%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.20%
  • nearNEAR Protocol(NEAR)$2.30-2.81%
  • suiSui(SUI)$0.71-2.52%
  • crypto-com-chainCronos(CRO)$0.0588291.36%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.77-0.04%
  • MemeCoreMemeCore(M)$1.15-2.38%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.49-0.68%
  • BittensorBittensor(TAO)$232.50-1.70%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.06%
  • aaveAave(AAVE)$124.15-2.92%
  • pax-goldPAX Gold(PAXG)$4,351.93-0.04%
  • AsterAster(ASTER)$0.690.70%
  • mantleMantle(MNT)$0.56-2.86%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0576301.03%
  • BitwayBitway(BTW)$0.6719.29%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research Uncovers the Mechanics of Dishonesty in Large Language Models: A Deep Dive into Prompt Engineering and Neural Network Analysis

December 7, 2023
in AI & Technology
Reading Time: 5 mins read
A A
This AI Research Uncovers the Mechanics of Dishonesty in Large Language Models: A Deep Dive into Prompt Engineering and Neural Network Analysis
ShareShareShareShareShare

Understanding large language models (LLMs) and promoting their honest conduct has become increasingly crucial as these models have demonstrated growing capabilities and started widely adopted by society. Researchers contend that new risks, such as scalable disinformation, manipulation, fraud, election tampering, or the speculative risk of loss of control, arise from the potential for models to be deceptive (which they define as “the systematic inducement of false beliefs in the pursuit of some outcome other than the truth”). Research indicates that even while the models’ activations have the necessary information, they may need more than misalignment to produce the right result. 

Previous studies have distinguished between truthfulness and honesty, saying that the former refrains from making false claims, while the latter refrains from making claims it does not “believe.” This distinction helps to make sense of it. Therefore, a model may generate misleading assertions owing to misalignment in the form of dishonesty rather than a lack of skill. Since then, several studies have tried to address LLM honesty by delving into a model’s internal state to find truthful representations. Proposals for recent black box techniques have also been made to identify and provoke massive language model lying. Notably, previous work demonstrates that improving the extraction of internal model representations may be achieved by forcing models to consider a notion actively. 

Furthermore, models include a “critical” intermediary layer in context-following environments, beyond which representations of true or incorrect responses in context-following tend to diverge a phenomenon known as “overthinking.” Motivated by previous studies, the researchers broadened the focus from incorrectly labeled in-context learning to deliberate dishonesty, in which they gave the model explicit instructions to lie. Using probing and mechanical interpretability methodologies, the research team from Cornell University, the University of Pennsylvania, and the University of Maryland hopes to identify and comprehend which layers and attention heads in the model are accountable for dishonesty in this context. 

The following are their contributions: 

1. The research team shows that, as determined by considerably below-chance accuracy on true/false questions, LLaMA-2-70b-chat can be trained to lie. According to the study team, this can be quite delicate and has to be carefully and quickly engineered. 

2. Using activation patching and probing, the research team finds independent evidence for five model layers critical to dishonest conduct. 

3. Only 46 attention heads, or 0.9% of all heads in the network, were effectively subjected to causal interventions by the study team, which forced deceptive models to respond truthfully. These treatments are resilient over several dataset splits and prompts. 

In a nutshell the research team looks at a straightforward case of lying, where they provide LLM instructions on whether to tell the truth or not. Their findings demonstrate that huge models can display dishonest behaviour, producing right answers when asked to be honest and erroneous responses if pushed to lie. These findings build on earlier research that suggests activation probing can generalize out-of-distribution when prompted. However, the research team does discover that this may necessitate lengthy prompt engineering due to problems like the model’s tendency to output the “False” token sooner in the sequence than the “True” token. 

By using prefix injection, the research team can consistently induce lying. Subsequently, the team compares the activations of the dishonest and honest models, localizing the layers and attention heads involved in lying. By employing linear probes to investigate this lying behavior, the research team discovers that early-to-middle layers see comparable model representations for honest and liar prompts before diverging drastically to become anti-parallel. This might show that prior layers should have a context-invariant representation of truth, as desired by a body of literature. Activation patching is another tool the research team uses to understand more about the workings of specific layers and heads. The researchers discovered that localized interventions could completely address the mismatch between the honest-prompted and liar models in either direction. 

Significantly, these interventions on a mere 46 attention heads demonstrate a solid degree of cross-dataset and cross-prompt resilience. The research team focuses on lying by utilizing an accessible dataset and specifically telling the model to lie, in contrast to earlier work that has largely examined the accuracy and integrity of models that are honest by default. Thanks to this context, researchers have learned a great deal about the subtleties of encouraging dishonest conduct and the methods by which big models engage in dishonest behavior. To guarantee the ethical and safe application of LLMs in the real world, the research team hopes that more work in this context will lead to new approaches to stopping LLM lying.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


✅ [Featured AI Model] Check out LLMWare and It’s RAG- specialized 7B Parameter LLMs

Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Why Do Routers Have So Many Antennas?
AI & Technology

Why Do Routers Have So Many Antennas?

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
Next Post
A Choice Apart: With Guests Max Bazerman & Vivienne Wagner

A Choice Apart: With Guests Max Bazerman & Vivienne Wagner

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Hunter Biden launches crypto meme coin making light of laptop scandal

Hunter Biden launches crypto meme coin making light of laptop scandal

September 7, 2026
Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

September 9, 2026
Cardiff Oncology, Inc. (CRDF) Presents at 8th Annual RAS-Targeted Drug Development Summit – Slideshow (NASDAQ:CRDF) 2026-09-10

Cardiff Oncology, Inc. (CRDF) Presents at 8th Annual RAS-Targeted Drug Development Summit – Slideshow (NASDAQ:CRDF) 2026-09-10

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!