• bitcoinBitcoin(BTC)$81,150.004.94%
  • ethereumEthereum(ETH)$2,621.515.90%
  • tetherTether(USDT)$1.000.05%
  • binancecoinBNB(BNB)$759.851.14%
  • rippleXRP(XRP)$1.416.81%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$113.338.82%
  • tronTRON(TRX)$0.3388650.90%
  • zcashZcash(ZEC)$1,533.821.89%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.41%
  • HyperliquidHyperliquid(HYPE)$93.708.15%
  • dogecoinDogecoin(DOGE)$0.0879564.41%
  • moneroMonero(XMR)$568.3110.23%
  • whitebitWhiteBIT Coin(WBT)$83.224.52%
  • USDSUSDS(USDS)$1.000.02%
  • RainRain(RAIN)$0.0134185.83%
  • chainlinkChainlink(LINK)$12.385.52%
  • cardanoCardano(ADA)$0.2316737.31%
  • leo-tokenLEO Token(LEO)$8.910.25%
  • stellarStellar(XLM)$0.1946083.13%
  • uniswapUniswap(UNI)$9.022.37%
  • bitcoin-cashBitcoin Cash(BCH)$247.920.23%
  • nearNEAR Protocol(NEAR)$3.7910.16%
  • Ethena USDeEthena USDe(USDE)$1.000.06%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$58.717.10%
  • CantonCanton(CC)$0.1128354.40%
  • USD1USD1(USD1)$1.000.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.50%
  • avalanche-2Avalanche(AVAX)$8.476.79%
  • hedera-hashgraphHedera(HBAR)$0.0794132.85%
  • suiSui(SUI)$0.825.58%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000051.44%
  • crypto-com-chainCronos(CRO)$0.0594921.44%
  • MemeCoreMemeCore(M)$1.294.49%
  • BittensorBittensor(TAO)$257.667.19%
  • paypal-usdPayPal USD(PYUSD)$1.000.04%
  • tether-goldTether Gold(XAUT)$4,374.680.40%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.862.33%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • aaveAave(AAVE)$143.436.85%
  • AsterAster(ASTER)$0.783.17%
  • mantleMantle(MNT)$0.627.67%
  • Pump.funPump.fun(PUMP)$0.0042834.13%
  • OndoOndo(ONDO)$0.4061933.87%
  • polkadotPolkadot(DOT)$1.141.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Anthropic Introduces Attribution Graphs: A New Interpretability Method to Trace Internal Reasoning in Claude 3.5 Haiku

April 6, 2025
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper from Anthropic Introduces Attribution Graphs: A New Interpretability Method to Trace Internal Reasoning in Claude 3.5 Haiku
ShareShareShareShareShare

While the outputs of large language models (LLMs) appear coherent and useful, the underlying mechanisms guiding these behaviors remain largely unknown. As these models are increasingly deployed in sensitive and high-stakes environments, it has become crucial to understand what they do and how they do it.

The main challenge lies in uncovering the internal steps that lead a model to a specific response. The computations happen across hundreds of layers and billions of parameters, making it difficult to isolate the processes involved. Without a clear understanding of these steps, trusting or debugging their behavior becomes harder, especially in tasks requiring reasoning, planning, or factual reliability. Researchers are thus focused on reverse-engineering these models to identify how information flows and decisions are made internally.

YOU MAY ALSO LIKE

How Focus Mode Has Changed In iOS 27

AI Almost Led The US Military To Start A War With China, Report Says

Existing interpretability methods like attention maps and feature attribution offer partial views into model behavior. While these tools help highlight which input tokens contribute to outputs, they often fail to trace the full chain of reasoning or identify intermediate steps. Moreover, these tools usually focus on surface-level behaviors and do not provide consistent insight into deeper computational structures. This has created the need for more structured, fine-grained methods to trace logic through internal representations over multiple steps.

To address this, researchers from Anthropic introduced a new technique called attribution graphs. These graphs allow researchers to trace the internal flow of information between features within a model during a single forward pass. By doing so, they attempt to identify intermediate concepts or reasoning steps that are not visible from the model’s outputs alone. The attribution graphs generate hypotheses about the computational pathways a model follows, which are then tested using perturbation experiments. This approach marks a significant step toward revealing the “wiring diagram” of large models, much like how neuroscientists map brain activity.

The researchers applied attribution graphs to Claude 3.5 Haiku, a lightweight language model released by Anthropic in October 2024. The method begins by identifying interpretable features activated by a specific input. These features are then traced to determine their influence on the final output. For example, when prompted with a riddle or poem, the model selects a set of rhyming words before writing lines, a form of planning. In another example, the model identifies “Texas” as an intermediate step to answer the question, “What’s the capital of the state containing Dallas?” which it correctly resolves as “Austin.” The graphs reveal the model outputs and how it internally represents and transitions between ideas.

The performance results from attribution graphs uncovered several advanced behaviors within Claude 3.5 Haiku. In poetry tasks, the model pre-plans rhyming words before composing each line, showing anticipatory reasoning. In multi-hop questions, the model forms internal intermediate representations, such as associating Dallas with Texas before determining Austin as the answer. It leverages both language-specific and abstract circuits for multilingual inputs, with the latter becoming more prominent in Claude 3.5 Haiku than in earlier models. Further, the model generates diagnoses internally in medical reasoning tasks and uses them to inform follow-up questions. These findings suggest that the model can abstract planning, internal goal-setting, and stepwise logical deductions without explicit instruction.

This research presents attribution graphs as a valuable interpretability tool that reveals the hidden layers of reasoning in language models. By applying this method, the team from Anthropic has shown that models like Claude 3.5 Haiku don’t merely mimic human responses—they compute through layered, structured steps. This opens the door to deeper audits of model behavior, allowing more transparent and responsible deployment of advanced AI systems.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.

🔥 [Register Now] miniCON Virtual Conference on OPEN SOURCE AI: FREE REGISTRATION + Certificate of Attendance + 3 Hour Short Event (April 12, 9 am- 12 pm PST) + Hands on Workshop [Sponsored]


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How Focus Mode Has Changed In iOS 27
AI & Technology

How Focus Mode Has Changed In iOS 27

September 18, 2026
AI Almost Led The US Military To Start A War With China, Report Says
AI & Technology

AI Almost Led The US Military To Start A War With China, Report Says

September 18, 2026
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
AI & Technology

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

September 18, 2026
Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright
AI & Technology

Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright

September 18, 2026
Next Post
Judge temporarily blocks deportation of Georgetown University researcher

Judge temporarily blocks deportation of Georgetown University researcher

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Air Force One emergency slide mistakenly deployed

Air Force One emergency slide mistakenly deployed

September 14, 2026
Tech company discloses first-of-its-kind A.I. cyberattack against a government

Tech company discloses first-of-its-kind A.I. cyberattack against a government

September 15, 2026
Fed hikes interest rates for first time in three years, drawing rebuke from Trump

Fed hikes interest rates for first time in three years, drawing rebuke from Trump

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!