• bitcoinBitcoin(BTC)$76,021.000.55%
  • ethereumEthereum(ETH)$2,405.330.36%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$719.971.11%
  • rippleXRP(XRP)$1.290.78%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$98.351.64%
  • tronTRON(TRX)$0.3358071.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.13%
  • zcashZcash(ZEC)$1,316.0518.09%
  • HyperliquidHyperliquid(HYPE)$78.232.16%
  • dogecoinDogecoin(DOGE)$0.0804700.73%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013211-6.00%
  • moneroMonero(XMR)$490.55-1.36%
  • whitebitWhiteBIT Coin(WBT)$78.000.34%
  • leo-tokenLEO Token(LEO)$8.86-0.20%
  • chainlinkChainlink(LINK)$10.930.15%
  • cardanoCardano(ADA)$0.194744-0.31%
  • stellarStellar(XLM)$0.1813643.59%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$218.131.22%
  • USD1USD1(USD1)$1.00-0.01%
  • uniswapUniswap(UNI)$6.473.36%
  • litecoinLitecoin(LTC)$51.230.29%
  • CantonCanton(CC)$0.0956924.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-1.04%
  • nearNEAR Protocol(NEAR)$2.5911.33%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.391.67%
  • hedera-hashgraphHedera(HBAR)$0.073271-2.09%
  • suiSui(SUI)$0.713.38%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.61%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0557250.99%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,273.24-0.39%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.12-0.11%
  • BittensorBittensor(TAO)$219.920.95%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$110.460.63%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.17%
  • BitwayBitway(BTW)$0.746.03%
  • AsterAster(ASTER)$0.692.21%
  • pax-goldPAX Gold(PAXG)$4,273.18-0.50%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0573140.70%
  • mantleMantle(MNT)$0.551.98%
  • aaveAave(AAVE)$117.18-3.55%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Advancing AI’s Causal Reasoning: Hong Kong Polytechnic University and Chongqing University Researchers Develop CausalBench for LLM Evaluation

April 13, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Advancing AI’s Causal Reasoning: Hong Kong Polytechnic University and Chongqing University Researchers Develop CausalBench for LLM Evaluation
ShareShareShareShareShare

Causal learning delves into the foundational principles governing data distributions in the real world, influencing the operational effectiveness of artificial intelligence. The capacity of AI models to comprehend causality impacts their abilities to justify decisions, adapt to new data, and hypothesize alternative realities. Despite the rising interest in large language models (LLMs), evaluating their ability to process causality remains challenging due to the need for a thorough benchmark.

Existing research includes basic benchmarks assessing LLMs like GPT-3 and its variants through simple correlation tasks, often using limited datasets with straightforward causal structures. Studies also routinely examine LLMs such as BERT, RoBERTa, and DeBERTa, but these typically need more diversity in task complexity and dataset variety. Previous frameworks have tried integrating structured data into evaluations yet have yet to combine it effectively with background knowledge. This has restricted their utility in fully exploring LLM capabilities across realistic and complex scenarios, highlighting the need for more comprehensive and diverse evaluation methods in causal learning.

Researchers from Hong Kong Polytechnic University and Chongqing University have introduced CausalBench, a novel benchmark developed to assess LLMs’ causal learning capabilities rigorously. CausalBench stands out due to its comprehensive approach, incorporating multiple layers of complexity and a wide array of tasks that test LLMs’ abilities to interpret and apply causal reasoning in varied contexts. This methodology ensures a robust evaluation of the models under conditions that closely simulate real-world scenarios.

The methodology of CausalBench involves testing LLMs with datasets such as Asia, Sachs, and Survey to assess causal understanding. The framework tasks LLMs to identify correlations, construct causal skeletons, and determine causality directions. Evaluations measure performance through F1 score, accuracy, Structural Hamming Distance (SHD), and Structural Intervention Distance (SID). These tests occur in a zero-shot scenario to gauge each model’s innate causal reasoning capabilities without prior fine-tuning. This ensures the evaluation reflects each LLM’s fundamental abilities to process and analyze causal relationships across increasingly complex scenarios.

Initial evaluations using CausalBench reveal noteworthy performance variations across different LLMs. For instance, specific models like GPT4-Turbo achieved F1 scores above 0.5 in correlation tasks on datasets such as Asia and Sachs. However, performance generally declined in more complex causality assessments involving the Survey dataset, with many models struggling to exceed F1 scores of 0.3. These results highlight the varying capacities of LLMs to handle different levels of causal complexity, providing a clear metric of success and areas for future improvement in model training and algorithm development.

To conclude, researchers from Hong Kong Polytechnic University and Chongqing University introduced CasualBench,  a comprehensive benchmark that effectively measures the causal learning capabilities of LLMs. By utilizing diverse datasets and complex evaluation tasks, this research provides crucial insights into the strengths and weaknesses of various LLMs in understanding causality. The findings underscore the need for continued development in model training to enhance AI’s causal reasoning, vital for real-world applications where accurate decision-making and logical inference based on causality are essential.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


Want to get in front of 1.5 Million AI Audience? Work with us here


YOU MAY ALSO LIKE

AI Safety Can’t Rely on an Honor Code – Unite.AI

Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI

Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Safety Can’t Rely on an Honor Code – Unite.AI
AI & Technology

AI Safety Can’t Rely on an Honor Code – Unite.AI

September 16, 2026
Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI
AI & Technology

Denise Ruffner, VP Business Development and Commercial Operations Worldwide, Haiqu – Interview Series – Unite.AI

September 16, 2026
MindsEye Developer Build A Rocket Boy Is Reportedly Shutting Down
AI & Technology

MindsEye Developer Build A Rocket Boy Is Reportedly Shutting Down

September 16, 2026
The Boox Note Air6C E Ink Tablet Flips Pages Nearly 40 Percent Faster
AI & Technology

The Boox Note Air6C E Ink Tablet Flips Pages Nearly 40 Percent Faster

September 16, 2026
Next Post
Popular diabetes and weight loss drugs often hard to get for people who need them

Popular diabetes and weight loss drugs often hard to get for people who need them

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
‘I thought I was going to die’: 9/11 survivor shares her escape from the Twin Towers

‘I thought I was going to die’: 9/11 survivor shares her escape from the Twin Towers

September 12, 2026
Did This AI Research Bot Just Find an EDGE on Polymarket/Kalshi?

Did This AI Research Bot Just Find an EDGE on Polymarket/Kalshi?

September 16, 2026
NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI

NVIDIA Adds RTX PRO 5500 Blackwell GPU with 84 GB GDDR7 Memory – Unite.AI

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!