• bitcoinBitcoin(BTC)$79,534.001.34%
  • ethereumEthereum(ETH)$2,511.471.57%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$749.94-0.29%
  • rippleXRP(XRP)$1.442.60%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$104.541.46%
  • tronTRON(TRX)$0.338836-0.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,266.029.23%
  • HyperliquidHyperliquid(HYPE)$86.523.93%
  • dogecoinDogecoin(DOGE)$0.0912412.06%
  • RainRain(RAIN)$0.016368-2.33%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$82.223.05%
  • moneroMonero(XMR)$496.76-1.31%
  • chainlinkChainlink(LINK)$12.20-2.68%
  • leo-tokenLEO Token(LEO)$9.180.02%
  • cardanoCardano(ADA)$0.2209360.97%
  • stellarStellar(XLM)$0.1892770.29%
  • bitcoin-cashBitcoin Cash(BCH)$259.741.34%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$54.35-2.25%
  • CantonCanton(CC)$0.1058421.72%
  • uniswapUniswap(UNI)$6.70-4.33%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.54%
  • hedera-hashgraphHedera(HBAR)$0.078940-1.59%
  • avalanche-2Avalanche(AVAX)$7.98-0.70%
  • nearNEAR Protocol(NEAR)$2.6013.13%
  • suiSui(SUI)$0.820.29%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000050.10%
  • crypto-com-chainCronos(CRO)$0.060227-0.97%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.191.29%
  • tether-goldTether Gold(XAUT)$4,407.410.12%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$269.436.50%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$114.19-0.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • mantleMantle(MNT)$0.641.44%
  • AsterAster(ASTER)$0.75-0.91%
  • aaveAave(AAVE)$129.390.17%
  • polkadotPolkadot(DOT)$1.178.61%
  • Pump.funPump.fun(PUMP)$0.0046447.96%
  • pax-goldPAX Gold(PAXG)$4,409.910.11%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

November 5, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
ShareShareShareShareShare

In conversational AI, evaluating the Theory of Mind (ToM) through question-answering has become an essential benchmark. However, passive narratives need to improve in assessing ToM capabilities. To address this limitation, diverse questions have been designed to necessitate the same reasoning skills. These questions have revealed the limited ToM capabilities of LLMs. Even with chain-of-thought reasoning or fine-tuning, state-of-the-art LLMs still require assistance when dealing with these questions and perform below human standards.

Researchers from different universities introduced FANToM, a benchmark for testing ToM in LLMs through conversational question answering. It incorporates psychological and empirical insights into LLM evaluation. FANToM proves challenging for top LLMs, which perform worse than humans even with advanced reasoning or fine-tuning. The benchmark evaluates LLMs by requiring binary responses to questions about characters’ knowledge and listing characters with specific information. Human performance was assessed with 11 student volunteers.

FANToM is a new English benchmark designed to assess machine ToM in conversational contexts, focusing on social interactions. It includes 10,000 questions within multiparty conversations, emphasizing information asymmetry and distinct mental states among characters. The goal is to measure models’ ability to track beliefs in discussions, testing their understanding of others’ mental states and identifying instances of illusory ToM. 

FANToM tests machine ToM in LLMs through question-answering in conversational contexts with information asymmetry. It includes 10,000 questions based on multiparty conversations where characters have distinct mental states due to inaccessible information. The benchmark assesses LLMs’ ability to track beliefs in discussions and identify illusory ToM. Despite chain-of-thought reasoning or fine-tuning, existing LLMs perform significantly worse on FANToM than humans, as evaluated results indicate.

The evaluation results of FANToM reveal that even with chain-of-thought reasoning or fine-tuning, existing LLMs perform significantly worse than humans. Some LLM ToM reasoning in FANToM is deemed illusory, indicating their inability to comprehend distinct character perspectives. While applying zero-shot chain-of-thought logic or fine-tuning improves LLM scores, substantial gaps compared to human performance persist. The findings underscore the challenges in developing models with coherent Theory of Mind reasoning, emphasizing the difficulty of achieving human-level understanding in LLMs.

In conclusion, FANToM is a valuable benchmark for assessing ToM in LLMs during conversational interactions, highlighting the need for more interaction-oriented standards that align better with real-world use cases. The measure has shown that current LLMs underperform compared to humans, even with advanced techniques. It has identified the issue of internal consistency in neural models and provided various approaches to address it. FANToM emphasizes distinguishing between accessible and inaccessible information in ToM reasoning. 

Future research directions include grounding ToM reasoning in pragmatics, visual information, and belief graphs. Evaluations can encompass diverse conversation scenarios beyond small talk on specific topics, and multi-modal aspects like visual information can be integrated. Addressing the issue of internal consistency in neural models is crucial. FANToM is now publicly available for further research, promoting the advancement of ToM understanding in LLMs. Future studies may consider incorporating relationship variables for more dynamic social reasoning.


Check out the Paper, Github, and Project. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

How To Take Full Advantage Of Gemini When Planning Your Next Trip

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?
AI & Technology

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

September 9, 2026
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
AI & Technology

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

September 9, 2026
Next Post
Biden To Sign Executive Order To Protect Abortion Access

Biden To Sign Executive Order To Protect Abortion Access

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Grupo Financiero Inbursa Adopts Harvey Across Its Legal Organization – Unite.AI

Grupo Financiero Inbursa Adopts Harvey Across Its Legal Organization – Unite.AI

September 7, 2026
Microsoft Tells Court Copilot Rarely Reproduces Books in AI Copyright MDL – Unite.AI

Microsoft Tells Court Copilot Rarely Reproduces Books in AI Copyright MDL – Unite.AI

September 4, 2026
Video shows Paris attacker armed with two knives

Video shows Paris attacker armed with two knives

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!