• bitcoinBitcoin(BTC)$77,099.000.23%
  • ethereumEthereum(ETH)$2,516.532.70%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$724.121.63%
  • rippleXRP(XRP)$1.350.24%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.932.41%
  • tronTRON(TRX)$0.338180-0.67%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.29%
  • zcashZcash(ZEC)$1,160.513.63%
  • HyperliquidHyperliquid(HYPE)$79.10-1.00%
  • dogecoinDogecoin(DOGE)$0.0838760.07%
  • RainRain(RAIN)$0.015464-2.16%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$515.960.86%
  • whitebitWhiteBIT Coin(WBT)$80.090.63%
  • chainlinkChainlink(LINK)$11.51-0.43%
  • leo-tokenLEO Token(LEO)$9.15-0.40%
  • cardanoCardano(ADA)$0.204775-1.35%
  • stellarStellar(XLM)$0.1776650.71%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.570.58%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$52.981.03%
  • CantonCanton(CC)$0.096810-1.77%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.32%
  • uniswapUniswap(UNI)$5.99-0.50%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.42-1.67%
  • hedera-hashgraphHedera(HBAR)$0.074091-1.73%
  • nearNEAR Protocol(NEAR)$2.41-3.14%
  • shiba-inuShiba Inu(SHIB)$0.0000050.81%
  • suiSui(SUI)$0.72-1.72%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0560850.00%
  • MemeCoreMemeCore(M)$1.181.78%
  • tether-goldTether Gold(XAUT)$4,346.800.63%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.831.82%
  • BittensorBittensor(TAO)$234.32-2.09%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • aaveAave(AAVE)$124.031.50%
  • mantleMantle(MNT)$0.581.25%
  • pax-goldPAX Gold(PAXG)$4,352.510.73%
  • AsterAster(ASTER)$0.68-3.67%
  • polkadotPolkadot(DOT)$1.03-7.39%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054492-3.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from China Propose ALCUNA: A Groundbreaking Artificial Intelligence Benchmark for Evaluating Large-Scale Language Models on New Knowledge Integration

November 1, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from China Propose ALCUNA: A Groundbreaking Artificial Intelligence Benchmark for Evaluating Large-Scale Language Models on New Knowledge Integration
ShareShareShareShareShare

Evaluating large-scale language models (LLMs) in handling new knowledge is challenging. Researchers from Peking University introduced KnowGen, a method to generate new knowledge by modifying existing entity attributes and relationships. A benchmark called ALCUNA assesses LLMs’ abilities in knowledge understanding and differentiation. Their study reveals that LLMs often struggle with reasoning about new versus internal knowledge. It highlights the importance of caution when applying LLMs to new scenarios and encourages LLM development in handling new knowledge.

LLMs like FLAN-T5, GPT-3, OPT, LLama, and GPT-4 have excelled in various natural language tasks with applications in commercial products. Existing benchmarks assess their performance but rely on existing knowledge. Researchers propose Know-Gen and the ALCUNA benchmark to evaluate LLMs handling new knowledge. It emphasizes the need for caution when using LLMs with new scenarios or expertise and aims to spur development in this context.

LLMs have excelled in various tasks, but existing benchmarks may need to measure their ability to handle new knowledge. New standards are proposed to address this gap. Evaluating LLMs’ performance with new knowledge is crucial due to evolving information. Overlapping training and test data can affect memory assessment. Constructing a new knowledge benchmark is challenging but necessary. 

Know-Gen is a method for generating new knowledge by modifying entity attributes and relationships. It evaluates LLMs using zero-shot and few-shot methods, with and without Chain-of-Thought reasoning forms. Their study explores the impact of artificial entity similarity to parent entities, assessing attribute and name similarity. Multiple LLMs are evaluated on these benchmarks, including ChatGPT, Alpaca-7B, Vicuna-13B, and ChatGLM-6B.

LLMs’ performance on the ALCUNA benchmark, assessing their handling of new knowledge, could be better, especially in reasoning between new and existing knowledge. ChatGPT performs the best, with Vicuna as the second-best model. The few-shot setting generally outperforms zero-shot, and the CoT reasoning form is superior. LLMs struggle most with knowledge association and multi-hop reasoning. Entity similarity has an impact on their understanding. Their method emphasizes the importance of evaluating LLMs on new knowledge and proposes the Know-Gen and ALCUNA benchmarks to facilitate progress in this area.

The proposed method is limited to biological data but has potential applicability in other domains adhering to ontological representation. Evaluation is constrained to a few LLM models due to closed-source models and scale, warranting assessment with a broader range of models. It emphasizes LLMs’ new knowledge handling but lacks an extensive analysis of current benchmark limitations. It also does not address potential biases or ethical implications related to generating new knowledge using the Know-Gen approach or the responsible use of LLMs in new knowledge contexts.

KnowGen and the ALCUNA benchmark can help to evaluate LLMs in handling new knowledge. While ChatGPT performs best and Vicuna is second best, LLMs’ performance in reasoning between new and existing knowledge could be better. Few-shot settings outperform zero-shot, and CoT reasoning is superior. LLMs struggle with knowledge association, emphasizing the need for further development. It calls for caution in using LLMs with new knowledge and anticipates these benchmarks will drive LLM development in this context.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
AI & Technology

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

September 11, 2026
Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak
AI & Technology

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

September 11, 2026
New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Next Post
DHS IG briefs Jan. 6 Committee On Erased Secret Service Texts

DHS IG briefs Jan. 6 Committee On Erased Secret Service Texts

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Unrepentant Elizabeth Holmes already plotting biotech comeback from behind bars

Unrepentant Elizabeth Holmes already plotting biotech comeback from behind bars

September 9, 2026
How To Watch The Flame Fatales 2026 Speedrunning Marathon

How To Watch The Flame Fatales 2026 Speedrunning Marathon

September 10, 2026
Marco Rubio describes Trump’s military strategy with Iran as a ‘head for an eye’

Marco Rubio describes Trump’s military strategy with Iran as a ‘head for an eye’

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!