• bitcoinBitcoin(BTC)$77,969.00-1.40%
  • ethereumEthereum(ETH)$2,468.60-0.93%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$717.51-4.41%
  • rippleXRP(XRP)$1.38-3.18%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.16-2.62%
  • tronTRON(TRX)$0.3403990.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,223.86-1.48%
  • HyperliquidHyperliquid(HYPE)$82.90-3.43%
  • dogecoinDogecoin(DOGE)$0.085069-6.09%
  • RainRain(RAIN)$0.0161921.14%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$506.932.10%
  • whitebitWhiteBIT Coin(WBT)$80.58-1.44%
  • chainlinkChainlink(LINK)$11.82-2.02%
  • leo-tokenLEO Token(LEO)$9.190.09%
  • cardanoCardano(ADA)$0.213292-2.46%
  • stellarStellar(XLM)$0.179388-4.64%
  • bitcoin-cashBitcoin Cash(BCH)$246.52-4.35%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.37-3.17%
  • CantonCanton(CC)$0.101359-5.36%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.10%
  • uniswapUniswap(UNI)$6.03-8.94%
  • avalanche-2Avalanche(AVAX)$7.73-2.49%
  • hedera-hashgraphHedera(HBAR)$0.075988-3.92%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.40-3.12%
  • suiSui(SUI)$0.76-6.10%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.94%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056530-6.04%
  • MemeCoreMemeCore(M)$1.200.76%
  • tether-goldTether Gold(XAUT)$4,394.810.06%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BittensorBittensor(TAO)$251.64-2.63%
  • okbOKB(OKB)$112.25-1.83%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.07%
  • mantleMantle(MNT)$0.59-7.49%
  • AsterAster(ASTER)$0.71-5.51%
  • aaveAave(AAVE)$123.55-4.10%
  • pax-goldPAX Gold(PAXG)$4,397.770.03%
  • polkadotPolkadot(DOT)$1.10-7.22%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0561310.73%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

How Do Large Language Models Perform in Long-Form Question Answering? A Deep Dive by Salesforce Researchers into LLM Robustness and Capabilities

September 24, 2023
in AI & Technology
Reading Time: 4 mins read
A A
How Do Large Language Models Perform in Long-Form Question Answering? A Deep Dive by Salesforce Researchers into LLM Robustness and Capabilities
ShareShareShareShareShare

While Large Language Models (LLMs) like ChatGPT and GPT-4 have demonstrated better performance across several benchmarks, open-source projects like MMLU and OpenLLMBoard have quickly progressed in catching up across multiple applications and benchmarks. Understanding their capabilities, constraints, and distinctions becomes more crucial as they enter the new era of LLMs with rapid advancements in new models and methodologies. Although LLMs have demonstrated their ability to generate coherent text in tasks like summarization, more is needed about how well they do on LFQA. 

One of the significant problems that still needs to be solved is long-form question answering (LFQA), which has numerous and significant real-world applications (such as support forums, troubleshooting, customer service, etc.). Answering such inquiries frequently calls for complicated thinking skills to comprehend the question and make sense of the material that is dispersed across the original paper. The main points of the articles are condensed into abstract summaries. They assume that follow-up inquiries from these summaries would necessitate a better comprehension of the subjects connecting various sections of the source material. Additionally, other researchers show that responses that call for comprehension of more than a third of a lengthy material are frequently evaluated as “HARD” by people. 

Researchers from Salesforce suggest a scalable assessment approach to compare and contrast the differences between huge LLMs and smaller yet successful basic LLMs (such as Llama-7B, 13B) and their distilled counterparts (such as Alpaca-7B, 13B). To do this, they indicate that ChatGPT be instructed explicitly to construct complicated questions from document summaries. Their empirical study reveals that follow-up questions created from summaries present a difficult but more realistic setup for assessing the reasoning skills of LLMs on two fronts (complexity of generated questions and response quality of open-source LLMs). They use GPT-4 to determine the response quality on coherence, relevance, factual consistency, and correctness under earlier works because entirely depending on human review for long-form QA is expensive and challenging to scale. They also do a smaller-scale human evaluation, demonstrating that GPT-4 strongly correlates with human evaluation, making their assessment credible. 

The following are their primary conclusions from this study: 

• They recommend inferring from lengthier contexts by making numerous runs through the context for > 20% of the time to generate questions from abstractive summaries. 

• Distilled LLMs (Alpaca-7B, 13B) often rely less on context when generating questions from the original material, but their ability to create questions from document summaries is greatly reduced. 

• For questions derived from summaries (> 16.8%), responses produced by distilled LLMs can be consistent across contexts, but they frequently go off-topic, produce redundant replies, and are only partially accurate. 

• Alpaca-7B and 13B are more sensitive to lengthier contexts (>1024 tokens) than base LLMs (Llama), although they typically produce sensible replies.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI

AppleCare One Now Has A $50 Tier Per Month For Families

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI
AI & Technology

Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI

September 10, 2026
AppleCare One Now Has A  Tier Per Month For Families
AI & Technology

AppleCare One Now Has A $50 Tier Per Month For Families

September 10, 2026
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
AI & Technology

DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

September 10, 2026
2028 Volvo XC40 First Look: Hello new tech, goodbye EV
AI & Technology

2028 Volvo XC40 First Look: Hello new tech, goodbye EV

September 10, 2026
Next Post
US provided Canada with intelligence on killing of KTF chief Hardeep Singh Nijjar: Report

US provided Canada with intelligence on killing of KTF chief Hardeep Singh Nijjar: Report

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Federal program cuts spark concern over FDA and CDC response to cyclospora

Federal program cuts spark concern over FDA and CDC response to cyclospora

September 5, 2026
Is The Steam Deck Still Worth It In 2026?

Is The Steam Deck Still Worth It In 2026?

September 5, 2026
GPT-6 Astra is a Monster… That Still Needs You!

GPT-6 Astra is a Monster… That Still Needs You!

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!