• bitcoinBitcoin(BTC)$79,886.000.39%
  • ethereumEthereum(ETH)$2,502.392.07%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$763.075.83%
  • rippleXRP(XRP)$1.421.41%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.302.38%
  • tronTRON(TRX)$0.3333880.48%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.061.66%
  • HyperliquidHyperliquid(HYPE)$85.601.74%
  • zcashZcash(ZEC)$1,062.913.58%
  • dogecoinDogecoin(DOGE)$0.0908507.26%
  • RainRain(RAIN)$0.0172134.70%
  • moneroMonero(XMR)$552.744.17%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.264.91%
  • whitebitWhiteBIT Coin(WBT)$73.680.74%
  • leo-tokenLEO Token(LEO)$9.331.22%
  • cardanoCardano(ADA)$0.2211074.91%
  • stellarStellar(XLM)$0.1858363.43%
  • bitcoin-cashBitcoin Cash(BCH)$261.546.04%
  • daiDai(DAI)$1.000.01%
  • uniswapUniswap(UNI)$7.2616.02%
  • CantonCanton(CC)$0.1106992.91%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$54.835.68%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.421.43%
  • hedera-hashgraphHedera(HBAR)$0.0812723.24%
  • avalanche-2Avalanche(AVAX)$7.693.93%
  • suiSui(SUI)$0.804.14%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000054.80%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.230.99%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0573552.84%
  • tether-goldTether Gold(XAUT)$4,423.79-0.08%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.13-0.60%
  • okbOKB(OKB)$115.866.18%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BittensorBittensor(TAO)$236.663.73%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.06%
  • AsterAster(ASTER)$0.786.83%
  • aaveAave(AAVE)$135.683.80%
  • mantleMantle(MNT)$0.592.35%
  • pax-goldPAX Gold(PAXG)$4,430.96-0.10%
  • OndoOndo(ONDO)$0.3752712.03%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0573271.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research Evaluates the Correctness and Faithfulness of Instruction-Following Models For Their Ability To Perform Question-Answering

August 5, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Research Evaluates the Correctness and Faithfulness of Instruction-Following Models For Their Ability To Perform Question-Answering
ShareShareShareShareShare

Recently introduced Large Language Models (LLMs) have taken the Artificial Intelligence (AI) community by storm. These models have been able to successfully imitate human beings by using super-good Natural Language Processing (NLP), Natural Language Generation (NLG) and Natural Language Understanding (NLU). LLMs have become famous for imitating humans for having realistic conversations and are capable of answering simple and complex questions, content generation, code completion, machine translation, and text summarization. The goal of NLP is to make it possible for computer systems to comprehend and react to commands given in natural language, enabling people to engage with them in a more natural and flexible way, the best example of which is the instruction following models.

These models are trained using LLMs, supervised examples, or other types of supervision, and exposure to thousands of tasks written as natural language instructions. In recent research, a team from Mila Quebec AI Institute, McGill University, and Facebook CIFAR AI Chair has researched evaluating the performance of instruction-following models for their ability to perform question-answering (QA) on a given set of text passages. These models can answer questions when provided with a prompt describing the task, the question, and relevant text passages retrieved by a retriever, and the responses produced by these models are known to be natural and informative, which helps build users’ trust and engagement. 

These models can respond to user queries naturally and fluently by only adding retrieved documents and instructions to their input. However, this extra verbosity makes it difficult for conventional QA evaluation metrics like exact match (EM) and F1 score to effectively quantify model performance. This is due to the possibility that the model’s response may include more details that the reference answer omits while still being accurate. The team has provided two criteria for measuring instruction-following models in retrieval-augmented quality assurance (QA) in order to overcome this problem.

  1. Regarding information necessity, accuracy: This dimension evaluates how well the model satisfies the informational requirements of a user. It is concerned with whether the generated response includes pertinent information, even if it goes beyond what is mentioned directly in the reference answer.
  1. Fidelity in relation to information provided: This dimension assesses how well the model grounds answers in the knowledge presented. A true model should refrain from responding when irrelevant information is presented, in addition to giving precise answers when it is accessible.

The authors have evaluated several recent instruction-following models on three diverse QA datasets: Natural Questions for open-domain QA, HotpotQA for multi-hop QA, and TopiOCQA for conversational QA. They analyzed 900 model responses manually and compared the results with different automatic metrics for accuracy and faithfulness. Their research has suggested that recall, which measures the percentage of tokens from the reference answer that are also present in the model response, correlates more strongly with correctness than lexical overlap metrics like EM or F1 score. Compared to other token-overlap metrics for faithfulness, K-Precision, which is the percentage of model answer tokens that exist in the knowledge snippet, has a stronger correlation with human judgments.

In conclusion, this study seeks to advance a more thorough assessment of instruction-following models for QA tasks, taking into account both their advantages and disadvantages. The team has promoted additional advancement in this area by making their code and data accessible on their GitHub repository


Check out the Paper, GitHub, and Tweet. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 27k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Is The Steam Deck Still Worth It In 2026?

How To Find Your MacBook’s Diagnostic Menu

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🔥 Use SQL to predict the future (Sponsored)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Is The Steam Deck Still Worth It In 2026?
AI & Technology

Is The Steam Deck Still Worth It In 2026?

September 5, 2026
How To Find Your MacBook’s Diagnostic Menu
AI & Technology

How To Find Your MacBook’s Diagnostic Menu

September 5, 2026
The Reasons Rugged Laptops Are Rarely Bought By Consumers
AI & Technology

The Reasons Rugged Laptops Are Rarely Bought By Consumers

September 5, 2026
GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI
AI & Technology

GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

September 5, 2026
Next Post
‘Bloomberg Technology’ Full Show (6/19/2019)

'Bloomberg Technology' Full Show (6/19/2019)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Video investigation: Medics under fire in Lebanon

Video investigation: Medics under fire in Lebanon

September 2, 2026
Aon close to acquiring USI Insurance from KKR in B deal: report

Aon close to acquiring USI Insurance from KKR in $17B deal: report

August 30, 2026
Full Episode: TODAY Show – July 24

Full Episode: TODAY Show – July 24

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!