• bitcoinBitcoin(BTC)$79,236.00-0.88%
  • ethereumEthereum(ETH)$2,493.78-0.65%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$739.77-1.75%
  • rippleXRP(XRP)$1.40-1.77%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$104.03-1.94%
  • tronTRON(TRX)$0.334683-0.41%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,156.45-5.59%
  • HyperliquidHyperliquid(HYPE)$85.39-2.96%
  • dogecoinDogecoin(DOGE)$0.0905740.21%
  • RainRain(RAIN)$0.016316-2.79%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$519.12-2.19%
  • chainlinkChainlink(LINK)$12.78-2.28%
  • whitebitWhiteBIT Coin(WBT)$76.663.93%
  • leo-tokenLEO Token(LEO)$9.16-1.82%
  • cardanoCardano(ADA)$0.221614-0.07%
  • stellarStellar(XLM)$0.1937984.09%
  • bitcoin-cashBitcoin Cash(BCH)$261.330.67%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • uniswapUniswap(UNI)$6.97-2.38%
  • litecoinLitecoin(LTC)$55.571.14%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.104391-5.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-2.02%
  • hedera-hashgraphHedera(HBAR)$0.0829231.60%
  • avalanche-2Avalanche(AVAX)$8.174.30%
  • suiSui(SUI)$0.832.13%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.44%
  • nearNEAR Protocol(NEAR)$2.32-4.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057090-1.07%
  • tether-goldTether Gold(XAUT)$4,411.91-0.17%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.152.35%
  • BittensorBittensor(TAO)$261.20-3.27%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$115.481.60%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.35%
  • AsterAster(ASTER)$0.77-1.40%
  • mantleMantle(MNT)$0.623.52%
  • aaveAave(AAVE)$132.78-1.50%
  • pax-goldPAX Gold(PAXG)$4,413.43-0.22%
  • OndoOndo(ONDO)$0.3856290.13%
  • polkadotPolkadot(DOT)$1.0710.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A New AI Research Introduces REV: A Game-Changer in AI Research – A New Information-Theoretic Measure Evaluating Novel, Label-Relevant Information in Free-Text Rationales

July 15, 2023
in AI & Technology
Reading Time: 5 mins read
A A
A New AI Research Introduces REV: A Game-Changer in AI Research – A New Information-Theoretic Measure Evaluating Novel, Label-Relevant Information in Free-Text Rationales
ShareShareShareShareShare

Model explanations have proved essential for trust and interpretability in natural language processing (NLP). Free-text rationales, which provide a natural language explanation of a model prediction, have gained popularity because of their adaptability in eliciting the thought process that went into the model’s choice, bringing them closer to human explanations. However, existing metrics for free-text explanation evaluation are still mostly accuracy-based and narrowly focused on how well a justification can assist a (proxy) model in predicting the label it explains. These metrics provide no insight into the new data given by the reason to the original input that would explain why the label was chosen—the precise function a justification is intended to fulfill. 

For instance, even though they provide differing amounts of fresh and pertinent information, the two rationales r*1 and r*1 in Fig. 1 would be deemed equally important under present measures. To address this issue, they introduce an automatic evaluation for free-text justifications along two dimensions in this paper: (1) whether the justification supports (i.e., is predictive of) the intended label, and (2) how much additional information it adds to the label justification beyond that which is already present in the input. 

For instance, the justification r^1,b in Fig. 1 contradicts (1) as it does not anticipate the label “enjoy nature.” Although rationale r^1,a does support the label, it does not provide any new information to what is already stated in input x to support it; as a result, it violates clause (2). Both requirements of the rationale r*1 are met: it provides additional and pertinent information that goes beyond the input to support the label. Both r^1,a and r^1,b will be penalized in their evaluation while r1,a and r1,b will be rewarded. Researchers from the University of Virginia, Allen Institute for AI, University of Southern California, and the University of Washington in this study provide REV2, an information-theoretic framework for assessing free-text justifications along the two previously described dimensions that they have modified. 

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!
Figure 1: The metric REV can distinguish all three rationales by measuring how much new and label-relevant information each adds over a vacuous rationale

REV is based on conditional V-information, which measures the extent to which a representation has information beyond that of a baseline representation and is available to a model family V. They treat any vacuous justification that does nothing more than (and declaratively) pair an input with a predetermined label without adding any new information that would shed light on the decision-making process behind the label as their baseline representation. When evaluating rationales, REV adapts conditional V-information. To do this, they compare two representations: one from an evaluation model trained to produce the label given the input and the rationale and the other from another evaluation model for the same task, but only considering the input (under the guise of a void rationale). 

Other metrics cannot assess fresh and label-relevant information in rationales because they do not account for empty justifications. For two reasoning tasks, commonsense question-answering and natural language inference, across four benchmarks, they offer evaluations with REV for justifications in their studies. Numerous quantitative assessments show how REV may provide ratings along new axes for free-text justifications while more aligned with human judgments than current measurements. They also provide comparisons to show how sensitive REV is to different levels of input disturbances. Additionally, evaluation with REV sheds light on why the performance of predictions is not always enhanced by the rationales discovered by chain-of-thought prompting.


Check out the Paper and GitHub link. Don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 800+ AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Capcom Is Reviving More Dormant Franchises After The Success Of Onimusha: Way Of The Sword

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Capcom Is Reviving More Dormant Franchises After The Success Of Onimusha: Way Of The Sword
AI & Technology

Capcom Is Reviving More Dormant Franchises After The Success Of Onimusha: Way Of The Sword

September 7, 2026
OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
AI & Technology

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

September 7, 2026
When Is It No Longer Worth Repairing Your Phone And Buying A New One Instead
AI & Technology

When Is It No Longer Worth Repairing Your Phone And Buying A New One Instead

September 7, 2026
Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI
AI & Technology

Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI

September 7, 2026
Next Post
The RealReal CEO Says They’re Focused on Keeping Fakes Off the Market

The RealReal CEO Says They're Focused on Keeping Fakes Off the Market

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Massive new SoCal Costco store sets opening date as completion nears

Massive new SoCal Costco store sets opening date as completion nears

August 31, 2026
In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars – Unite.AI

In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars – Unite.AI

September 6, 2026
Enterprises put non-Nvidia chips 14 points ahead of Nvidia’s next-gen GPUs on their evaluation lists

Enterprises put non-Nvidia chips 14 points ahead of Nvidia’s next-gen GPUs on their evaluation lists

September 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!