• bitcoinBitcoin(BTC)$83,480.00-1.15%
  • ethereumEthereum(ETH)$2,657.87-1.46%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$772.790.13%
  • rippleXRP(XRP)$1.50-1.07%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$120.09-0.41%
  • tronTRON(TRX)$0.3339670.26%
  • zcashZcash(ZEC)$1,566.86-4.49%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.06-0.38%
  • HyperliquidHyperliquid(HYPE)$90.06-3.83%
  • dogecoinDogecoin(DOGE)$0.094400-1.71%
  • chainlinkChainlink(LINK)$13.97-1.06%
  • moneroMonero(XMR)$537.41-4.44%
  • whitebitWhiteBIT Coin(WBT)$83.25-1.19%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.250433-0.54%
  • RainRain(RAIN)$0.012571-1.62%
  • leo-tokenLEO Token(LEO)$9.01-0.26%
  • stellarStellar(XLM)$0.211969-1.08%
  • nearNEAR Protocol(NEAR)$5.224.04%
  • bitcoin-cashBitcoin Cash(BCH)$316.72-4.61%
  • uniswapUniswap(UNI)$9.34-4.41%
  • CantonCanton(CC)$0.1435054.24%
  • litecoinLitecoin(LTC)$70.88-0.97%
  • suiSui(SUI)$1.246.56%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.73-0.19%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.645.05%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0967203.84%
  • quant-networkQuant(QNT)$262.7246.95%
  • BittensorBittensor(TAO)$309.54-3.26%
  • BitwayBitway(BTW)$1.2724.04%
  • shiba-inuShiba Inu(SHIB)$0.000006-1.89%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.064792-2.22%
  • OndoOndo(ONDO)$0.6012.06%
  • tether-goldTether Gold(XAUT)$4,202.22-1.76%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.2700600.61%
  • MemeCoreMemeCore(M)$1.18-3.63%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$118.23-2.22%
  • Pump.funPump.fun(PUMP)$0.00523119.82%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$150.68-2.83%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

The RAG reality check: New open-source framework lets enterprises scientifically measure AI performance

April 8, 2025
in AI & Technology
Reading Time: 6 mins read
A A
The RAG reality check: New open-source framework lets enterprises scientifically measure AI performance
ShareShareShareShareShare

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


Enterprises are spending time and money building out retrieval-augmented generation (RAG) systems. The goal is to have an accurate enterprise AI system, but are those systems actually working?

YOU MAY ALSO LIKE

20 Agentic Use Cases of TypeSafe AI’s Jev

Which Is Better To Use?

The inability to objectively measure whether RAG systems are actually working is a critical blind spot. One potential solution to that challenge is launching today with the debut of the Open RAG Eval open-source framework. The new framework was developed by enterprise RAG platform provider Vectara working together with Professor Jimmy Lin and his research team at the University of Waterloo.

Open RAG Eval transforms the currently subjective ‘this looks better than that’ comparison approach into a rigorous, reproducible evaluation methodology that can measure retrieval accuracy, generation quality and hallucination rates across enterprise RAG deployments.

The framework assesses response quality using two major metric categories: retrieval metrics and generation metrics. It allows organizations to apply this evaluation to any RAG pipeline, whether using Vectara’s platform or custom-built solutions. For technical decision-makers, this means finally having a systematic way to identify exactly which components of their RAG implementations need optimization.

“If you can’t measure it, you can’t improve it,” Jimmy Lin, professor at the University of Waterloo, told VentureBeat in an exclusive interview. “In information retrieval and dense vectors, you could measure lots of things, ndcg [Normalized Discounted Cumulative Gain], precision, recall…but when it came to right answers, we had no way, that’s why we started on this path.”

Why RAG evaluation has become the bottleneck for enterprise AI adoption

Vectara was an early pioneer in the RAG space. The company launched in October 2022, before ChatGPT was a household name. Vectara actually debuted technology it originally referred to as grounded AI back in May 2023, as a way to limit hallucinations, before the RAG acronym was commonly used.

Over the last few months, for many enterprises, RAG implementations have grown increasingly complex and difficult to assess. A key challenge is that organizations are moving beyond simple question-answering to multi-step agentic systems.

“In the agentic world, evaluation is doubly important, because these AI agents tend to be multi-step,” Am Awadallah, Vectara CEO and cofounder told VentureBeat. “If you don’t catch hallucination the first step, then that compounds with the second step, compounds with the third step, and you end up with the wrong action or answer at the end of the pipeline.”

How Open RAG Eval works: Breaking the black box into measurable components

The Open RAG Eval framework approaches evaluation through a nugget-based methodology. 

Lin explained that the nugget approach  breaks responses down into essential facts, then measures how effectively a system captures the nuggets.

The framework evaluates RAG systems across four specific metrics:

  1. Hallucination detection – Measures the degree to which generated content contains fabricated information not supported by source documents.
  2. Citation – Quantifies how well citations in the response are supported by source documents.
  3. Auto nugget – Evaluates the presence of essential information nuggets from source documents in generated responses.
  4. UMBRELA (Unified Method for Benchmarking Retrieval Evaluation with LLM Assessment) – A holistic method for assessing overall retriever performance

Importantly, the framework evaluates the entire RAG pipeline end-to-end, providing visibility into how embedding models, retrieval systems, chunking strategies, and LLMs interact to produce final outputs.

The technical innovation: Automation through LLMs

What makes Open RAG Eval technically significant is how it uses large language models to automate what was previously a manual, labor-intensive evaluation process.

“The state of the art before we started, was left versus right comparisons,” Lin explained. “So this is, do you like the left one better? Do you like the right one better? Or they’re both good, or they’re both bad? That was sort of one way of doing things.”

Lin noted that the nugget-based evaluation approach itself isn’t new, but its automation through LLMs represents a breakthrough.

The framework uses Python with sophisticated prompt engineering to get LLMs to perform evaluation tasks like identifying nuggets and assessing hallucinations, all wrapped in a structured evaluation pipeline.

Competitive landscape: How Open RAG Eval fits into the evaluation ecosystem

As enterprise use of AI continues to mature, there is a growing number of evaluation frameworks. Just last week, Hugging Face launched Yourbench to test models against the company’s internal data. At the end of January, Galileo launched its Agentic Evaluations technology.

The Open RAG Eval is different in that it is strongly focussed on the RAG pipeline, not just LLM outputs.. The framework also has a strong academic foundation and is built on established information retrieval science rather than ad-hoc methods.

The framework builds on Vectara’s previous contributions to the open-source AI community, including its Hughes Hallucination Evaluation Model (HHEM), which has been downloaded over 3.5 million times on Hugging Face and has become a standard benchmark for hallucination detection.

“We’re not calling it the Vectara eval framework, we’re calling it the Open RAG Eval framework because we really want other companies and other institutions to start helping build this out,” Awadallah emphasized. “We need something like that in the market, for all of us, to make these systems evolve in the right way.”

What Open RAG Eval means in the real world

While still an early stage effort, Vectara at least already has multiple users interested in using the Open RAG Eval framework.

Among them is Jeff Hummel, SVP of Product and Technology at real estate firm Anywhere.re. Hummel expects that partnering with Vectara will allow him to streamline his company’s RAG evaluation process.

Hummel noted that scaling his RAG deployment introduced significant challenges around infrastructure complexity, iteration velocity and rising costs. 

“Knowing the benchmarks and expectations in terms of performance and accuracy helps our team be predictive in our scaling calculations,” Hummel said. “To be frank, there weren’t a ton of frameworks for setting benchmarks on these attributes; we relied heavily on user feedback, which was sometimes objective and did translate to success at scale.”

From measurement to optimization: Practical applications for RAG implementers

For technical decision-makers, Open RAG Eval can help answer crucial questions about RAG deployment and configuration:

  • Whether to use fixed token chunking or semantic chunking
  • Whether to use hybrid or vector search, and what values to use for lambda in hybrid search
  • Which LLM to use and how to optimize RAG prompts
  • What thresholds to use for hallucination detection and correction

In practice, organizations can establish baseline scores for their existing RAG systems, make targeted configuration changes, and measure the resulting improvement. This iterative approach replaces guesswork with data-driven optimization.

While this initial release focuses on measurement, the roadmap includes optimization capabilities that could automatically suggest configuration improvements based on evaluation results. Future versions might also incorporate cost metrics to help organizations balance performance against operational expenses.

For enterprises looking to lead in AI adoption, Open RAG Eval means they can implement a scientific approach to evaluation rather than relying on subjective assessments or vendor claims. For those earlier in their AI journey, it provides a structured way to approach evaluation from the beginning, potentially avoiding costly missteps as they build out their RAG infrastructure.

Daily insights on business use cases with VB Daily

If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.

Read our Privacy Policy

Thanks for subscribing. Check out more VB newsletters here.

An error occured.

Credit: Source link
ShareTweetSendSharePin

Related Posts

20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Which Is Better To Use?
AI & Technology

Which Is Better To Use?

September 28, 2026
Are 3D Printers Worth Buying In 2026?
AI & Technology

Are 3D Printers Worth Buying In 2026?

September 28, 2026
Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Next Post
Hundreds were killed as Israel resumes airstrikes in Gaza

Hundreds were killed as Israel resumes airstrikes in Gaza

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Russian arena says Kanye West concerts will not happen

Russian arena says Kanye West concerts will not happen

September 24, 2026
Associated Banc-Corp: The Easy Money Has Already Been Made (Rating Downgrade)

Associated Banc-Corp: The Easy Money Has Already Been Made (Rating Downgrade)

September 28, 2026
New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!