• bitcoinBitcoin(BTC)$86,477.006.55%
  • ethereumEthereum(ETH)$2,773.525.05%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$799.383.49%
  • rippleXRP(XRP)$1.549.11%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$118.907.22%
  • tronTRON(TRX)$0.3444680.56%
  • zcashZcash(ZEC)$1,470.79-2.69%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.35%
  • HyperliquidHyperliquid(HYPE)$93.66-0.13%
  • dogecoinDogecoin(DOGE)$0.10001814.50%
  • moneroMonero(XMR)$596.825.99%
  • whitebitWhiteBIT Coin(WBT)$87.075.10%
  • RainRain(RAIN)$0.013945-1.19%
  • chainlinkChainlink(LINK)$13.155.29%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2443396.95%
  • leo-tokenLEO Token(LEO)$8.970.31%
  • stellarStellar(XLM)$0.2159949.17%
  • uniswapUniswap(UNI)$8.992.44%
  • nearNEAR Protocol(NEAR)$4.262.80%
  • bitcoin-cashBitcoin Cash(BCH)$269.095.76%
  • avalanche-2Avalanche(AVAX)$11.25-0.40%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$61.955.26%
  • CantonCanton(CC)$0.1181818.52%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.0415.96%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.454.71%
  • hedera-hashgraphHedera(HBAR)$0.0925657.08%
  • BittensorBittensor(TAO)$313.8019.54%
  • shiba-inuShiba Inu(SHIB)$0.0000069.81%
  • MemeCoreMemeCore(M)$1.46-0.73%
  • crypto-com-chainCronos(CRO)$0.06601810.65%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,360.09-0.33%
  • okbOKB(OKB)$123.184.07%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.49%
  • aaveAave(AAVE)$146.186.21%
  • OndoOndo(ONDO)$0.4545724.72%
  • BitwayBitway(BTW)$0.818.39%
  • mantleMantle(MNT)$0.657.20%
  • EthenaEthena(ENA)$0.212073-2.58%
  • polkadotPolkadot(DOT)$1.217.16%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

OLAPH: A Simple and Novel AI Framework that Enables the Improvement of Factuality through Automatic Evaluations

May 27, 2024
in AI & Technology
Reading Time: 5 mins read
A A
OLAPH: A Simple and Novel AI Framework that Enables the Improvement of Factuality through Automatic Evaluations
ShareShareShareShareShare

Large Language Models (LLMs) are stepping into clinical and medical fields as they grow in capability and versatility. These models have a number of benefits, including the capacity to supplement or even replace the work that doctors typically do. This include providing medical information, keeping track of patient information, and holding consultations with patients.

In the medical profession, one of the main advantages of LLMs is their capacity to produce long-form text, which is necessary for giving thorough responses to patient inquiries. Responses that are accurate and instructive are essential, particularly in medical situations when providing false information might have detrimental effects. For instance, when a patient asks about the origins of a white tongue, the LLM must answer truthfully about possible causes, including bacterial accumulation, without spreading myths, such as the idea that the condition is invariably dangerous and irreversible.

✅ [Featured Article] LLMWare.ai Selected for 2024 GitHub Accelerator: Enabling the Next Wave of Innovation in Enterprise RAG with Small Specialized Language Models

In the medical area, there are numerous scenarios in which producing comprehensive, extended answers is necessary. This is particularly crucial when answering inquiries from patients, as the details given must be true and factual. To ensure the accuracy and consistency of these answers, an automated process for assessing the assertions made by LLMs is required. 

To dive into this, in a recent study, a team of researchers has produced MedLFQA, a specialized benchmark dataset derived from pre-existing long-form question-answering datasets in the biomedical area. The goal of MedLFQA is to make it easier to automatically assess the factual accuracy of responses produced by LLMs. This dataset helps in determining the accuracy and dependability of the facts offered in these lengthy responses.

The team has offered a unique framework called OLAPH (Optimizing Large language models’ Answers with Preferences of reducing Hallucination). OLAPH uses a series of automated assessments to improve the factual accuracy of LLMs. The methodology uses an iterative training process to teach the LLM to favor responses with the greatest factual and assessment metrics scores. 

For each question, the OLAPH framework generates several response samples. Then, using predetermined assessment criteria, the response with the greatest score is chosen. The LLM is then further trained using this preferred response, bringing its subsequent responses closer to the correct and preferred answers. The model would otherwise produce false information, but this iterative approach helps to limit the issue of hallucinations.

The results have shown considerable improvements in factual accuracy for LLMs trained with the OLAPH framework, even when measured against measures not expressly included in the training procedure. A 7-billion parameter LLM trained with OLAPH produced long-form responses on par with professional medical responses in terms of quality.

The team has summarized their primary contributions as follows.

  1. The team has released MedLFQA, a reorganized benchmark dataset for automated assessment of the long-text generation produced by LLMs in the biomedical field. 
  1. In order to evaluate the veracity of medical claims provided in long-form responses, the team has developed two distinct statements that offer a comprehensive picture of the LLMs’ capacity to produce accurate data.
  1. OLAPH framework has been introduced, which enhances LLM replies through iterative learning and automatic evaluation. 
  1. It has been demonstrated that LLMs with 7 billion parameters when trained using the OLAPH framework, can produce long-form answers that are comparable in factual accuracy to those provided by medical experts.

In conclusion, this study proposes the OLAPH architecture to enhance long-form medical responses by iterative training, and it introduces MedLFQA as a baseline for assessing the factual accuracy of these responses produced by LLMs. The findings show that OLAPH has the potential to greatly improve LLMs’ dependability in producing accurate medical information, which could be crucial for a number of medical applications.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


[Free AI Webinar] ‘How to Build Personalized Marketing Chatbots (Gemini vs LoRA)’.


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Here’s Why Apple’s Mac Studio Has Become So Expensive
AI & Technology

Here’s Why Apple’s Mac Studio Has Become So Expensive

September 21, 2026
Tesla Will Soon Roll Out FSD Supervised In The Czech Republic
AI & Technology

Tesla Will Soon Roll Out FSD Supervised In The Czech Republic

September 21, 2026
Next Post
My 35-Year-Old Keeps Asking Me for Help

My 35-Year-Old Keeps Asking Me for Help

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Cohere CEO Warns Against an AI Safety ‘Cartel’

Cohere CEO Warns Against an AI Safety ‘Cartel’

September 16, 2026
Iran says U.S. strike on a wedding party killed civilians

Iran says U.S. strike on a wedding party killed civilians

September 18, 2026
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!