• bitcoinBitcoin(BTC)$84,465.00-2.02%
  • ethereumEthereum(ETH)$2,683.87-2.65%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$765.64-2.92%
  • rippleXRP(XRP)$1.50-4.83%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$114.94-3.07%
  • tronTRON(TRX)$0.341864-0.05%
  • zcashZcash(ZEC)$1,501.24-7.77%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.54%
  • HyperliquidHyperliquid(HYPE)$94.01-2.94%
  • dogecoinDogecoin(DOGE)$0.092575-7.88%
  • moneroMonero(XMR)$551.52-3.60%
  • whitebitWhiteBIT Coin(WBT)$84.78-2.23%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.35-5.42%
  • cardanoCardano(ADA)$0.238468-6.99%
  • RainRain(RAIN)$0.012244-6.42%
  • leo-tokenLEO Token(LEO)$8.96-0.16%
  • stellarStellar(XLM)$0.202331-6.56%
  • bitcoin-cashBitcoin Cash(BCH)$336.86-1.94%
  • uniswapUniswap(UNI)$9.30-7.76%
  • nearNEAR Protocol(NEAR)$4.30-2.63%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$62.04-2.24%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.32-8.23%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.109900-4.63%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-4.05%
  • hedera-hashgraphHedera(HBAR)$0.090575-8.80%
  • suiSui(SUI)$0.96-6.64%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.88%
  • BittensorBittensor(TAO)$287.60-9.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061142-8.75%
  • BitwayBitway(BTW)$1.0417.17%
  • MemeCoreMemeCore(M)$1.22-7.07%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,289.52-1.65%
  • okbOKB(OKB)$118.89-3.41%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.23%
  • mantleMantle(MNT)$0.65-2.51%
  • aaveAave(AAVE)$139.09-5.91%
  • EthenaEthena(ENA)$0.205530-6.19%
  • OndoOndo(ONDO)$0.413498-6.66%
  • AsterAster(ASTER)$0.69-5.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

The Impact of Questionable Research Practices on the Evaluation of Machine Learning (ML) Models

July 27, 2024
in AI & Technology
Reading Time: 5 mins read
A A
The Impact of Questionable Research Practices on the Evaluation of Machine Learning (ML) Models
ShareShareShareShareShare

Evaluating model performance is essential in the significantly advancing fields of Artificial Intelligence and Machine Learning, especially with the introduction of Large Language Models (LLMs). This review procedure helps understand these models’ capabilities and create dependable systems based on them. However, what is referred to as Questionable Research Practices (QRPs) frequently jeopardize the integrity of these assessments. These methods have the potential to greatly exaggerate published results, deceiving the scientific community and the general public about the actual effectiveness of ML models.

The primary driving force for QRPs is the ambition to publish in esteemed journals or to attract funding and users. Due to the intricacy of ML research, which includes pre-training, post-training, and evaluation stages, there is much potential for QRPs. Contamination, cherrypicking, and misreporting are the three basic categories these actions fall into.

YOU MAY ALSO LIKE

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

Contamination

When data from the test set is used for training, assessment, or even model prompts, this is known as contamination. High-capacity models such as LLMs can remember test data that is exposed during training. Researchers have provided extensive documentation on this problem, detailing cases in which models were purposefully or unintentionally trained using test data. There are various ways that contamination can occur, which are as follows.

  1. Training on the Test Set: This results in unduly optimistic performance predictions when test data is unintentionally added to the training set.
  1. Prompt Contamination: During few-shot evaluations, using test data in the prompt gives the model an unfair advantage.
  1. Retrieval Augmented Generation (RAG) Contamination: Data leakage via retrieval systems using benchmarks.
  1. Dirty Paraphrases and Contaminated Models: Rephrased test data and contaminated models are used to train models, while contaminated models are used to generate training data.
  1. Over-hyping and Meta-contamination: Exaggerating and meta-contaminating designs by recycling contaminated designs or fine-tuning hyperparameters after test results are obtained.

Cherrypicking

Cherrypicking is the practice of adjusting experimental conditions to support the intended result. It is possible for researchers to test their models several times under different scenarios and only publish the best outcomes. This comprises of the following.

  1. Baseline Nerfing: It is the deliberate under-optimization of baseline models to give the impression that the new model is better.
  1. Runtime Hacking: It includes modifying inference parameters after the fact to improve performance metrics.
  1. Benchmark Hacking Choosing simpler benchmarks or subsets of benchmarks to make sure the model runs well is known as benchmark hacking.
  1. Golden Seed: Reporting the top-performing seed after training with several random seeds.

Misreporting

A variety of techniques are included in misreporting when researchers present generalizations based on skewed or limited benchmarks. For example, consider the following:

  1. Superfluous Cog: Claiming originality by adding unnecessary modules.
  1. Whack-a-mole: Keeping an eye on and adjusting certain malfunctions as needed.
  1. P-hacking: The selective presentation of statistically significant findings.
  1. Point Scores: Ignoring variability by reporting results from a single run without error bars.
  1. Outright Lies and Over/Underclaiming: Creating fake outcomes or making incorrect assertions regarding the capabilities of the model.

Irreproducible Research Practices (IRPs), in addition to QRPs, add to the complexity of the ML evaluation environment. It is challenging for subsequent researchers to duplicate, expand upon, or examine earlier research because of IRPs. One common instance is dataset concealing, in which researchers withhold information about the training datasets they utilize, including metadata. The competitive nature of ML research and worries about copyright infringement frequently motivate this technique. The validation and replication of discoveries, which are essential to the advancement of science, are hampered by the lack of transparency in dataset sharing.

In conclusion, the integrity of ML research and assessment is critical. Although QRPs and IRPs may benefit companies and researchers in the near term, they damage the field’s credibility and dependability over the long run. Setting up and upholding strict guidelines for research processes is essential as ML models are used more often and have a greater impact on society. The full potential of ML models can only be attained by openness, responsibility, and a dedication to moral research. It is imperative that the community collaborates to recognize and address these practices, guaranteeing that the progress in ML is grounded in honesty and fairness.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 47k+ ML SubReddit

Find Upcoming AI Webinars here


Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses
AI & Technology

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips
AI & Technology

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

September 23, 2026
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
AI & Technology

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

September 23, 2026
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Next Post
Sec. Blinken says Israel has the ‘will’ and ‘means to police itself’: Full interview

Sec. Blinken says Israel has the 'will' and 'means to police itself': Full interview

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Near-total lunar eclipse seen over Southern California

Near-total lunar eclipse seen over Southern California

September 22, 2026
Fermi Explorer co-founder talks mission to send interstellar spacecraft to neighboring star

Fermi Explorer co-founder talks mission to send interstellar spacecraft to neighboring star

September 19, 2026
ATI Inc. (ATI) Presents at Morgan Stanley’s 14th Annual Laguna Conference Transcript

ATI Inc. (ATI) Presents at Morgan Stanley’s 14th Annual Laguna Conference Transcript

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!