• bitcoinBitcoin(BTC)$86,458.001.22%
  • ethereumEthereum(ETH)$2,753.460.79%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$791.180.53%
  • rippleXRP(XRP)$1.637.12%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.921.98%
  • tronTRON(TRX)$0.343978-1.41%
  • zcashZcash(ZEC)$1,619.158.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.77%
  • HyperliquidHyperliquid(HYPE)$97.113.08%
  • dogecoinDogecoin(DOGE)$0.1023392.23%
  • moneroMonero(XMR)$568.79-1.21%
  • whitebitWhiteBIT Coin(WBT)$86.881.14%
  • chainlinkChainlink(LINK)$13.050.82%
  • cardanoCardano(ADA)$0.2585645.37%
  • USDSUSDS(USDS)$1.000.01%
  • RainRain(RAIN)$0.013099-4.22%
  • leo-tokenLEO Token(LEO)$8.980.26%
  • stellarStellar(XLM)$0.2216094.59%
  • bitcoin-cashBitcoin Cash(BCH)$343.9229.72%
  • uniswapUniswap(UNI)$10.4117.34%
  • nearNEAR Protocol(NEAR)$4.483.76%
  • litecoinLitecoin(LTC)$64.395.85%
  • avalanche-2Avalanche(AVAX)$11.164.42%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • CantonCanton(CC)$0.114605-3.71%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.0998186.97%
  • suiSui(SUI)$1.031.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.472.71%
  • shiba-inuShiba Inu(SHIB)$0.0000062.47%
  • BittensorBittensor(TAO)$314.67-0.81%
  • crypto-com-chainCronos(CRO)$0.0684743.98%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.30-5.54%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,334.130.35%
  • okbOKB(OKB)$125.323.10%
  • BitwayBitway(BTW)$0.9316.62%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • aaveAave(AAVE)$152.296.83%
  • mantleMantle(MNT)$0.698.30%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.20%
  • EthenaEthena(ENA)$0.2177103.49%
  • OndoOndo(ONDO)$0.4394521.59%
  • pepePepe(PEPE)$0.000005-3.66%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

LongPiBench: A Comprehensive Benchmark that Explores How Even the Top Large Language Models have Relative Positional Biases

October 25, 2024
in AI & Technology
Reading Time: 5 mins read
A A
LongPiBench: A Comprehensive Benchmark that Explores How Even the Top Large Language Models have Relative Positional Biases
ShareShareShareShareShare

Accurate assessment of Large Language Models is best done with complex tasks involving long input sequences. Input sequence can exceed even 200,000 tokens in complex tasks such as repository analysis and information retrieval.LLMs, in response, have evolved, too, to accommodate context lengths of up to 1 million tokens. While examining the performance of capable LLMs on tasks involving long context lengths, researchers noticed a few underlying problems. Models exhibited difficulties while processing an input’s middle information, commonly called the “Lost in the Middle Effect.” Earlier research in LLM assessment had absolute positional biases that presumed relevant information concentration in specific locations. However, realistically, the information is scattered as multiple pertinent chunks of the text, which brings in the view of relative positional biases where the performance is examined with respect to the relative distance between chunks. Relative position introduces a bias in LLMs, thus affecting their performance. This article explains the latest research that systematically investigates positional biases in large language models.

Researchers from Tsinghua University and ModelBest Inc. introduced LongPiBench, a comprehensive benchmark to isolate and assess positional biases of LLMs. LongPiBench allows assessment concerning absolute and relative information positions with tasks ranging from easy to complex and 32k to 256k tokens. It contains three different tasks spanning four different context lengths-32k, 64k, 128k, and 256k. Furthermore, it has 16 different levels of absolute and relative positions. LongPiBench is collocated in two steps. Manual annotation of multiple seed examples is succeeded by augmentations to vary the positions of relevant information. The authors assessed multiple LLMs on this dataset, and it helped them to unravel the significant shortcomings of the latest models. 

YOU MAY ALSO LIKE

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

The Pros And Cons Of Using A Password Manager Over An Authenticator App

LongPiBench was developed by labeling seed points from Table SQL, Timeline Reordering, and Equation Solving tasks. This was followed by augmentation or rearrangement of relevant information. The context was decomposed into elements for each task based on respective units. Table SQL units were table entries, event entries for timeline reordering, and equation lines for equation solving. Every element was further annotated for relevance by forming queries around relevant items and adding irrelevant ones. The authors further implemented quality control checks to ensure integrity.

The research team evaluated 11 renowned LLMs on LongPiBench. They found that newer models are somewhat immune to the “Lost in Middle Effect,” but they still exhibit biases related to the spacing of relevant information. Six of the 11 LLMs were open-sourced models, and the remaining were commercial models. Llama-3.1-Instruct series, GPT-4o-mini, Claude-3-Haiku, and Gemini-1.5-Flash were some of the models assessed. During the preliminary tests, authors found that timeline reordering and equation solving were rigorous and challenging, and even top-performing models could have at most 20 % accuracy. Therefore, further analysis was performed on the Table SQL task. In tasks with absolute positioning, commercial and larger open-sourced models showed excellent robustness against the ‘lost in the middle effect ‘. For relative positioning, all models exhibited biases in different positions. Their performance sharply decreased with variations in relative distance. The issue of relative positioning bias is so severe that it reduced the recall rate by 30 %, even in the most straightforward tasks of retrieval. This highlights the necessity of continuously mitigating positional biases in long-text models

LongPiBench highlights the importance of relative positioning biases in modern LLMs and how they remain unresolved.It is essential to analyze this bias in more tasks to understand and solve the challenge because, if unresolved, this issue may substantially undermine the effectiveness of long-text language models in practical applications.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 55k+ ML SubReddit.

[Upcoming Live Webinar- Oct 29, 2024] The Best Platform for Serving Fine-Tuned Models: Predibase Inference Engine (Promoted)


Adeeba Alam Ansari is currently pursuing her Dual Degree at the Indian Institute of Technology (IIT) Kharagpur, earning a B.Tech in Industrial Engineering and an M.Tech in Financial Engineering. With a keen interest in machine learning and artificial intelligence, she is an avid reader and an inquisitive individual. Adeeba firmly believes in the power of technology to empower society and promote welfare through innovative solutions driven by empathy and a deep understanding of real-world challenges.

Listen to our latest AI podcasts and AI research videos here ➡️


Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
AI & Technology

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

September 23, 2026
The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
How To Hide Or Replace The Audio Button In iMessages
AI & Technology

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026
Improve Your Apple CarPlay Experience By Doing These Simple Things
AI & Technology

Improve Your Apple CarPlay Experience By Doing These Simple Things

September 22, 2026
Next Post
High school student’s report shines light on Mexican Repatriation 1930s

High school student's report shines light on Mexican Repatriation 1930s

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
United Kingdom: The Bank Of England Faces A Difficult Choice

United Kingdom: The Bank Of England Faces A Difficult Choice

September 17, 2026
Prosecution closes with argument Lindsay Clancy was a ‘functioning mom’ who knew ‘right and wrong’

Prosecution closes with argument Lindsay Clancy was a ‘functioning mom’ who knew ‘right and wrong’

September 22, 2026
She lived with headaches and brain fog for years. Then she tried magic mushrooms. – The Washington Post

She lived with headaches and brain fog for years. Then she tried magic mushrooms. – The Washington Post

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!