• bitcoinBitcoin(BTC)$86,173.001.04%
  • ethereumEthereum(ETH)$2,744.420.77%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$788.290.48%
  • rippleXRP(XRP)$1.626.75%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.981.53%
  • tronTRON(TRX)$0.343726-1.31%
  • zcashZcash(ZEC)$1,624.707.48%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.77%
  • HyperliquidHyperliquid(HYPE)$96.652.29%
  • dogecoinDogecoin(DOGE)$0.1010112.18%
  • moneroMonero(XMR)$569.01-0.37%
  • whitebitWhiteBIT Coin(WBT)$86.610.99%
  • chainlinkChainlink(LINK)$12.930.38%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2557744.36%
  • RainRain(RAIN)$0.013031-4.65%
  • leo-tokenLEO Token(LEO)$8.98-0.02%
  • stellarStellar(XLM)$0.2183803.14%
  • bitcoin-cashBitcoin Cash(BCH)$350.5431.45%
  • uniswapUniswap(UNI)$10.4217.29%
  • nearNEAR Protocol(NEAR)$4.401.00%
  • litecoinLitecoin(LTC)$63.424.51%
  • avalanche-2Avalanche(AVAX)$11.113.72%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.000.01%
  • CantonCanton(CC)$0.112649-5.76%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.0979273.36%
  • suiSui(SUI)$1.021.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.461.79%
  • shiba-inuShiba Inu(SHIB)$0.0000062.65%
  • BittensorBittensor(TAO)$311.30-1.92%
  • crypto-com-chainCronos(CRO)$0.0670672.46%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.29-4.28%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,324.950.45%
  • okbOKB(OKB)$124.582.18%
  • BitwayBitway(BTW)$0.9416.06%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • aaveAave(AAVE)$149.836.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • mantleMantle(MNT)$0.697.29%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.04%
  • EthenaEthena(ENA)$0.2159792.10%
  • OndoOndo(ONDO)$0.4345470.57%
  • pepePepe(PEPE)$0.000005-4.17%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

SynDL: A Synthetic Test Collection Utilizing Large Language Models to Revolutionize Large-Scale Information Retrieval Evaluation and Relevance Assessment

September 2, 2024
in AI & Technology
Reading Time: 6 mins read
A A
SynDL: A Synthetic Test Collection Utilizing Large Language Models to Revolutionize Large-Scale Information Retrieval Evaluation and Relevance Assessment
ShareShareShareShareShare

Information retrieval (IR) is a fundamental aspect of computer science, focusing on efficiently locating relevant information within large datasets. As data grows exponentially, the need for advanced retrieval systems becomes increasingly critical. These systems use sophisticated algorithms to match user queries with relevant documents or passages. Recent developments in machine learning, particularly in natural language processing (NLP), have significantly enhanced the capabilities of IR systems. By employing techniques such as dense passage retrieval and query expansion, researchers aim to improve the accuracy and relevance of search results. These advancements are pivotal in fields ranging from academic research to commercial search engines, where the ability to quickly & accurately retrieve information is essential.

A persistent challenge in information retrieval is the creation of large-scale test collections that can accurately model the complex relationships between queries and documents. Traditional test collections often rely on human assessors to judge the relevance of records, a process that is not only time-consuming but also costly. This reliance on human judgment limits the scale of test collections and hampers the developing and evaluation of more advanced retrieval systems. For instance, existing collections like MS MARCO include over 1 million questions, but for each query, only an average of 10 passages are deemed relevant, leaving approximately 8.8 million passages as non-relevant. This significant imbalance highlights the difficulty in capturing the full complexity of query-document relationships, particularly in large datasets.

YOU MAY ALSO LIKE

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

Researchers have explored methods to enhance the effectiveness of IR systems. One approach uses large language models (LLMs), which have shown promise in generating relevance judgments that align closely with human assessments. The TREC Deep Learning Tracks, organized from 2019 to 2023, have been instrumental in advancing this research. These tracks have provided test collections that include queries with varying degrees of relevance labels. However, even these efforts have been constrained by the limited number of queries, only 82 in the 2023 track, used for evaluation. This limitation has sparked interest in developing new methods to scale the evaluation process while maintaining high accuracy and relevance.

Researchers from University College London, University of Sheffield, Amazon, and Microsoft introduced a new test collection named SynDL. SynDL represents a significant advancement in the field of IR by leveraging LLMs to generate a large-scale synthetic dataset. This collection extends the existing TREC Deep Learning Tracks by incorporating over 1,900 test queries and generating 637,063 query-passage pairs for relevance assessment. The development process of SynDL involved aggregating initial queries from the five years of TREC Deep Learning Tracks, including 500 synthetic queries generated by GPT-4 and T5 models. These synthetic queries allow for a more extensive analysis of query-document relationships and provide a robust framework for evaluating the performance of retrieval systems.

The core innovation of SynDL lies in its use of LLMs to annotate query-passage pairs with detailed relevance labels. Unlike previous collections, SynDL offers a deep and wide relevance assessment by associating each query with an average of 320 passages. This approach increases the scale of the evaluation and provides a more nuanced understanding of the relevance of each passage to a given query. SynDL effectively bridges the gap between human and machine-generated relevance judgments by leveraging LLMs’ advanced natural language comprehension capabilities. The use of GPT-4 for annotation has been particularly noteworthy, as it enables high granularity in labeling passages as irrelevant, related, highly relevant, or perfectly relevant.

The evaluation of SynDL has demonstrated its effectiveness in providing reliable and consistent system rankings. In comparative studies, SynDL highly correlated with human judgments, with Kendall’s Tau coefficients of 0.8571 for NDCG@10 and 0.8286 for NDCG@100. Moreover, the top-performing systems from the TREC Deep Learning Tracks maintained their rankings when evaluated using SynDL, indicating the robustness of the synthetic dataset. The inclusion of synthetic queries also allowed researchers to analyze potential biases in LLM-generated text, particularly regarding the use of similar language models in both query generation and system evaluation. Despite these concerns, SynDL exhibited a balanced evaluation environment, where GPT-based systems did not receive undue advantages.

In conclusion, SynDL represents a major advancement in information retrieval by addressing the limitations of existing test collections. Through the innovative use of large language models, SynDL provides a large-scale, synthetic dataset that enhances the evaluation of retrieval systems. With its detailed relevance labels and extensive query coverage, SynDL offers a more comprehensive framework for assessing the performance of IR systems. The successful correlation with human judgments and the inclusion of synthetic queries make SynDL a valuable resource for future research.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 50k+ ML SubReddit

Here is a highly recommended webinar from our sponsor: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’


Aswin AK is a consulting intern at MarkTechPost. He is pursuing his Dual Degree at the Indian Institute of Technology, Kharagpur. He is passionate about data science and machine learning, bringing a strong academic background and hands-on experience in solving real-life cross-domain challenges.

▶• ılıılıılıılıılı Upcoming Live Session: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
AI & Technology

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

September 23, 2026
The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
How To Hide Or Replace The Audio Button In iMessages
AI & Technology

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026
Next Post
Can You Still Build Wealth and Rent?

Can You Still Build Wealth and Rent?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Indonesian volcano spews ash and smoke during eruption

Indonesian volcano spews ash and smoke during eruption

September 20, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
An iOS 27 Bug Can Temporarily Freeze Your iPhone

An iOS 27 Bug Can Temporarily Freeze Your iPhone

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!