• bitcoinBitcoin(BTC)$85,368.004.57%
  • ethereumEthereum(ETH)$2,730.402.34%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$786.671.79%
  • rippleXRP(XRP)$1.515.92%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.804.35%
  • tronTRON(TRX)$0.3486681.56%
  • zcashZcash(ZEC)$1,494.64-1.57%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$93.83-0.26%
  • dogecoinDogecoin(DOGE)$0.09998612.87%
  • moneroMonero(XMR)$572.95-4.88%
  • whitebitWhiteBIT Coin(WBT)$85.863.05%
  • RainRain(RAIN)$0.013662-3.62%
  • chainlinkChainlink(LINK)$12.912.73%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2435014.92%
  • leo-tokenLEO Token(LEO)$8.950.41%
  • stellarStellar(XLM)$0.2120437.04%
  • nearNEAR Protocol(NEAR)$4.434.82%
  • uniswapUniswap(UNI)$8.994.42%
  • bitcoin-cashBitcoin Cash(BCH)$265.154.52%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$10.71-4.40%
  • litecoinLitecoin(LTC)$60.904.52%
  • CantonCanton(CC)$0.1173544.60%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.027.24%
  • hedera-hashgraphHedera(HBAR)$0.0926836.55%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.71%
  • BittensorBittensor(TAO)$323.0619.82%
  • shiba-inuShiba Inu(SHIB)$0.0000068.72%
  • crypto-com-chainCronos(CRO)$0.0658507.14%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.36-10.36%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,323.13-0.69%
  • okbOKB(OKB)$121.121.33%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • aaveAave(AAVE)$143.323.93%
  • EthenaEthena(ENA)$0.2164231.18%
  • pepePepe(PEPE)$0.00000528.58%
  • BitwayBitway(BTW)$0.803.32%
  • mantleMantle(MNT)$0.645.19%
  • OndoOndo(ONDO)$0.4328870.86%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

LOFT: A Comprehensive AI Benchmark for Evaluating Long-Context Language Models

June 23, 2024
in AI & Technology
Reading Time: 5 mins read
A A
LOFT: A Comprehensive AI Benchmark for Evaluating Long-Context Language Models
ShareShareShareShareShare

Long-context language models (LCLMs) have emerged as a promising technology with the potential to revolutionize artificial intelligence. These models aim to tackle complex tasks and applications while eliminating the need for intricate pipelines that were previously necessary due to context length limitations. However, the development and evaluation of LCLMs face significant challenges. Current evaluation methods rely on synthetic tasks or fixed-length datasets that fail to adequately assess the true capabilities of these models in real-world scenarios. The lack of rigorous benchmarks for truly long-context tasks hinders the ability to stress-test LCLMs on paradigm-shifting applications. Addressing these limitations is crucial for realizing the full potential of LCLMs and their impact on AI development.

Researchers have made several attempts to evaluate LCLMs, but each approach has limitations. While scalable, synthetic tasks like “Needle-in-A-Haystack” retrieval and multi-hop QA fail to capture the complexities of real-world scenarios. Other benchmarks using existing NLP datasets for extreme summarization and multi-document QA lack dynamic scaling capabilities, making them unsuitable for very long contexts. Instruction-following evaluations like LongAlpaca and LongBench-Chat offer limited task diversity and context lengths. Ada-LEval proposes a length-adaptable benchmark but relies on somewhat synthetic tasks. Studies on long-context QA using retrieved documents have shown promising results but are limited to contexts under 10,000 tokens. These existing methods fall short of comprehensively evaluating LCLMs on diverse, real-world tasks with truly long contexts, highlighting the need for more robust evaluation frameworks.

YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

DeepMind Researchers introduce the Long-Context Frontiers (LOFT) to overcome the limitations of existing evaluation methods for LCLMs. LOFT comprises six tasks across 35 datasets, encompassing text, visual, and audio modalities. This comprehensive benchmark is designed to push LCLMs to their limits and assess their real-world impact. Unlike previous evaluations, LOFT allows for the automatic creation of increasing context lengths, currently extending to one million tokens with the potential for further expansion. The benchmark focuses on four key areas where LCLMs have disruptive potential: retrieval across multiple modalities, retrieval-augmented generation (RAG), SQL-free database querying, and many-shot in-context learning. By targeting these areas, LOFT aims to provide a rigorous and scalable evaluation framework that can keep pace with the evolving capabilities of LCLMs.

The LOFT benchmark encompasses a diverse range of real-world applications to evaluate LCLMs comprehensively. It features six main tasks: retrieval, RAG, SQL-like reasoning, and many-shot in-context learning (ICL), spanning 35 datasets across text, visual, and audio modalities. The benchmark is designed with three context length limits: 32k, 128k, and 1M tokens, with the potential to scale further. For retrieval and RAG tasks, LOFT creates shared corpora containing gold passages and random samples, ensuring smaller corpora are subsets of larger ones. Many-shot ICL tasks adapt datasets from Big-Bench Hard and LongICLBench, while SQL tasks use Spider and SparC datasets with associated databases. This structure allows for rigorous evaluation of LCLMs’ performance across various context lengths and task types.

The LOFT benchmark evaluates Gemini 1.5 Pro, GPT-4, and Claude 3 Opus across various tasks and context lengths. Gemini 1.5 Pro performs well in text retrieval, visual retrieval, and audio retrieval, often matching or exceeding specialized models. It excels in multi-hop RAG tasks but struggles with multi-target datasets at larger scales. SQL-like reasoning tasks show potential but require improvement. Many-shot ICL results vary, with Gemini 1.5 Pro and Claude 3 Opus performing strongly in different areas. The benchmark highlights LCLMs’ growing capabilities across diverse tasks and modalities, while also identifying areas for improvement, particularly in scaling to larger contexts and complex reasoning.

In this study LOFT benchmark has been introduced to assess the evolving capabilities of Large Context Language Models (LCLMs) as they scale to handle increasingly long contexts. LOFT comprises tasks designed to evaluate LCLMs on potential paradigm-shifting applications: retrieval, retrieval-augmented generation, SQL-like reasoning, and in-context learning. With dynamic scaling up to 1 million tokens and the potential to extend to 1 billion, LOFT ensures ongoing relevance as LCLMs advance. Initial results show LCLMs demonstrating competitive retrieval capabilities compared to specialized systems, despite lacking specific training. However, the benchmark also reveals significant room for improvement in long-context reasoning, particularly as models access even longer context windows.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

[Announcing Gretel Navigator] Create, edit, and augment tabular data with the first compound AI system trusted by EY, Databricks, Google, and Microsoft


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
Car Dealer Chaos Arises From Cyberattack on .2T Market

Car Dealer Chaos Arises From Cyberattack on $1.2T Market

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Vanderbilt recovers fumble with 1 second left to stun NC State – ESPN

Vanderbilt recovers fumble with 1 second left to stun NC State – ESPN

September 19, 2026
Bayeux Tapestry on display in London after nearly 1,000 years

Bayeux Tapestry on display in London after nearly 1,000 years

September 15, 2026
White House launches video games that promote Trump’s agenda

White House launches video games that promote Trump’s agenda

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!