• bitcoinBitcoin(BTC)$77,258.00-0.05%
  • ethereumEthereum(ETH)$2,524.840.52%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$727.670.08%
  • rippleXRP(XRP)$1.370.71%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.69-0.29%
  • tronTRON(TRX)$0.3401030.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.08%
  • zcashZcash(ZEC)$1,125.98-2.66%
  • HyperliquidHyperliquid(HYPE)$79.160.44%
  • dogecoinDogecoin(DOGE)$0.0848150.71%
  • RainRain(RAIN)$0.0157242.56%
  • moneroMonero(XMR)$539.103.31%
  • USDSUSDS(USDS)$1.00-0.02%
  • whitebitWhiteBIT Coin(WBT)$80.270.10%
  • chainlinkChainlink(LINK)$11.51-0.05%
  • leo-tokenLEO Token(LEO)$9.14-0.15%
  • cardanoCardano(ADA)$0.2079360.71%
  • stellarStellar(XLM)$0.1800220.33%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.27-1.37%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.730.99%
  • uniswapUniswap(UNI)$6.426.75%
  • CantonCanton(CC)$0.0976730.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.15%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0748030.50%
  • avalanche-2Avalanche(AVAX)$7.39-0.92%
  • shiba-inuShiba Inu(SHIB)$0.0000052.14%
  • nearNEAR Protocol(NEAR)$2.371.59%
  • suiSui(SUI)$0.730.18%
  • crypto-com-chainCronos(CRO)$0.0606507.31%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-2.03%
  • tether-goldTether Gold(XAUT)$4,349.95-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.610.27%
  • BittensorBittensor(TAO)$232.27-1.63%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.06%
  • aaveAave(AAVE)$125.940.57%
  • pax-goldPAX Gold(PAXG)$4,354.66-0.05%
  • AsterAster(ASTER)$0.691.30%
  • mantleMantle(MNT)$0.55-4.50%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0570523.88%
  • polkadotPolkadot(DOT)$1.01-3.23%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes Infini-Gram: A Groundbreaking Approach to Scale and Enhance N-Gram Models Beyond Traditional Limits

February 12, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Proposes Infini-Gram: A Groundbreaking Approach to Scale and Enhance N-Gram Models Beyond Traditional Limits
ShareShareShareShareShare

Pretrained on trillion-token corpora, large neural language models (LLMs) have achieved remarkable performance strides (Touvron et al., 2023a; Geng & Liu, 2023). However, the scalability benefits of such data for traditional n-gram language models (LMs) still need to be explored. This paper from the University of Washington and Allen Institute for Artificial Intelligence delves into the relevance of n-gram LMs in the era of neural LLMs and introduces groundbreaking advancements in their modernization.

The authors affirm the continued utility of n-gram LMs in text analysis and enhancing neural LLMs. To address this, they modernized traditional n-gram LMs by scaling training data to an unprecedented 1.4 trillion tokens, rivaling the size of major open-source text corpora (Together, 2023; Soldaini et al., 2023). This represents the largest n-gram LM to date. Departing from historical constraints on n (e.g., n ≤ 5), the authors highlight the advantages of larger n’s value. Figure 1 illustrates the enhanced predictive capacity of n-gram LMs with larger n values, challenging conventional limitations. Consequently, they introduce the concept of an ∞-gram LM, with unbounded n, utilizing a backoff variant (Jurafsky & Martin, 2000) for improved accuracy.

The ∞-gram LM leverages a suffix array, replacing impractical n-gram count tables. This implementation, referred to as the infini-gram engine, achieves remarkable efficiency with 7 bytes of storage per token. The suffix array, built on 1.4 trillion tokens using an 80-core CPU node in under three days, ensures low-latency, resource-efficient querying at less than 20 milliseconds for n-gram counting. The ∞-gram engine, a testament to innovation, makes on-disk indexes integral to inference.

The ∞-gram LM, a conceptual extension of n-gram LMs, employs backoff judiciously to enhance predictive accuracy. Sparsity in ∞-gram estimates necessitate interpolation with neural LMs, addressing perplexity concerns. The paper introduces query types supported by Infini-gram, showcasing impressive latency benchmarks in Table 1.

Building on the suffix array implementation, the paper outlines efficient methods for n-gram counting, occurrence position retrieval, and document identification. Sharding strategies reduce latency proportional to the number of shards, optimizing processing times. Clever optimizations, such as reusing search results and on-disk search, further enhance the speed of ∞-gram computation.

Infini-gram’s application across diverse neural LMs, including GPT-2, GPT-Neo, LLaMA-2, and SILO, demonstrates consistent perplexity improvements (Table 2). The paper underscores the significance of data diversity, revealing ∞-gram’s efficacy in complementing neural LMs across different model series.

Analyses with ∞-gram shed light on human-written and machine-generated text. Notably, ∞-gram exhibits high accuracy in predicting the next token based on human-written document prefixes. The paper establishes a positive correlation between neural LMs and ∞-gram, suggesting the latter’s potential to enhance LM performance in predicting human-written text.

The paper concludes with a visionary outlook, presenting preliminary applications of the Infini-gram engine. From understanding text corpora to mitigating copyright infringement, the possibilities are diverse. The authors anticipate further insightful analyses and innovative applications fueled by Infini-gram.


Check out the Paper and Model. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

Is There Any Benefit To Restarting Your Gaming Handheld Regularly?

Vineet Kumar is a consulting intern at MarktechPost. He is currently pursuing his BS from the Indian Institute of Technology(IIT), Kanpur. He is a Machine Learning enthusiast. He is passionate about research and the latest advancements in Deep Learning, Computer Vision, and related fields.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Is There Any Benefit To Restarting Your Gaming Handheld Regularly?
AI & Technology

Is There Any Benefit To Restarting Your Gaming Handheld Regularly?

September 12, 2026
Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait
AI & Technology

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

September 12, 2026
Diablo V Is Coming Out In Spring 2029
AI & Technology

Diablo V Is Coming Out In Spring 2029

September 12, 2026
Next Post
Midnight plans to aggregate games inside Web3 Evergreen MMORPG

Midnight plans to aggregate games inside Web3 Evergreen MMORPG

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
My Wife Is Blowing All Of Our Money On Parties

My Wife Is Blowing All Of Our Money On Parties

September 9, 2026
OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI

OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI

September 9, 2026
Stay Tuned NOW Streaming Behind The Scenes! – Sept 11

Stay Tuned NOW Streaming Behind The Scenes! – Sept 11

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!