• bitcoinBitcoin(BTC)$84,376.00-2.15%
  • ethereumEthereum(ETH)$2,675.02-2.74%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$767.04-2.35%
  • rippleXRP(XRP)$1.49-5.85%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$114.44-3.03%
  • tronTRON(TRX)$0.340072-0.40%
  • zcashZcash(ZEC)$1,513.430.03%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.30%
  • HyperliquidHyperliquid(HYPE)$93.43-2.62%
  • dogecoinDogecoin(DOGE)$0.092342-7.46%
  • moneroMonero(XMR)$551.11-2.33%
  • whitebitWhiteBIT Coin(WBT)$84.62-2.39%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.28-5.35%
  • cardanoCardano(ADA)$0.238156-5.31%
  • RainRain(RAIN)$0.012262-6.57%
  • leo-tokenLEO Token(LEO)$9.010.33%
  • stellarStellar(XLM)$0.201936-6.46%
  • bitcoin-cashBitcoin Cash(BCH)$344.001.61%
  • nearNEAR Protocol(NEAR)$4.444.31%
  • uniswapUniswap(UNI)$9.18-1.26%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$61.09-2.37%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.32-5.81%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.109344-4.14%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-2.72%
  • hedera-hashgraphHedera(HBAR)$0.090289-9.17%
  • suiSui(SUI)$0.96-4.06%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.16%
  • BittensorBittensor(TAO)$288.37-6.84%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.061158-8.21%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.20-7.56%
  • BitwayBitway(BTW)$1.0015.35%
  • tether-goldTether Gold(XAUT)$4,289.24-1.62%
  • okbOKB(OKB)$118.00-3.48%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.19%
  • mantleMantle(MNT)$0.65-1.51%
  • aaveAave(AAVE)$139.12-3.28%
  • EthenaEthena(ENA)$0.2077260.47%
  • OndoOndo(ONDO)$0.412632-5.14%
  • AsterAster(ASTER)$0.69-4.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from China Proposes Continuity-Relativity indExing with gAussian Middle (CREAM): A Simple yet Effective AI Method to Extend the Context of Large Language Models

June 16, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper from China Proposes Continuity-Relativity indExing with gAussian Middle (CREAM): A Simple yet Effective AI Method to Extend the Context of Large Language Models
ShareShareShareShareShare

Large language models (LLMs) like transformers are typically pre-trained with a fixed context window size, such as 4K tokens. However, many applications require processing much longer contexts, up to 256K tokens. Extending the context length of these models poses challenges, particularly in ensuring efficient use of information from the middle part of the context, often referred to as the “Lost-in-the-Middle” problem. Existing methods that extend context length often need extensive fine-tuning at the target length and struggle to effectively handle information from the middle of the context. 

Researchers from the Beijing Institute for General Artificial Intelligence (BIGAI), Beijing, China, and the National Key Laboratory of General Artificial Intelligence, Beijing, China, introduce CREAM, ContinuityRelativity indExing with gAussian Middle, to address the challenges in extending the context window of pre-trained LLMs. Current methods to extend the context window of pre-trained LLMs include positional encoding (PE)-based approaches. These methods are based on interpolated positional encodings that require fine-tuning on the target context length, resulting in high computational overhead. Methods like efficient transformers and memory augmentation modify the model architecture or add supplementary modules, complicating implementation and adaptation. 

In contrast, CREAM is designed to extend LLMs to significantly longer context lengths efficiently. It manipulates position indices to interpolate positional encodings within the pre-trained context window size and introduces a truncated Gaussian sampling method to focus on the middle part of the context during fine-tuning. This approach allows the model to be fine-tuned within its pre-trained window size while achieving effective performance on extended contexts up to 256K tokens.

CREAM’s methodology involves two main strategies: ensuring continuity and relativity in positional encoding. For continuity, CREAM manipulates position indices to generate shorter sequences within the pre-trained context window, maintaining densely connected positional indices. For relativity, it leverages rotary positional encoding (RoPE) to learn relative positions between token pairs. Additionally, CREAM divides the pre-trained context window into three segments (head, middle, tail) and uses a truncated Gaussian function to prioritize the middle segment during fine-tuning.

Experiments with Llama-2-7B and Llama-2-7B-Chat models demonstrated CREAM’s efficiency and effectiveness. CREAM extended the context window from 4K up to 256K tokens and showed superior performance in long-context understanding tasks. Specifically, CREAM outperformed existing methods in retrieving information from long contexts and alleviating the “Lost-in-the-Middle” issue. It also achieved promising results in long-context question-answering and summarization tasks, outperforming strong baselines with minimal fine-tuning steps.

In conclusion, CREAM addresses the limitations of current methods by efficiently extending the context length of LLMs while focusing on middle-context information. The proposed method successfully balances continuity and relativity in positional encoding and employs a truncated Gaussian sampling approach to enhance middle-content understanding. Experimental results validate CREAM’s effectiveness in extending context windows and improving performance in long-context scenarios, offering a practical solution to the “Lost-in-the-Middle” problem.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit


YOU MAY ALSO LIKE

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

Disney+ And Hulu Are Getting Even More Expensive (Again)

Pragati Jhunjhunwala is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Kharagpur. She is a tech enthusiast and has a keen interest in the scope of software and data science applications. She is always reading about the developments in different field of AI and ML.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Disney+ And Hulu Are Getting Even More Expensive (Again)
AI & Technology

Disney+ And Hulu Are Getting Even More Expensive (Again)

September 23, 2026
Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age
AI & Technology

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

September 23, 2026
Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Next Post
Amazon workers say they were exploited by labor supply and recruiting firms

Amazon workers say they were exploited by labor supply and recruiting firms

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
XTN: Transportation Likely To Lag Into 2027 Amid Macro Pressures And Factor Weaknesses

XTN: Transportation Likely To Lag Into 2027 Amid Macro Pressures And Factor Weaknesses

September 22, 2026
I Gave GPT-6 & Claude ,000 Each to Trade on Kalshi

I Gave GPT-6 & Claude $1,000 Each to Trade on Kalshi

September 22, 2026
Far-right commentator Milo Yiannopoulos arrested by ICE

Far-right commentator Milo Yiannopoulos arrested by ICE

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!