• bitcoinBitcoin(BTC)$81,270.000.00%
  • ethereumEthereum(ETH)$2,627.400.31%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.55-0.11%
  • rippleXRP(XRP)$1.41-0.07%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$110.96-1.99%
  • tronTRON(TRX)$0.3398780.46%
  • zcashZcash(ZEC)$1,473.81-6.39%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.49%
  • HyperliquidHyperliquid(HYPE)$92.09-1.02%
  • dogecoinDogecoin(DOGE)$0.087631-0.29%
  • moneroMonero(XMR)$541.82-5.84%
  • whitebitWhiteBIT Coin(WBT)$82.88-0.53%
  • RainRain(RAIN)$0.0137362.56%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.400.38%
  • cardanoCardano(ADA)$0.2288790.46%
  • leo-tokenLEO Token(LEO)$8.91-0.06%
  • stellarStellar(XLM)$0.1958820.62%
  • uniswapUniswap(UNI)$8.63-3.18%
  • bitcoin-cashBitcoin Cash(BCH)$254.40-0.95%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.6427.80%
  • nearNEAR Protocol(NEAR)$3.55-4.56%
  • daiDai(DAI)$1.00-0.01%
  • litecoinLitecoin(LTC)$58.12-1.37%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.108156-3.27%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.11%
  • hedera-hashgraphHedera(HBAR)$0.0822033.65%
  • suiSui(SUI)$0.875.22%
  • MemeCoreMemeCore(M)$1.5316.86%
  • shiba-inuShiba Inu(SHIB)$0.0000061.26%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BittensorBittensor(TAO)$267.525.86%
  • crypto-com-chainCronos(CRO)$0.059665-0.62%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • tether-goldTether Gold(XAUT)$4,373.74-0.03%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$117.490.60%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.27%
  • aaveAave(AAVE)$140.72-1.71%
  • EthenaEthena(ENA)$0.20745721.75%
  • mantleMantle(MNT)$0.62-0.90%
  • AsterAster(ASTER)$0.76-2.22%
  • OndoOndo(ONDO)$0.4216504.86%
  • Pump.funPump.fun(PUMP)$0.004214-0.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Presents a Direct Experimental Comparison between 8B-Parameter Mamba, Mamba-2, Mamba-2-Hybrid, and Transformer Models Trained on Upto 3.5T Tokens

June 19, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Presents a Direct Experimental Comparison between 8B-Parameter Mamba, Mamba-2, Mamba-2-Hybrid, and Transformer Models Trained on Upto 3.5T Tokens
ShareShareShareShareShare

Transformer-based Large Language Models (LLMs) have emerged as the backbone of Natural Language Processing (NLP). These models have shown remarkable performance over a variety of NLP tasks. The creative self-attention mechanism that enables effective all-to-all communication between tokens in a sequence is primarily responsible for their success. Transformers have become a leading NLP research tool because of this approach and its capacity to expand both model and dataset sizes.

However, self-attention layers are not without restrictions, especially when working with lengthy sequences. The self-attention computational load grows quadratically with the sequence length during training. A large key-value cache is required to hold the state since the memory demand at inference time increases linearly with the number of previous tokens. Numerous attempts have been made to optimize self-attention layers in response to these efficiency difficulties. Still, these attempts are not up to the language modeling power of conventional self-attention.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

Selective state-space models (SSMs) such as Mamba solve some of the fundamental limitations associated with Transformers. Because of the key-value cache, transformers have quadratic computational complexity in relation to sequence length and high memory requirements during inference. SSMs provide a better, more effective solution by reducing these problems. Recent studies have shown that SSMs can compete with Transformers, if not outperform them, in language modeling tasks, making them a reasonable alternative.

Previous studies comparing SSMs and Transformers have mostly focused on small-scale trials using models with less than 3 billion parameters and training on datasets smaller than 1 trillion tokens, despite the good results. A team of researchers has recently performed a thorough comparison using 8-billion-parameter models of Mamba, Mamba-2, and Transformers, all trained on datasets up to 3.5 trillion tokens, in order to properly comprehend the performance of these architectures at greater sizes. 

The team has also incorporated an 8-billion-parameter hybrid model, called Mamba-2-Hybrid that consists of 50% MLP layers, 7% self-attention, and 43% Mamba-2. To find out if Mamba models could compete with Transformer models when given more training resources, the team evaluated them across a wide range of natural language tasks. The results showed that on several tasks, pure SSM models, including Mamba and Mamba-2, either matched or outperformed Transformers. 

However, these models failed on tasks that required considerable long-context reasoning and tasks that required strong copying or in-context learning, like the five-shot MMLU and Phonebook Lookup tasks. On all 12 assessed standard tasks, the 8-billion-parameter Mamba-2-Hybrid model outperformed the 8-billion-parameter Transformer, with an average improvement of 2.65 points. During inference, the hybrid model demonstrated the capacity to generate tokens up to eight times faster.

The team has expanded their studies to incorporate versions of the Mamba-2-Hybrid and Transformer models that allow sequence lengths of 16K, 32K, and 128K in order to evaluate long-context capabilities further. The hybrid model continued to perform on par with or better than the Transformer on average across 23 additional long-context tasks.  As part of NVIDIA’s Megatron-LM project, the team has released code.


Check out the Paper and Code. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit


Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
AI & Technology

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI
AI & Technology

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
Next Post
Meet the Press NOW — Dec. 5

Meet the Press NOW — Dec. 5

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
President Trump calls Clancy case a ‘horrible tragedy’

President Trump calls Clancy case a ‘horrible tragedy’

September 17, 2026
Raymond Horsch posted ads on Craigslist, preyed on prostitutes: Victim’s cousin – newsnationnow.com

Raymond Horsch posted ads on Craigslist, preyed on prostitutes: Victim’s cousin – newsnationnow.com

September 18, 2026
Hurricane Lowell barrels towards Hawaii

Hurricane Lowell barrels towards Hawaii

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!