• bitcoinBitcoin(BTC)$85,876.005.75%
  • ethereumEthereum(ETH)$2,761.795.35%
  • tetherTether(USDT)$1.000.04%
  • binancecoinBNB(BNB)$798.595.16%
  • rippleXRP(XRP)$1.496.59%
  • usd-coinUSDC(USDC)$1.000.04%
  • solanaSolana(SOL)$117.757.65%
  • tronTRON(TRX)$0.3448940.15%
  • zcashZcash(ZEC)$1,496.003.21%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • HyperliquidHyperliquid(HYPE)$92.730.12%
  • dogecoinDogecoin(DOGE)$0.09650111.48%
  • moneroMonero(XMR)$567.824.36%
  • whitebitWhiteBIT Coin(WBT)$86.484.46%
  • RainRain(RAIN)$0.0140944.66%
  • chainlinkChainlink(LINK)$13.004.88%
  • USDSUSDS(USDS)$1.000.04%
  • cardanoCardano(ADA)$0.2443317.69%
  • leo-tokenLEO Token(LEO)$8.89-0.43%
  • stellarStellar(XLM)$0.2093037.13%
  • uniswapUniswap(UNI)$8.941.14%
  • bitcoin-cashBitcoin Cash(BCH)$264.395.73%
  • nearNEAR Protocol(NEAR)$4.06-2.09%
  • avalanche-2Avalanche(AVAX)$11.06-0.61%
  • Ethena USDeEthena USDe(USDE)$1.000.06%
  • litecoinLitecoin(LTC)$62.818.66%
  • CantonCanton(CC)$0.1174259.20%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.03%
  • suiSui(SUI)$1.0320.05%
  • hedera-hashgraphHedera(HBAR)$0.0929388.09%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.433.63%
  • shiba-inuShiba Inu(SHIB)$0.0000068.50%
  • MemeCoreMemeCore(M)$1.49-3.44%
  • BittensorBittensor(TAO)$289.3711.19%
  • Global DollarGlobal Dollar(USDG)$1.000.03%
  • crypto-com-chainCronos(CRO)$0.0634317.46%
  • paypal-usdPayPal USD(PYUSD)$1.000.05%
  • tether-goldTether Gold(XAUT)$4,351.13-0.51%
  • okbOKB(OKB)$122.654.47%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9025.87%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • aaveAave(AAVE)$142.994.44%
  • OndoOndo(ONDO)$0.4490825.97%
  • EthenaEthena(ENA)$0.212960-1.39%
  • mantleMantle(MNT)$0.646.27%
  • pepePepe(PEPE)$0.00000523.27%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Adam-mini: A Memory-Efficient Optimizer Revolutionizing Large Language Model Training with Reduced Memory Usage and Enhanced Performance

July 2, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Adam-mini: A Memory-Efficient Optimizer Revolutionizing Large Language Model Training with Reduced Memory Usage and Enhanced Performance
ShareShareShareShareShare

The field of research focuses on optimizing algorithms for training large language models (LLMs), which are essential for understanding and generating human language. These models are critical for various applications, including natural language processing and artificial intelligence. Training LLMs requires significant computational resources and memory, making optimizing these processes a high-priority area for researchers.

The primary problem addressed by this paper is the high memory demand of optimization algorithms used in training large language models. Specifically, the Adam optimizer, a standard in the field due to its superior performance, requires substantial memory to store optimizer states such as first-order and second-order momentum values. This memory demand doubles the necessary resources compared to the model size, creating a significant burden. As a result, training large models becomes expensive and less accessible to researchers with limited resources. Alternative methods like Adafactor attempt to reduce memory usage but often compromise performance, highlighting the need for more efficient solutions.

YOU MAY ALSO LIKE

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

How To Choose The Right USB To USB-C Adapter

The Adam optimizer is widely used for training LLMs because of its ability to handle various model sizes and tasks effectively. However, Adam’s requirement for extensive memory to store its optimizer states, particularly the first-order and second-order momentums, poses a considerable challenge. For instance, training a 7 billion parameter model with Adam requires about 56 GB per card for these states alone, totaling 86 GB when gradients are included. This makes training prohibitively expensive, even with advanced graphical cards like the A100-80GB. Additionally, CPU-offloading and sharding are employed to manage this high memory requirement, increasing latency and slowing down the training process.

Researchers from The Chinese University of Hong Kong, Shenzhen, Shenzhen Research Institute of Big Data, Duke University, and Stanford University introduced Adam-mini, an optimizer designed to achieve similar or better performance than Adam while reducing memory usage by 45% to 50%. Adam-mini accomplishes this by partitioning model parameters into blocks based on the Hessian structure of transformers. Each block is then assigned a single high-quality learning rate, significantly reducing the number of learning rates from billions to a manageable number. This approach allows Adam-mini to maintain or even improve performance with a fraction of the memory required by Adam.

Adam-mini works by leveraging the near-block diagonal structure of transformers’ Hessians, partitioning parameters into blocks such as Query, Key, Value, and MLP layers. For each block, a single effective learning rate is calculated using the average of Adam’s second-order momentum values in that block. This method reduces the memory footprint and simplifies the learning rate assignment process. For example, during the pre-training of Llama2-7B on two A800-80GB GPUs, Adam-mini achieved a throughput of 5572.19 tokens per second, compared to 3725.59 tokens per second with AdamW, representing a 49.6% increase. This efficiency results in a 33% reduction in wall-clock time for processing the same number of tokens.

The researchers validated Adam-mini’s performance across various language models ranging from 125 million to 7 billion parameters, including pre-training, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF). The optimizer demonstrated on-par or superior performance to AdamW, with notable improvements in memory efficiency and training speed. For instance, in supervised fine-tuning and reinforcement learning tasks, Adam-mini consistently outperformed AdamW, achieving higher evaluation scores and faster convergence.

In conclusion, the Adam-mini optimizer addresses the significant memory inefficiencies of traditional optimization methods like Adam by introducing a novel partitioning strategy based on the Hessian structure of models. This innovative approach results in substantial memory savings and improved training efficiency, making it a valuable tool for researchers working with large-scale language models. By reducing the memory footprint by up to 50% and increasing throughput by nearly 50%, Adam-mini not only enhances the feasibility of training large models but also encourages broader participation from researchers with limited GPU resources.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI
AI & Technology

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

September 21, 2026
How To Choose The Right USB To USB-C Adapter
AI & Technology

How To Choose The Right USB To USB-C Adapter

September 21, 2026
A Laptop That Works Better With Your Android Phone
AI & Technology

A Laptop That Works Better With Your Android Phone

September 21, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI
AI & Technology

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
Next Post
Vector database company Qdrant wants RAG to be more cost-effective

Vector database company Qdrant wants RAG to be more cost-effective

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Steve Kornacki previews generational change fights in Massachusetts primary

Steve Kornacki previews generational change fights in Massachusetts primary

September 20, 2026
Frozen food giant shuts California factory, eliminating 260 jobs

Frozen food giant shuts California factory, eliminating 260 jobs

September 19, 2026
Snap Makes Its Case for Wearing A Computer On Your Face

Snap Makes Its Case for Wearing A Computer On Your Face

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!