• bitcoinBitcoin(BTC)$85,510.005.26%
  • ethereumEthereum(ETH)$2,729.342.68%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$788.541.64%
  • rippleXRP(XRP)$1.516.77%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.504.79%
  • tronTRON(TRX)$0.3471541.26%
  • zcashZcash(ZEC)$1,447.37-4.32%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$92.27-0.66%
  • dogecoinDogecoin(DOGE)$0.09860511.92%
  • moneroMonero(XMR)$577.82-0.06%
  • whitebitWhiteBIT Coin(WBT)$85.973.54%
  • RainRain(RAIN)$0.013827-2.30%
  • chainlinkChainlink(LINK)$12.872.24%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2444386.15%
  • leo-tokenLEO Token(LEO)$8.960.40%
  • stellarStellar(XLM)$0.2118437.22%
  • nearNEAR Protocol(NEAR)$4.443.19%
  • uniswapUniswap(UNI)$9.094.10%
  • bitcoin-cashBitcoin Cash(BCH)$263.593.82%
  • avalanche-2Avalanche(AVAX)$11.120.11%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • litecoinLitecoin(LTC)$60.212.26%
  • CantonCanton(CC)$0.1164934.53%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.0411.70%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.443.69%
  • hedera-hashgraphHedera(HBAR)$0.0913466.09%
  • BittensorBittensor(TAO)$315.6318.17%
  • shiba-inuShiba Inu(SHIB)$0.0000067.15%
  • crypto-com-chainCronos(CRO)$0.0670068.79%
  • MemeCoreMemeCore(M)$1.44-3.77%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,343.72-0.44%
  • okbOKB(OKB)$121.891.92%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.53%
  • BitwayBitway(BTW)$0.8210.06%
  • aaveAave(AAVE)$141.922.78%
  • mantleMantle(MNT)$0.645.39%
  • OndoOndo(ONDO)$0.4367502.13%
  • EthenaEthena(ENA)$0.209406-1.77%
  • pepePepe(PEPE)$0.00000525.32%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Can You Build Large Language Models Like ChatGPT At Half Cost?

May 11, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Can You Build Large Language Models Like ChatGPT At Half Cost?
ShareShareShareShareShare

Large Language Models (LLMs) like GPT-3 and ChatGPT have revolutionized AI by offering Natural Language Understanding and content generation capabilities. But their development comes at a hefty price limiting accessibility and further research. Researchers estimate that training GPT-3 cost OpenAI around $5 million. Nevertheless, Microsoft recognized the potential and invested $1 billion in 2019 and $10 billion in 2023 in OpenAI’s GPT-3 and ChatGPT venture.

LLMs are machine learning models trained on extensive textual data for NLP applications. They are based on transformer architecture and utilize attention mechanisms for NLP tasks like question-answering, machine translation, sentiment analysis, etc.

YOU MAY ALSO LIKE

Why It’s Important To Unplug Your PC During A Power Outage

Why Is Your Laptop Fan So Loud?

The question arises: can the efficiency of these large models be increased while simultaneously reducing computational cost and training time?

Several approaches, like Progressive Neural Networks, Network Morphism, intra-layer model parallelism, knowledge inheritance, etc., have been developed to reduce the computational cost of training neural networks. The novel LiGO (Linear Growth Operator) approach we will discuss is setting a new benchmark. It halves the computational cost of training LLMs.

Before discussing this technique, examining the factors contributing to the high price of making LLMs is essential.

Cost of Building Large Language Models

Three major expenses for developing LLMs are as follows:

1. Computational Resources

Building LLMs require massive computational resources to train on large datasets. They must process billions of parameters and learn complex patterns from massive textual data.

Investment in specialized hardware such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs) is required for building and training LLMs to achieve state-of-the-art performance.

For instance, GPT-3 was trained on a supercomputer with 10000 enterprise-grade GPUs (H100 and A100) and 285,000 CPU cores.

2. Energy Consumption

The intensive computational resources required for building LLMs result in significant energy consumption. For instance, training 175 billion parameters GPT-3 took 14.8 days using 10,000 V100 GPUs, equivalent to 3.55 million GPU hours. Such a high level of energy consumption has significant environmental effects as well.

3. Data Storage & Management

LLMs are trained on large datasets. For instance, GPT-3 was trained on a vast corpus of textual data, including Common Crawl, WebText2, Books1, Books2, and Wikipedia, among other sources. Significant infrastructure investment is required to collect, curate and store these datasets.

Also, cloud storage is required for data storage, and human expertise for data preprocessing and version control. Moreover, ensuring that your data strategy complies with regulations like GDPR also adds to the cost.

LiGO Technique: Reduce the Cost of Building Large Language Models to Half

LiGO (Linear Growth Operator) is a novel technique developed by researchers at MIT to reduce the computational cost of training LLMs by 50%. The method involves initializing the weights of larger models from those of smaller pre-trained models, enabling efficient scaling of neural networks.

Image from the Paper: Learning to Grow Pretrained Models For Efficient Transformer Training

Yoon Kim, the senior author of the paper, says:

“It’s been estimated that training models at the scale of what ChatGPT is hypothesized to run on could take millions of dollars just for a single training run. Can we improve the efficiency of these training methods, so we can still get good models in less time and for less money? We propose to do this by leveraging smaller language models that have previously been trained.”

This method maintains the performance benefits of larger models with reduced computational cost and training time compared to training a large model from scratch. LiGO utilizes a data-driven linear growth operator that combines depth and width operators for optimum performance.

The paper utilized various datasets to conduct text-based experiments, including the English Wikipedia corpus for training BERT and RoBERTa models and the C4 dataset for training GPT2.

The LiGO technique experimentation included growing BERT-Small to BERT-Base, BERT-Base to BERT-Large, RoBERTaSmall to RoBERTa-Base, GPT2-Base to GPT2-Medium, and CaiT-XS to CaiT-S.

The researchers compared their approach with several other baselines, including training from scratch, progressive training, bert2BERT, and KI.

LiGO technique offered 44.7% savings in FLOPs (floating-point operations per second) and 40.7% savings in wall time compared to training BERT-Base from scratch by reusing the BERT-Small model. LiGO growth operator outperforms StackBERT, MSLT, bert2BERT, and KI in efficient training.

Benefits of Using a Training Optimization Technique Like LiGO

LiGO is an efficient neural network training method that has various benefits listed as follows:

1. Faster Training

As stated earlier, faster training is the main advantage of the LiGO technique. It trains LLMs in half the time, increasing productivity and reducing costs.

2. Resource Efficient

LiGO is resource-efficient since it minimizes wall time and FLOPs, leading to a more cost-effective and eco-friendly approach to training large transformer models.

3. Generalization

The LiGO technique has improved the performance of both language and vision transformers suggesting that it is a generalizable technique that can be applied to various tasks.

Building commercial AI products is just one facet of the overall expenses associated with AI systems. Another significant component of costs comes from daily operations. For instance, it costs OpenAI about $700,000 every day to answer queries using ChatGPT. Researchers are expected to continue exploring approaches that make LLMs cost-effective during training and more accessible on runtime.

For more AI-related content, visit unite.ai.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Next Post
FTX Contagion | Bloomberg Technology  11/14/2022

FTX Contagion | Bloomberg Technology 11/14/2022

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
SpaceX soars 5% after Starship reusable rocket launch date revealed

SpaceX soars 5% after Starship reusable rocket launch date revealed

September 16, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Current with Christine Romans – Sept. 2 | NBC News NOW

Current with Christine Romans – Sept. 2 | NBC News NOW

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!