• bitcoinBitcoin(BTC)$76,634.000.68%
  • ethereumEthereum(ETH)$2,454.931.92%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$734.242.12%
  • rippleXRP(XRP)$1.30-0.42%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.182.74%
  • tronTRON(TRX)$0.335223-0.09%
  • zcashZcash(ZEC)$1,494.6413.90%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.26%
  • HyperliquidHyperliquid(HYPE)$83.446.45%
  • dogecoinDogecoin(DOGE)$0.0818361.67%
  • moneroMonero(XMR)$515.734.43%
  • USDSUSDS(USDS)$1.000.04%
  • whitebitWhiteBIT Coin(WBT)$78.991.17%
  • RainRain(RAIN)$0.012864-3.29%
  • chainlinkChainlink(LINK)$11.363.59%
  • leo-tokenLEO Token(LEO)$8.920.70%
  • cardanoCardano(ADA)$0.2025203.77%
  • stellarStellar(XLM)$0.1867392.62%
  • uniswapUniswap(UNI)$7.7419.75%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • bitcoin-cashBitcoin Cash(BCH)$233.076.81%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.905.22%
  • CantonCanton(CC)$0.0991995.11%
  • nearNEAR Protocol(NEAR)$3.0016.55%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.53%
  • avalanche-2Avalanche(AVAX)$7.613.58%
  • hedera-hashgraphHedera(HBAR)$0.0756662.89%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000055.54%
  • suiSui(SUI)$0.733.28%
  • crypto-com-chainCronos(CRO)$0.0576982.81%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.196.17%
  • tether-goldTether Gold(XAUT)$4,339.981.79%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$230.144.29%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.581.94%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • AsterAster(ASTER)$0.747.80%
  • aaveAave(AAVE)$129.6110.75%
  • BitwayBitway(BTW)$0.71-3.30%
  • pax-goldPAX Gold(PAXG)$4,339.411.74%
  • mantleMantle(MNT)$0.573.83%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0582381.94%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

‘Inheritune’ by UT Austin Assists Efficient Language Model Training: Leveraging Inheritance and Reduced Data for Comparable Performance

April 21, 2024
in AI & Technology
Reading Time: 4 mins read
A A
‘Inheritune’ by UT Austin Assists Efficient Language Model Training: Leveraging Inheritance and Reduced Data for Comparable Performance
ShareShareShareShareShare

Scaling up LLMs presents significant challenges due to the immense computational resources needed and the need for high-quality datasets. Typically, the pre-training process involves utilizing models with billions of parameters and training them on datasets containing trillions of tokens. This intricate procedure demands substantial computational power and access to high-quality data to achieve better performance in language understanding and generation tasks.

Researchers from UT Austin have developed “Inheritune,” a method to distinguish smaller base LMs from larger ones. They inherit a few transformer blocks from a larger LM and then train the smaller model on a tiny fraction (0.1%) of the original pretraining data. This approach efficiently creates LMs with 1.5 billion parameters using just 1 billion tokens, leveraging a single GPU in under 12 hours. Despite using significantly less data, the resulting models perform comparably to publicly available LMs trained on larger datasets, demonstrating efficacy across various settings.

Previous approaches to training small-base LMs involve extensive training from scratch with trillions of tokens or utilizing high-quality synthetic data. For instance, tinyllama-1B is trained from scratch with 3 trillion tokens over 90 days. In contrast, the Inheritune, efficiently trains small base LMs by inheriting transformer blocks from larger models and training on a small subset of data, achieving comparable performance with significantly fewer computational resources. While model compression techniques have been successful in other domains, such as neural networks, they have yet to be as effective in the complex functions of large LMs.

In the Inheritune approach, a small base LM is crafted by inheriting a fraction of pre-training data and a few layers from an existing large LM. Firstly, the first n layers of the reference model are inherited, initializing the target model. Then, the target model is trained on the available subset of training data for a specified number of epochs. In the experiments, the researchers use a 1 billion token subset of the Redpajama v1 dataset to train a 1.5 billion parameter LM, achieving competitive performance compared to scratch-trained and derived LMs. The researchers evaluate the approach using various baseline models, primarily considering their pre-training data quality for fair comparison.

Inheritance enables the extraction of smaller target LMs without sacrificing performance, showcasing comparable zero-shot performance on relevant downstream tasks. Moreover, these LMs outperform similar-sized models trained from scratch, surpassing them after fewer training steps. Experimentation with GPT2-medium models demonstrates that initialization with Inheritune, particularly with attention and MLP weights, yields superior convergence speed and final validation loss performance. Surprisingly, initializing either attention or MLP weights produces similar improvements in convergence speed and validation loss.

Also, Limitations of the Inheritune method include its inability to modify the architectural design beyond changing the number of transformer blocks, potentially limiting flexibility in customizing hidden sizes and attention heads. Sensitivity to the quality of the training dataset is another concern due to its small size. Additionally, selecting blocks to retain, dataset curation, and hyperparameter tuning still need to explore avenues for improvement. Nevertheless, the study concludes that Inheritune effectively pre-trains small base language models with minimal data and computational resources, offering a straightforward approach to model reduction from large reference models.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


For Content Partnership, Please Fill Out This Form Here..


YOU MAY ALSO LIKE

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

Candy Crush Developers Are Planning A Strike For Next Week

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI
AI & Technology

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

September 17, 2026
Candy Crush Developers Are Planning A Strike For Next Week
AI & Technology

Candy Crush Developers Are Planning A Strike For Next Week

September 17, 2026
Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI
AI & Technology

Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI

September 17, 2026
Lofi Girl Returns With A New House Music Station And Vinyl Compilation
AI & Technology

Lofi Girl Returns With A New House Music Station And Vinyl Compilation

September 17, 2026
Next Post
University of Florida cuts all DEI roles across campus

University of Florida cuts all DEI roles across campus

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Is The Samsung Galaxy S24 Still Worth Buying?

Is The Samsung Galaxy S24 Still Worth Buying?

September 14, 2026
Critics say ‘fashion slop’ is coming for the industry’s design creativity

Critics say ‘fashion slop’ is coming for the industry’s design creativity

September 14, 2026
Raskin says its more important for Democrats to address healthcare

Raskin says its more important for Democrats to address healthcare

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!