• bitcoinBitcoin(BTC)$77,134.000.06%
  • ethereumEthereum(ETH)$2,544.043.38%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$724.511.86%
  • rippleXRP(XRP)$1.360.65%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.291.97%
  • tronTRON(TRX)$0.336615-0.81%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.43%
  • zcashZcash(ZEC)$1,174.183.86%
  • HyperliquidHyperliquid(HYPE)$80.871.26%
  • dogecoinDogecoin(DOGE)$0.0845091.19%
  • RainRain(RAIN)$0.015650-1.77%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$513.980.36%
  • whitebitWhiteBIT Coin(WBT)$80.350.72%
  • chainlinkChainlink(LINK)$11.650.15%
  • leo-tokenLEO Token(LEO)$9.16-0.30%
  • cardanoCardano(ADA)$0.205217-1.75%
  • stellarStellar(XLM)$0.1787340.22%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • bitcoin-cashBitcoin Cash(BCH)$228.181.22%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.05%
  • litecoinLitecoin(LTC)$53.422.41%
  • CantonCanton(CC)$0.098264-0.55%
  • uniswapUniswap(UNI)$6.070.68%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.361.04%
  • nearNEAR Protocol(NEAR)$2.583.19%
  • avalanche-2Avalanche(AVAX)$7.49-1.27%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074283-1.21%
  • shiba-inuShiba Inu(SHIB)$0.0000052.23%
  • suiSui(SUI)$0.73-1.29%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0565380.71%
  • tether-goldTether Gold(XAUT)$4,347.670.09%
  • MemeCoreMemeCore(M)$1.170.93%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.271.92%
  • BittensorBittensor(TAO)$235.98-1.58%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • mantleMantle(MNT)$0.593.00%
  • aaveAave(AAVE)$124.731.94%
  • pax-goldPAX Gold(PAXG)$4,352.200.16%
  • AsterAster(ASTER)$0.69-2.32%
  • polkadotPolkadot(DOT)$1.05-4.39%
  • OndoOndo(ONDO)$0.3535281.59%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Can Large Language Models Retain Old Skills While Learning New Ones? This Paper Introduces LLaMA Pro-8.3B: A New Frontier in AI Adaptability

January 9, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Can Large Language Models Retain Old Skills While Learning New Ones? This Paper Introduces LLaMA Pro-8.3B: A New Frontier in AI Adaptability
ShareShareShareShareShare

Large Language Models (LLMs) have transformed the field of Natural Language Processing (NLP) and the way humans interact with machines. From question answering and text generation to text summarization and code completion, these models have extended their capabilities in a variety of tasks. 

Though LLMs are highly adaptable, their potential as universal language agents is limited in programming, mathematics, the biomedical sciences, and finance. Methods like domain-adaptive pretraining improve LLMs using domain-specific corpora following their first pretraining with a lower computation cost. 

YOU MAY ALSO LIKE

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

However, catastrophic forgetting presents a major obstacle, as post-pretraining causes the model’s initial general abilities to deteriorate. This makes it difficult for the model to function at its optimal level on various tasks. Hence, a technique that adds domain-specific knowledge to LLMs without compromising their overall capabilities is required.

To address this issue, a team of researchers has suggested a new post-pretraining technique called block expansion for LLMs that involves extending Transformer blocks. With this method, the model’s information can be effectively and efficiently added without any catastrophic forgetting. Using duplicate Transformer blocks, this technique includes growing a pre-trained LLM that is available off the shelf. 

While the remaining blocks stay frozen, the recently inserted blocks are exclusively fine-tuned using domain-specific corpora and feature zero-initialized linear layers to aid in identity mapping. An extended pre-trained model that performs well in both general and domain-specific tasks is the outcome of this method.

The team has introduced the family of LLAMA PRO in this study. By experimenting with code and math corpora, LLAMA PRO-8.3B has been developed. Initialized from LLaMA2-7B, this adaptable foundation model performs exceptionally well on a wide range of general tasks, programming, and mathematics. The possibility of catastrophic forgetting has been reduced by fine-tuning the extended blocks only with fresh corpus data, guaranteeing the model’s flexibility and proficiency with both newly learned and pre-existing knowledge.

LLAMA PRO has demonstrated superior performance on multiple benchmarks, as does its instruction-following equivalent, LLAMA PRO – INSTRUCT. They have significantly outperformed current open models in the LLaMA family, demonstrating the models’ great potential for reasoning and handling a variety of tasks as intelligent agents.

The team has summarized their primary contributions as follows.

  1. A new technique called block expansion has been presented for LLMs, making it easier to incorporate new information without sacrificing existing capabilities.
  1. Flexible models like LLAMA PRO and LLAMA PRO – INSTRUCT, which smoothly combine programming and natural languages, have been introduced.
  1. These have excelled in math, programming, and general jobs, demonstrating the models’ adaptability.
  1. LLAMA PRO family has been thoroughly benchmarked on a variety of datasets that include both agent-oriented and traditional workloads.
  1. LLAMA PRO’s superiority and enormous potential have been demonstrated in handling more complicated and wide-ranging applications.

In conclusion, this study has provided important new insights into the interplay between programming and natural languages, providing a solid basis for creating sophisticated language agents that can function well in various settings. The results have highlighted how crucial it is to overcome the flaws in LLMs’ processes for learning new skills and point the way towards a viable path for developing more flexible and powerful language models.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables
AI & Technology

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

September 11, 2026
Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI
AI & Technology

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

September 11, 2026
Upgraded In All The Right Places
AI & Technology

Upgraded In All The Right Places

September 11, 2026
Apple’s iPhone Handoff Feature Will Cost You  A Month On T-Mobile
AI & Technology

Apple’s iPhone Handoff Feature Will Cost You $5 A Month On T-Mobile

September 11, 2026
Next Post
Can Large Language Models Learn New Tricks? This Machine Learning Research from Google Introduces ‘CALM’: A Novel Approach for Enhancing AI Capabilities Through Composition

Can Large Language Models Learn New Tricks? This Machine Learning Research from Google Introduces 'CALM': A Novel Approach for Enhancing AI Capabilities Through Composition

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
TikTok influencer killed by husband in murder-suicide

TikTok influencer killed by husband in murder-suicide

September 4, 2026
IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

September 10, 2026
Berlin police kill suspect in Pride festival attack

Berlin police kill suspect in Pride festival attack

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!