• bitcoinBitcoin(BTC)$76,466.001.19%
  • ethereumEthereum(ETH)$2,448.982.23%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$727.941.82%
  • rippleXRP(XRP)$1.290.77%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$100.723.23%
  • tronTRON(TRX)$0.333660-0.53%
  • zcashZcash(ZEC)$1,463.8311.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.022.14%
  • HyperliquidHyperliquid(HYPE)$82.184.98%
  • dogecoinDogecoin(DOGE)$0.0814082.42%
  • moneroMonero(XMR)$510.603.01%
  • USDSUSDS(USDS)$1.000.04%
  • whitebitWhiteBIT Coin(WBT)$78.841.63%
  • RainRain(RAIN)$0.0129181.54%
  • chainlinkChainlink(LINK)$11.284.04%
  • leo-tokenLEO Token(LEO)$8.920.74%
  • cardanoCardano(ADA)$0.2012994.93%
  • stellarStellar(XLM)$0.1847982.53%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • uniswapUniswap(UNI)$7.6020.58%
  • bitcoin-cashBitcoin Cash(BCH)$231.826.72%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.455.80%
  • CantonCanton(CC)$0.0991404.77%
  • nearNEAR Protocol(NEAR)$2.9818.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.343.79%
  • avalanche-2Avalanche(AVAX)$7.584.09%
  • hedera-hashgraphHedera(HBAR)$0.0751213.17%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000057.25%
  • suiSui(SUI)$0.734.83%
  • crypto-com-chainCronos(CRO)$0.0574082.75%
  • paypal-usdPayPal USD(PYUSD)$1.000.03%
  • MemeCoreMemeCore(M)$1.197.10%
  • tether-goldTether Gold(XAUT)$4,346.942.29%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$229.055.32%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.131.83%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.05%
  • AsterAster(ASTER)$0.748.05%
  • aaveAave(AAVE)$127.389.88%
  • BitwayBitway(BTW)$0.71-3.93%
  • pax-goldPAX Gold(PAXG)$4,346.402.25%
  • mantleMantle(MNT)$0.574.41%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0581031.89%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Huawei Introduces a Theoretical Framework Focused on the Memorization Process and Performance Dynamics of Transformer-based Language Models (LMs)

May 19, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper from Huawei Introduces a Theoretical Framework Focused on the Memorization Process and Performance Dynamics of Transformer-based Language Models (LMs)
ShareShareShareShareShare

Transformer-based neural networks have shown great ability to handle multiple tasks like text generation, editing, and question-answering. In many cases, models that use more parameters show better performance measured by perplexity and high accuracies of end tasks. This is the main reason for the development of larger models in industries. However, larger models sometimes result in a bad performance, for example,  the 2B model MiniCPM exhibits comparable capabilities to larger language models, such as Llama2-7B, Mistral-7B, Gemma-7B, and Llama-13B. Moreover, the size of high-quality data available may not keep pace as the computational resources for training larger models increase. 

Current methods to overcome such shortcomings include Scaling laws, Energy-based models, and Hopfield models. In scaling laws, the performance of models increases when there is a scale-up in the models’ size and volume of training data. Energy-based models have become famous as a fundamental modeling tool in different areas of machine learning over the past few decades. The main idea of this method is to model the neural network using a parameterized probability density function to present the distribution in terms of a learnable energy function. The last one is the Hopfield model, in which the classical Hopfield networks were developed as an example of associative memory. 

Researchers from Central Research Institute, 2012 Laboratories Huawei Technologies Co., Ltd. introduced a theoretical framework focused on the memorization process and performance dynamics of transformer-based language models (LMs). Researchers carried out a series of experiments using GPT-2 across different data sizes to overcome the signs of saturation and, at the same time, trained vanilla Transformer models on a dataset consisting of 2M tokens. The results of these experiments validated the theoretical results, offering important theoretical insights on the optimal cross-entropy-loss that can guide and improve decision-making in model training. 

A 12-layer transformer LM is trained using the GPT-2 small tokenizer and architecture on the OpenWebText dataset. This dataset is similar to the WebText dataset used for original GPT-2 model training, which contains 9B tokens from 8,013,769 documents. Using different amounts of data, three models are trained where a subset containing the first 1% (90M) and 0.1% (9M) of the OpenWebText data is created. Further, vanilla transformer models are trained using a small amount of high-quality data that contains pairs of English sentences in declarative formation and is context-free with a vocabulary size of 68 words, where the task is to convert declarative sentences into questions.

The training with 0.1% (9M) of the OpenWebText data shows over-fitting, and the training loss disappears over iterations. This happens because the training samples are not well-separated due to which the model energy decreases to a sum of some delta functions. When the model size is about the order O(D2) and trained on 90M tokens, the model can achieve similar training and validation loss compared to the setting with 9B tokens. Two vanilla Transformers of 6 and 10 layers are trained using a batch size of 8, and the training losses stabilize at a value of around 1 as predicted in Proposition.

In conclusion, researchers presented a theoretical framework focused on the memorization process and performance dynamics of transformer-based language models LMs. In this paper, transformer-based networks are modeled using associative memory, and cross-entropy loss is highlighted for model and data sizes. Also, experiments are carried out by (a) utilizing GPT-2 of different data sizes and (b) training vanilla Transformer models on a dataset of 2M tokens. Finally, a global energy function is created for the layered structure of the transformer models using the majorization-minimization technique.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI

Lofi Girl Returns With A New House Music Station And Vinyl Compilation

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI
AI & Technology

Anthropic Launches Life Sciences Verification Program in Beta – Unite.AI

September 17, 2026
Lofi Girl Returns With A New House Music Station And Vinyl Compilation
AI & Technology

Lofi Girl Returns With A New House Music Station And Vinyl Compilation

September 17, 2026
Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches
AI & Technology

Razer Refreshes The One-Handed Tartarus Pro Keyboard With Improved Switches

September 17, 2026
OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI
AI & Technology

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs – Unite.AI

September 17, 2026
Next Post
Medical world concerned after U.S. sells helium stockpile

Medical world concerned after U.S. sells helium stockpile

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Chick-fil-A is one of America’s toughest franchises to get — and the startup cost may surprise you

Chick-fil-A is one of America’s toughest franchises to get — and the startup cost may surprise you

September 14, 2026
Katie Stein, CEO of ASAPP – Interview Series – Unite.AI

Katie Stein, CEO of ASAPP – Interview Series – Unite.AI

September 15, 2026
Widow breaks tradition of keeping politics out of 9/11 remembrances

Widow breaks tradition of keeping politics out of 9/11 remembrances

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!