• bitcoinBitcoin(BTC)$77,214.000.19%
  • ethereumEthereum(ETH)$2,504.61-0.43%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$721.07-0.58%
  • rippleXRP(XRP)$1.35-0.59%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.89-0.43%
  • tronTRON(TRX)$0.3413180.49%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-0.31%
  • zcashZcash(ZEC)$1,089.35-2.22%
  • HyperliquidHyperliquid(HYPE)$78.11-1.72%
  • dogecoinDogecoin(DOGE)$0.084149-0.61%
  • RainRain(RAIN)$0.015294-2.66%
  • moneroMonero(XMR)$531.50-0.26%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.120.07%
  • chainlinkChainlink(LINK)$11.41-0.50%
  • leo-tokenLEO Token(LEO)$9.06-0.52%
  • cardanoCardano(ADA)$0.2081420.52%
  • stellarStellar(XLM)$0.179252-0.35%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.71-0.76%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.742.10%
  • uniswapUniswap(UNI)$6.25-0.57%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-1.45%
  • CantonCanton(CC)$0.095615-1.23%
  • hedera-hashgraphHedera(HBAR)$0.0760962.40%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.400.40%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.87%
  • nearNEAR Protocol(NEAR)$2.33-1.01%
  • suiSui(SUI)$0.72-0.19%
  • crypto-com-chainCronos(CRO)$0.058172-0.39%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.39-0.09%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-3.20%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.590.21%
  • BittensorBittensor(TAO)$234.471.31%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.06%
  • aaveAave(AAVE)$126.230.35%
  • BitwayBitway(BTW)$0.7127.33%
  • AsterAster(ASTER)$0.702.06%
  • pax-goldPAX Gold(PAXG)$4,348.83-0.16%
  • mantleMantle(MNT)$0.56-0.85%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056929-1.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Tsinghua University and Microsoft AI Unveil a Breakthrough in Language Model Training: The Path to Optimal Learning Efficiency

March 4, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from Tsinghua University and Microsoft AI Unveil a Breakthrough in Language Model Training: The Path to Optimal Learning Efficiency
ShareShareShareShareShare

With the rise of language models, there has been an enormous focus on improving the learning of  LMs to accelerate the learning speed and achieve a certain model performance with as few training steps as possible. This emphasis aids humans in understanding the boundaries of LMs amidst their escalating computational requirements. It also advances the democratization of large language models (LLMs), benefiting research and industry communities.

Prior works like Pre-Trained Models, Past, Present, and Future, focus on designing effective architectures, utilizing rich contexts, and improving computational efficiency. In h2oGPT: Democratizing Large Language Models, the researchers have tried to create open-source alternatives to the closed-source approaches. In Large Batch Optimization for Deep Learning: Training BERT in 76 minutes, they tried to overcome the computational challenge of LLMs.  These prior works explore practical acceleration methods at the model, optimizer, or data levels.

The researchers from the CoAI Group, Tsinghua University, and Microsoft Research have proposed a theory for optimizing LM learning, beginning with maximizing the data compression ratio. They derive the Learning Law theorem to elucidate optimal learning dynamics. Validation experiments on linear classification and language modeling tasks confirm the theorem’s properties. Results indicate that optimal LM learning enhances coefficients in LM scaling laws, offering promising implications for practical learning acceleration methods.

In their method (Optimal Learning of Language Models), the researchers demonstrated the principles of optimizing the LM learning speed, including the optimization objective, the property of optimal learning dynamics, and the essential improvement of the learning acceleration. For the optimization objective, they have proposed to minimize the area under the curve (AUC), a learning process with the smallest loss AUC corresponds to the highest compression ratio. Then, they derived the Learning Law theorem that characterizes the property of dynamics in the LM learning process that achieves the optimum of their objective. Here, a learning policy induces a learning process that determines which data points the LM learns as the training progresses.

After conducting experiments on linear classification with Perceptron and language modeling with Transformer, researchers optimized learning policies and validated them empirically. Near-optimal policies significantly accelerated learning, improving loss AUC by 5.50× and 2.41× for Perceptron and Transformer, respectively. Results confirmed theoretical predictions, demonstrating improved scaling law coefficients by up to 96.6% and 21.2%, promising faster LM training with practical significance.

In conclusion, researchers from the CoAI Group, Tsinghua University, and Microsoft Research have proposed a theory for optimizing LM learning to maximize compression ratio. They derive the Learning Law theorem, confirming that all examples contribute equally to optimal learning, validated in experiments. The optimal process improves LM scaling law coefficients, guiding future acceleration methods. 


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 38k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel

You may also like our FREE AI Courses….


YOU MAY ALSO LIKE

Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why
AI & Technology

Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why

September 13, 2026
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth
AI & Technology

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

September 13, 2026
If Your Laptop Trackpad Is Popping Out, Stop Using It Immediately
AI & Technology

If Your Laptop Trackpad Is Popping Out, Stop Using It Immediately

September 13, 2026
How To Get Your Cut Of PlayStation’s .85 Million Settlement
AI & Technology

How To Get Your Cut Of PlayStation’s $7.85 Million Settlement

September 13, 2026
Next Post
Is This Hacky or Tacky?

Is This Hacky or Tacky?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Apple Set for Major Product Launch with Foldable iPhone

Apple Set for Major Product Launch with Foldable iPhone

September 12, 2026
TikTok rejects Meta ads urging firm to join landmark child safety settlement: report

TikTok rejects Meta ads urging firm to join landmark child safety settlement: report

September 11, 2026
New details reveal Leon Black’s friendship with Jeffrey Epstein

New details reveal Leon Black’s friendship with Jeffrey Epstein

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!