• bitcoinBitcoin(BTC)$77,332.000.42%
  • ethereumEthereum(ETH)$2,394.11-0.64%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$686.741.17%
  • rippleXRP(XRP)$1.34-0.92%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$99.43-0.02%
  • tronTRON(TRX)$0.3247130.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-0.11%
  • HyperliquidHyperliquid(HYPE)$81.630.12%
  • zcashZcash(ZEC)$812.70-1.73%
  • dogecoinDogecoin(DOGE)$0.081454-0.21%
  • RainRain(RAIN)$0.0167692.09%
  • moneroMonero(XMR)$527.226.57%
  • USDSUSDS(USDS)$1.000.01%
  • leo-tokenLEO Token(LEO)$9.29-0.84%
  • whitebitWhiteBIT Coin(WBT)$70.880.01%
  • chainlinkChainlink(LINK)$11.14-0.53%
  • cardanoCardano(ADA)$0.1970740.44%
  • stellarStellar(XLM)$0.174004-0.92%
  • bitcoin-cashBitcoin Cash(BCH)$244.14-0.65%
  • daiDai(DAI)$1.00-0.02%
  • CantonCanton(CC)$0.109218-4.64%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$49.720.94%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.330.72%
  • uniswapUniswap(UNI)$5.812.64%
  • hedera-hashgraphHedera(HBAR)$0.073935-0.41%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • avalanche-2Avalanche(AVAX)$7.17-0.58%
  • shiba-inuShiba Inu(SHIB)$0.0000050.26%
  • suiSui(SUI)$0.731.31%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • tether-goldTether Gold(XAUT)$4,373.630.95%
  • crypto-com-chainCronos(CRO)$0.054424-0.51%
  • nearNEAR Protocol(NEAR)$1.86-2.88%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.04-2.12%
  • okbOKB(OKB)$106.14-3.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.22%
  • BittensorBittensor(TAO)$217.97-1.67%
  • AsterAster(ASTER)$0.735.67%
  • aaveAave(AAVE)$127.160.95%
  • pax-goldPAX Gold(PAXG)$4,384.110.98%
  • mantleMantle(MNT)$0.564.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056339-1.04%
  • MorphoMorpho(MORPHO)$2.51-0.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Unlocking AI Potential with MINILLM: A Deep Dive into Knowledge Distillation from Larger Language Models to Smaller Counterparts

June 21, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Unlocking AI Potential with MINILLM: A Deep Dive into Knowledge Distillation from Larger Language Models to Smaller Counterparts
ShareShareShareShareShare

Knowledge distillation which involves training a small student model under the supervision of a big teacher model is a typical strategy to decrease excessive computational resource demand due to the fast development of large language models. Black-box KD, in which only the teacher’s predictions are accessible, and white-box KD, in which the teacher’s parameters are used, are the two kinds of KD that are often used. Black-box KD has recently demonstrated encouraging outcomes in optimizing tiny models on the prompt-response pairs produced by LLM APIs. White-box KD becomes increasingly helpful for research communities and industrial sectors when more open-source LLMs are developed since student models get better signals from white-box instructor models, potentially leading to improved performance. 

While white-box KD for generative LLMs has not yet been investigated, it is mostly examined for small (1B parameters) language understanding models. They look into white-box KD of LLMs in this paper. They contend that the common KD could be better for LLMs that carry out tasks generatively. Standard KD objectives (including several variants for sequence-level models) essentially minimize the approximated forward Kullback-Leibler divergence (KLD) between the teacher and the student distribution, known as KL, forcing p to cover all the modes of q given the teacher distribution p(y|x) and the student distribution q(y|x)parameterized by. KL performs well for text classification problems because the output space often contains finite-number classes, ensuring that both p(y|x) and q(y|x) have a small number of modes. 

However, for open text generation problems, where the output spaces are far more complicated, p(y|x) may represent a substantially wider range of modes than q(y|x). During free-run generation, minimizing forward KLD can lead to q giving the void regions of p excessively high probability and producing highly improbable samples under p. They suggest minimizing the reverse KLD, KL, which is commonly employed in computer vision and reinforcement learning, to solve this issue. A pilot experiment shows how underestimating KL drives q to seek the major modes of p and give its vacant areas a low probability. 

🚀 JOIN the fastest ML Subreddit Community

This means that in the language generation of LLMs, the student model avoids learning too many long-tail versions of the instructor distribution and concentrates on the produced response’s accuracy, which is crucial in real-world situations where honesty and dependability are required. They generate the gradient of the objective with Policy Gradient to optimize min KL. Recent studies have demonstrated the effectiveness of policy optimization in optimizing PLMs. However, they also discovered that training the model still suffers from excessive variation, reward hacking, and generation length bias. As a result, they include:

  1. Single-step regularisation to lessen variation.
  2. Teacher-mixed sampling to lessen reward hacking.
  3. Length normalization to reduce length bias. 

In the instruction-following setting, which encompasses a wide range of NLP tasks, researchers from The CoAI Group, Tsinghua University, and Microsoft Research offer a novel technique called MINILLM, which they then apply to several generative language models with parameter sizes ranging from 120M to 13B. Five instruction-following datasets and Rouge-L and GPT-4 feedback for assessment are used. Their tests demonstrate that MINILM scales up successfully from 120M to 13B models and consistently beats the baseline standard KD models on all datasets (see Figure 1). More research reveals that MINILLM works better at producing lengthier replies with more variety and has reduced exposure bias and better calibration. The models are available on GitHub.

Figure 1 shows a comparison of the average GPT-4 feedback score on their assessment sets between MINILLM and the sequence-level KD (SeqKD). GPT-2-1.5B is seen on the left with GPT-2 125M, 340M, and 760M acting as the pupils. Middle: GPT-2 760M, 1.5B, and GPT-Neo 2.7B are the pupils, while GPT-J 6B is the instructor. OPT 13B is seen on the right with OPT 1.3B, 2.7B, and 6.7B as the students.

Check Out The Paper and Github link. Don’t forget to join our 24k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools From AI Tools Club

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs

Modding Platform Nexus Mods Now Owns SteamDB

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs
AI & Technology

Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs

September 2, 2026
Modding Platform Nexus Mods Now Owns SteamDB
AI & Technology

Modding Platform Nexus Mods Now Owns SteamDB

September 2, 2026
Huskeys Raises M Series A to Build the Security Control Layer for the AI Driven Network – Unite.AI
AI & Technology

Huskeys Raises $27M Series A to Build the Security Control Layer for the AI Driven Network – Unite.AI

September 2, 2026
Forward-deployed engineering is how enterprise AI learns
AI & Technology

Forward-deployed engineering is how enterprise AI learns

September 2, 2026
Next Post
SoftBank CEO Masayoshi Son had a crisis of confidence that left him in tears for days

SoftBank CEO Masayoshi Son had a crisis of confidence that left him in tears for days

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Neighbors react to teen suspect in clown costume

Neighbors react to teen suspect in clown costume

August 27, 2026
New video shows moment hero fires at shooter at Idaho In-and-Out

New video shows moment hero fires at shooter at Idaho In-and-Out

August 30, 2026
Nearly 25K pounds of frozen buffalo chicken recalled over inspection lapse

Nearly 25K pounds of frozen buffalo chicken recalled over inspection lapse

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!