• bitcoinBitcoin(BTC)$76,598.000.49%
  • ethereumEthereum(ETH)$2,449.261.42%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$739.132.18%
  • rippleXRP(XRP)$1.300.69%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.513.00%
  • tronTRON(TRX)$0.334862-0.25%
  • zcashZcash(ZEC)$1,464.208.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.07%
  • HyperliquidHyperliquid(HYPE)$86.6710.98%
  • dogecoinDogecoin(DOGE)$0.0819021.65%
  • moneroMonero(XMR)$517.503.30%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$78.890.81%
  • RainRain(RAIN)$0.012703-1.33%
  • chainlinkChainlink(LINK)$11.444.10%
  • leo-tokenLEO Token(LEO)$8.90-0.37%
  • cardanoCardano(ADA)$0.2055915.31%
  • stellarStellar(XLM)$0.183788-0.31%
  • uniswapUniswap(UNI)$7.8016.97%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$235.036.65%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$54.425.04%
  • nearNEAR Protocol(NEAR)$3.1819.56%
  • CantonCanton(CC)$0.1033504.27%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.09%
  • avalanche-2Avalanche(AVAX)$7.641.97%
  • hedera-hashgraphHedera(HBAR)$0.0750322.01%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.756.11%
  • shiba-inuShiba Inu(SHIB)$0.0000055.44%
  • crypto-com-chainCronos(CRO)$0.0580162.80%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.218.39%
  • tether-goldTether Gold(XAUT)$4,357.621.33%
  • BittensorBittensor(TAO)$234.504.23%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$111.701.49%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.18%
  • AsterAster(ASTER)$0.755.26%
  • aaveAave(AAVE)$129.037.78%
  • BitwayBitway(BTW)$0.71-3.07%
  • pax-goldPAX Gold(PAXG)$4,357.251.36%
  • Pump.funPump.fun(PUMP)$0.0040269.51%
  • mantleMantle(MNT)$0.573.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Cohere Enhances Language Model Stability with Automated Detection of Under-trained Tokens in LLMs

May 14, 2024
in AI & Technology
Reading Time: 6 mins read
A A
This AI Paper from Cohere Enhances Language Model Stability with Automated Detection of Under-trained Tokens in LLMs
ShareShareShareShareShare

Tokenization is essential in computational linguistics, particularly in the training and functionality of large language models (LLMs). This process involves dissecting text into manageable pieces or tokens, which is foundational for model training and operations. While effective tokenization can significantly enhance a model’s performance, issues arise when tokens within the model’s vocabulary are underrepresented or absent in the training datasets, leading to what researchers term ‘glitch tokens.’ When encountered in new input data, these tokens can destabilize a model and produce unpredictable outputs.

A prevalent issue in LLMs is the misalignment between tokenizer training and model training. Often, tokenizers are trained separately using distinct datasets, which can differ significantly from the data used to train the model. This disjoint can lead to some of the vocabulary glitch tokens being under-trained. The infamous “_SolidGoldMagikarp” token is a notorious glitch token that can induce unwanted model behaviors, such as hallucinations or producing nonsensical outputs.

Conventional methods for identifying under-trained tokens typically involve manual checks of the tokenizer’s behavior, examining how tokens are encoded and decoded, or analyzing their frequency in the training data. However, these methods are not scalable for the increasingly large and complex LLMs being developed today.

Researchers from Cohere introduce a novel approach that utilizes the model’s embedding weights to automate and scale the detection of under-trained tokens. The researchers developed a method to analyze these weights to spot anomalies indicative of insufficient training. By assessing the embedding matrix of a model, the research identifies tokens whose embedding weights deviate significantly from those of well-represented tokens. This method provides a systematic way to pinpoint glitch tokens by calculating the variance and distribution of embedding weights and comparing them against a normative model of adequately trained tokens.

The study demonstrated the effectiveness of this new method by applying it to several well-known models, including variations of Google’s BERT and OpenAI’s GPT series. The analysis identified a substantial percentage of the tokenizer’s vocabulary, up to 10% in some cases, as under-trained. These tokens were often specialized or infrequently used words, which exhibited the most significant discrepancies in embedding weight patterns.

This research has significant implications for the development and maintenance of LLMs. By employing automated techniques to detect and rectify under-trained tokens, developers can enhance the accuracy and robustness of language models. This advancement is crucial as LLMs are increasingly used in various applications, from automated writing aids to sophisticated conversational agents.

In conclusion, this research highlights a critical vulnerability in LLM training and presents a scalable solution to mitigate this issue. Implementing automated methods for detecting under-trained tokens allows for more robust training processes, ensuring that all tokens in a model’s vocabulary are adequately prepared to handle real-world applications. This research improves the efficacy and reliability of language models, paving the way for more reliable and effective natural language processing tools.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

Aswin AK is a consulting intern at MarkTechPost. He is pursuing his Dual Degree at the Indian Institute of Technology, Kharagpur. He is passionate about data science and machine learning, bringing a strong academic background and hands-on experience in solving real-life cross-domain challenges.


[Recommended Read] Rightsify’s GCX: Your Go-To Source for High-Quality, Ethically Sourced, Copyright-Cleared AI Music Training Datasets with Rich Metadata


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
AI & Technology

Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

September 17, 2026
Next Post
Stay Tuned NOW with Gadi Schwartz – Feb. 6 | NBC News  NOW

Stay Tuned NOW with Gadi Schwartz - Feb. 6 | NBC News NOW

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Thieves crawl through North Carolina restaurant

Thieves crawl through North Carolina restaurant

September 16, 2026
Ex-CNN anchor Brooke Baldwin reveals dad beat her with belt until she was ‘black and blue on my backside’

Ex-CNN anchor Brooke Baldwin reveals dad beat her with belt until she was ‘black and blue on my backside’

September 14, 2026
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!