• bitcoinBitcoin(BTC)$80,468.00-0.68%
  • ethereumEthereum(ETH)$2,577.04-1.91%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$750.97-1.61%
  • rippleXRP(XRP)$1.38-2.96%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$108.74-2.71%
  • tronTRON(TRX)$0.3401910.76%
  • zcashZcash(ZEC)$1,451.06-6.42%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.31%
  • HyperliquidHyperliquid(HYPE)$91.04-1.01%
  • dogecoinDogecoin(DOGE)$0.085372-2.25%
  • moneroMonero(XMR)$521.53-9.13%
  • whitebitWhiteBIT Coin(WBT)$81.85-1.55%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.0134890.26%
  • chainlinkChainlink(LINK)$12.00-2.73%
  • cardanoCardano(ADA)$0.221018-1.45%
  • leo-tokenLEO Token(LEO)$8.900.09%
  • stellarStellar(XLM)$0.191076-1.98%
  • uniswapUniswap(UNI)$8.88-1.15%
  • bitcoin-cashBitcoin Cash(BCH)$247.460.30%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.47-5.42%
  • litecoinLitecoin(LTC)$57.22-0.43%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6113.33%
  • CantonCanton(CC)$0.105245-4.16%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.48%
  • MemeCoreMemeCore(M)$1.6225.41%
  • hedera-hashgraphHedera(HBAR)$0.0810102.22%
  • suiSui(SUI)$0.820.89%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.35%
  • crypto-com-chainCronos(CRO)$0.058722-0.17%
  • BittensorBittensor(TAO)$254.420.29%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,368.55-0.13%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.80-0.36%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.34%
  • aaveAave(AAVE)$137.99-3.03%
  • AsterAster(ASTER)$0.74-4.02%
  • OndoOndo(ONDO)$0.4089952.43%
  • EthenaEthena(ENA)$0.19581911.78%
  • mantleMantle(MNT)$0.59-2.69%
  • pax-goldPAX Gold(PAXG)$4,361.02-0.13%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet mmBERT: An Encoder-only Language Model Pretrained on 3T Tokens of Multilingual Text in over 1800 Languages and 2–4× Faster than Previous Models

September 11, 2025
in AI & Technology
Reading Time: 10 mins read
A A
Meet mmBERT: An Encoder-only Language Model Pretrained on 3T Tokens of Multilingual Text in over 1800 Languages and 2–4× Faster than Previous Models
ShareShareShareShareShare

Why was a new multilingual encoder needed?

XLM-RoBERTa (XLM-R) has dominated multilingual NLP for more than 5 years, an unusually long reign in AI research. While encoder-only models like BERT and RoBERTa were central to early progress, most research energy shifted toward decoder-based generative models. Encoders, however, remain more efficient and often outperform decoders on embedding, retrieval, and classification tasks. Despite this, multilingual encoder development stalled.

A team of researchers from Johns Hopkins University propose mmBERT that addresses this gap by delivering a modern encoder, surpassesing XLM-R and rivals recent large-scale models such as OpenAI’s o3 and Google’s Gemini 2.5 Pro.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

How To Record Audio On Your iPhone

Understanding the architecture of mmBERT

mmBERT comes in two main configurations:

  • Base model: 22 transformer layers, 1152 hidden dimension, ~307M parameters (110M non-embedding).
  • Small model: ~140M parameters (42M non-embedding).

It adopts the Gemma 2 tokenizer with a 256k vocabulary, rotary position embeddings (RoPE), and FlashAttention2 for efficiency. Sequence length is extended from 1024 to 8192 tokens, using unpadded embeddings and sliding-window attention. This allows mmBERT to process contexts nearly an order of magnitude longer than XLM-R while maintaining faster inference.

What training data and phases were used?

mmBERT was trained on 3 trillion tokens spanning 1,833 languages. Data sources include FineWeb2, Dolma, MegaWika v2, ProLong, StarCoder, and others. English makes up only ~10–34% of the corpus depending on the phase.

Training was done in three stages:

  1. Pre-training: 2.3T tokens across 60 languages and code.
  2. Mid-training: 600B tokens across 110 languages, focused on higher-quality sources.
  3. Decay phase: 100B tokens covering 1,833 languages, emphasizing low-resource adaptation.

What new training strategies were introduced?

Three main innovations drive mmBERT’s performance:

  • Annealed Language Learning (ALL): Languages are introduced gradually (60 → 110 → 1833). Sampling distributions are annealed from high-resource to uniform, ensuring low-resource languages gain influence during later stages without overfitting limited data.
  • Inverse Masking Schedule: The masking ratio starts at 30% and decays to 5%, encouraging coarse-grained learning early and fine-grained refinements later.
  • Model Merging Across Decay Variants: Multiple decay-phase models (English-heavy, 110-language, and 1833-language) are combined via TIES merging, leveraging complementary strengths without retraining from scratch.

How does mmBERT perform on benchmarks?

  • English NLU (GLUE): mmBERT base achieves 86.3, surpassing XLM-R (83.3) and nearly matching ModernBERT (87.4), despite allocating >75% of training to non-English data.
  • Multilingual NLU (XTREME): mmBERT base scores 72.8 vs. XLM-R’s 70.4, with gains in classification and QA tasks.
  • Embedding tasks (MTEB v2): mmBERT base ties ModernBERT in English (53.9 vs. 53.8) and leads in multilingual (54.1 vs. 52.4 for XLM-R).
  • Code retrieval (CoIR): mmBERT outperforms XLM-R by ~9 points, though EuroBERT remains stronger on proprietary data.

How does mmBERT handle low-resource languages?

The annealed learning schedule ensures that low-resource languages benefit during later training. On benchmarks like Faroese FoQA and Tigrinya TiQuAD, mmBERT significantly outperforms both o3 and Gemini 2.5 Pro. These results demonstrate that encoder models, if trained carefully, can generalize effectively even in extreme low-resource scenarios.

What efficiency gains does mmBERT achieve?

mmBERT is 2–4× faster than XLM-R and MiniLM while supporting 8192-token inputs. Notably, it remains faster at 8192 tokens than older encoders were at 512 tokens. This speed boost derives from the ModernBERT training recipe, efficient attention mechanisms, and optimized embeddings.

Summary

mmBERT comes as the long-overdue replacement for XLM-R, redefining what a multilingual encoder can deliver. It runs 2–4× faster, handles sequences up to 8K tokens, and outperforms prior models on both high-resource benchmarks and low-resource languages that were underserved in the past. Its training recipe—3 trillion tokens paired with annealed language learning, inverse masking, and model merging—shows how careful design can unlock broad generalization without excessive redundancy. The result is an open, efficient, and scalable encoder that not only fills the six-year gap since XLM-R but also provides a robust foundation for the next generation of multilingual NLP systems.


Check out the Paper, Model on Hugging Face, GitHub and Technical details. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
How To Record Audio On Your iPhone
AI & Technology

How To Record Audio On Your iPhone

September 20, 2026
What Is The Difference Between Apple CarPlay And CarPlay Ultra?
AI & Technology

What Is The Difference Between Apple CarPlay And CarPlay Ultra?

September 19, 2026
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
AI & Technology

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
Next Post
Dassault Aviation: This Fighter And Business Jet Maker Has A Big Runway Ahead (DUAVF)

Dassault Aviation: This Fighter And Business Jet Maker Has A Big Runway Ahead (DUAVF)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Movie theaters unveil special screenings to entice audiences back into seats

Movie theaters unveil special screenings to entice audiences back into seats

September 16, 2026
Officer charged in shooting death of college student

Officer charged in shooting death of college student

September 19, 2026
At least 10 killed in central Mexico fireworks blast

At least 10 killed in central Mexico fireworks blast

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!