• bitcoinBitcoin(BTC)$84,823.000.79%
  • ethereumEthereum(ETH)$2,708.980.83%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$778.490.63%
  • rippleXRP(XRP)$1.54-0.90%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$123.922.81%
  • tronTRON(TRX)$0.333764-0.92%
  • zcashZcash(ZEC)$1,659.588.22%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.68%
  • HyperliquidHyperliquid(HYPE)$93.091.13%
  • dogecoinDogecoin(DOGE)$0.097685-0.39%
  • chainlinkChainlink(LINK)$14.340.56%
  • moneroMonero(XMR)$556.470.68%
  • whitebitWhiteBIT Coin(WBT)$84.610.77%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.255958-0.59%
  • RainRain(RAIN)$0.0126826.10%
  • leo-tokenLEO Token(LEO)$9.020.53%
  • stellarStellar(XLM)$0.217136-1.06%
  • nearNEAR Protocol(NEAR)$5.226.46%
  • bitcoin-cashBitcoin Cash(BCH)$339.11-0.10%
  • uniswapUniswap(UNI)$9.983.19%
  • litecoinLitecoin(LTC)$71.73-2.34%
  • CantonCanton(CC)$0.1373571.25%
  • suiSui(SUI)$1.255.09%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$11.022.36%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.6412.85%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.0947130.21%
  • BittensorBittensor(TAO)$329.903.39%
  • shiba-inuShiba Inu(SHIB)$0.0000060.11%
  • crypto-com-chainCronos(CRO)$0.0692265.18%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BitwayBitway(BTW)$1.1225.18%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.22-0.77%
  • EthenaEthena(ENA)$0.271741-4.16%
  • tether-goldTether Gold(XAUT)$4,280.05-0.02%
  • OndoOndo(ONDO)$0.55-2.13%
  • okbOKB(OKB)$121.870.28%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • quant-networkQuant(QNT)$171.1360.52%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.590.64%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • mantleMantle(MNT)$0.69-1.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.03%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

IBM AI Research Releases Two English Granite Embedding Models, Both Based on the ModernBERT Architecture

September 13, 2025
in AI & Technology
Reading Time: 10 mins read
A A
IBM AI Research Releases Two English Granite Embedding Models, Both Based on the ModernBERT Architecture
ShareShareShareShareShare

IBM has quietly built a strong presence in the open-source AI ecosystem, and its latest release shows why it shouldn’t be overlooked. The company has introduced two new embedding models—granite-embedding-english-r2 and granite-embedding-small-english-r2—designed specifically for high-performance retrieval and RAG (retrieval-augmented generation) systems. These models are not only compact and efficient but also licensed under Apache 2.0, making them ready for commercial deployment.

What Models Did IBM Release?

The two models target different compute budgets. The larger granite-embedding-english-r2 has 149 million parameters with an embedding size of 768, built on a 22-layer ModernBERT encoder. Its smaller counterpart, granite-embedding-small-english-r2, comes in at just 47 million parameters with an embedding size of 384, using a 12-layer ModernBERT encoder.

YOU MAY ALSO LIKE

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

Despite their differences in size, both support a maximum context length of 8192 tokens, a major upgrade from the first-generation Granite embeddings. This long-context capability makes them highly suitable for enterprise workloads involving long documents and complex retrieval tasks.

https://arxiv.org/abs/2508.21085

What’s Inside the Architecture?

Both models are built on the ModernBERT backbone, which introduces several optimizations:

  • Alternating global and local attention to balance efficiency with long-range dependencies.
  • Rotary positional embeddings (RoPE) tuned for positional interpolation, enabling longer context windows.
  • FlashAttention 2 to improve memory usage and throughput at inference time.

IBM also trained these models with a multi-stage pipeline. The process started with masked language pretraining on a two-trillion-token dataset sourced from web, Wikipedia, PubMed, BookCorpus, and internal IBM technical documents. This was followed by context extension from 1k to 8k tokens, contrastive learning with distillation from Mistral-7B, and domain-specific tuning for conversational, tabular, and code retrieval tasks.

How Do They Perform on Benchmarks?

The Granite R2 models deliver strong results across widely used retrieval benchmarks. On MTEB-v2 and BEIR, the larger granite-embedding-english-r2 outperforms similarly sized models like BGE Base, E5, and Arctic Embed. The smaller model, granite-embedding-small-english-r2, achieves accuracy close to models two to three times larger, making it particularly attractive for latency-sensitive workloads.

https://arxiv.org/abs/2508.21085

Both models also perform well in specialized domains:

  • Long-document retrieval (MLDR, LongEmbed) where 8k context support is critical.
  • Table retrieval tasks (OTT-QA, FinQA, OpenWikiTables) where structured reasoning is required.
  • Code retrieval (CoIR), handling both text-to-code and code-to-text queries.

Are They Fast Enough for Large-Scale Use?

Efficiency is one of the standout aspects of these models. On an Nvidia H100 GPU, the granite-embedding-small-english-r2 encodes nearly 200 documents per second, which is significantly faster than BGE Small and E5 Small. The larger granite-embedding-english-r2 also reaches 144 documents per second, outperforming many ModernBERT-based alternatives.

Crucially, these models remain practical even on CPUs, allowing enterprises to run them in less GPU-intensive environments. This balance of speed, compact size, and retrieval accuracy makes them highly adaptable for real-world deployment.

What Does This Mean for Retrieval in Practice?

IBM’s Granite Embedding R2 models demonstrate that embedding systems don’t need massive parameter counts to be effective. They combine long-context support, benchmark-leading accuracy, and high throughput in compact architectures. For companies building retrieval pipelines, knowledge management systems, or RAG workflows, Granite R2 provides a production-ready, commercially viable alternative to existing open-source options.

https://arxiv.org/abs/2508.21085

Summary

In short, IBM’s Granite Embedding R2 models strike an effective balance between compact design, long-context capability, and strong retrieval performance. With throughput optimized for both GPU and CPU environments, and an Apache 2.0 license that enables unrestricted commercial use, they present a practical alternative to bulkier open-source embeddings. For enterprises deploying RAG, search, or large-scale knowledge systems, Granite R2 stands out as an efficient and production-ready option.


Check out the Paper, granite-embedding-small-english-r2 and granite-embedding-english-r2. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
AI & Technology

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

September 27, 2026
Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?
AI & Technology

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

September 27, 2026
Next Post
How Adam Khoo uses StockOracle™ to analyze ASML Stock

How Adam Khoo uses StockOracle™ to analyze ASML Stock

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
STIP: Simple TIPS ETF, For Risk-Averse Investors Concerned About Inflation (NYSEARCA:STIP)

STIP: Simple TIPS ETF, For Risk-Averse Investors Concerned About Inflation (NYSEARCA:STIP)

September 25, 2026
‘I am not running for president,’ Attorney General Todd Blanche says at Iowa State Fair

‘I am not running for president,’ Attorney General Todd Blanche says at Iowa State Fair

September 26, 2026
New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

New report: Utah Valley University tried to warn Charlie Kirk’s staff of risks, but team ignored concerns – The Salt Lake Tribune

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!