• bitcoinBitcoin(BTC)$66,185.002.96%
  • ethereumEthereum(ETH)$1,928.133.00%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$576.571.74%
  • usd-coinUSDC(USDC)$1.000.01%
  • rippleXRP(XRP)$1.133.32%
  • solanaSolana(SOL)$78.102.15%
  • tronTRON(TRX)$0.3273360.47%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.68%
  • HyperliquidHyperliquid(HYPE)$62.703.26%
  • dogecoinDogecoin(DOGE)$0.0734651.70%
  • USDSUSDS(USDS)$1.000.01%
  • RainRain(RAIN)$0.013886-2.45%
  • zcashZcash(ZEC)$538.521.12%
  • leo-tokenLEO Token(LEO)$9.710.35%
  • whitebitWhiteBIT Coin(WBT)$57.702.81%
  • stellarStellar(XLM)$0.1916012.57%
  • chainlinkChainlink(LINK)$8.693.49%
  • cardanoCardano(ADA)$0.1743796.93%
  • moneroMonero(XMR)$345.312.97%
  • CantonCanton(CC)$0.1261191.40%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$224.425.45%
  • USD1USD1(USD1)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.450.98%
  • litecoinLitecoin(LTC)$47.350.62%
  • Global DollarGlobal Dollar(USDG)$1.00-0.16%
  • suiSui(SUI)$0.773.04%
  • hedera-hashgraphHedera(HBAR)$0.0683113.48%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • avalanche-2Avalanche(AVAX)$6.620.95%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0579810.83%
  • nearNEAR Protocol(NEAR)$1.992.60%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000041.33%
  • tether-goldTether Gold(XAUT)$4,054.940.95%
  • uniswapUniswap(UNI)$3.696.75%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.34%
  • OndoOndo(ONDO)$0.40051214.83%
  • BittensorBittensor(TAO)$199.912.76%
  • pax-goldPAX Gold(PAXG)$4,051.310.94%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056604-0.48%
  • okbOKB(OKB)$81.851.92%
  • AsterAster(ASTER)$0.631.43%
  • HTX DAOHTX DAO(HTX)$0.0000020.84%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.17-5.07%
  • usddUSDD(USDD)$1.000.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB

July 17, 2026
in AI & Technology
Reading Time: 15 mins read
A A
NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB
ShareShareShareShareShare

Embedding models decide which passages an agent ever sees. NVIDIA released Nemotron 3 Embed model to work on that layer. It targets production-scale RAG, agentic retrieval, code retrieval, and agent memory.

What is Nemotron 3 Embed?

The model collection includes three open checkpoints. Nemotron-3-Embed-8B-BF16 is the accuracy-first option. Nemotron-3-Embed-1B-BF16 carries the same design into a smaller footprint. Nemotron-3-Embed-1B-NVFP4 is the Blackwell-optimized 4-bit path.

YOU MAY ALSO LIKE

Sony Files Another Lawsuit Against AI Music Generator Udio

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

All three are transformer encoders trained with bidirectional attention masking. The final embedding comes from average pooling over token-level representations. Maximum sequence length is 32,768 tokens on every checkpoint.

Each model was evaluated across 34 languages. All three carry the OpenMDW License Agreement, version 1.1 (OpenMDW-1.1). Notably, the bases are Mistral models. The 8B is built with Ministral-3-8B-Instruct-2512. Both 1B variants use Ministral-3-3B-Instruct-2512.

Performance

Nemotron-3-Embed-8B-BF16 ranks #1 overall on RTEB (as of July 17 2026), the Retrieval Embedding Benchmark. Evaluation covers its 16 public tasks. Every figure below is average NDCG@10, at model sequence length 4096.

Model Params Emb dim RTEB ViDoRe-V3 text MMTEB (Retrieval)
Nemotron-3-Embed-8B-BF16 ~8B 4096 78.46 60.60 75.45
Nemotron-3-Embed-1B-BF16 1.14B 2048 72.38 57.74 71.04
Nemotron-3-Embed-1B-NVFP4 1.14B 2048 72.00 — —
llama-nemotron-embed-vl-1b-v2 — — 61.98 52.54 59.71
llama-nemotron-embed-1b-v2 — — 60.47 52.10 59.58

Two gaps are worth noting. The 1B gains 10.4 RTEB points over llama-nemotron-embed-vl-1b-v2, the prior-generation baseline. Separately, NVFP4 costs 0.38 RTEB points against its BF16 parent, or 99.5% retention.

How the 1B Model was Built?

Those 1B scores come from a compression pipeline, not a smaller training run. The parent was nemotron-3-embed-3b, pruned and distilled across two iterative rounds.

First, the 3B parent was pruned to 2B using NVIDIA ModelOpt mcore_minitron Neural Architecture Search (NAS). The search covers hidden width, FFN size, attention heads, and depth. It then picks the best candidate from the top-10 Pareto front. A 50k in-domain calibration corpus scored those candidates.

Next, the 2B model was distilled from the fine-tuned 8B embedding teacher. Distillation combined cosine distance loss (COS) and mean squared error (MSE) loss. The data blend was multilingual and in-domain. Finally, the same procedure repeated to produce the 1.14B checkpoint.

The NVFP4 Serving Tradeoff

Compression then continues into the serving format. Quantization hit weights and activations of linear layers only, targeting the NVFP4 data type. The research team used nvidia-modelopt v0.45.0. Quantization-Aware Distillation (QAD) followed, primarily to recover accuracy on long inputs.

Calibration used 512 samples: 256 queries and 256 passages from abisee/cnn_dailymail. QAD training used 20k samples.

The rsesearch team reports NVFP4 on Blackwell delivers up to 2x higher throughput than BF16. It retains 99%+ of BF16 retrieval accuracy. The NVFP4 card also documents dynamic embedding sizes. You can slice the 2048-d vector from the start to 1024 or 512 dimensions. Re-normalize afterward.


Interactive Explainer: The Five-Stage Retrieval Path

Before touching code, watch the path run. It animates prefixing, bidirectional encoding, average pooling, L2 normalization, and dot-product scoring. Scores come from each card’s published expected output.

Deployment Matrix

As that walkthrough implies, the checkpoints do not share runtime paths.

Feature 8B-BF16 1B-BF16 1B-NVFP4
Transformers / Sentence Transformers Yes Yes No
vLLM for /v2/embed 0.25.0 0.25.0 0.25.0
Microarchitectures Ampere, Hopper, Blackwell Ampere, Hopper, Blackwell Ampere, Hopper, Lovelace, Blackwell
Test hardware A100 80GB, H100 80GB A100 80GB, H100 80GB GB200, RTX 6000 PRO, A100, H100, L40, L4
Training data 50M+ samples 8.5M+ (distillation) 20k (QAD)

Alongside the checkpoints, NVIDIA research team released an optimized NIM microservice for the 1B model. The Rust-based NIM matches or outperforms the vLLM checkpoint on GB200 and RTX PRO 6000. NVIDIA tested input sequence lengths of 256 and 1024. Separately, NVIDIA NeMo AutoModel recipes cover fine-tuning and distillation.

Using It in Code

With those paths in mind, prefixes come first. Queries take query: and documents take passage: . Embeddings are L2-normalized, so dot product equals cosine similarity.

# pip install --upgrade "transformers>=5.2.0" "sentence-transformers>=5.4.1"
import torch
from sentence_transformers import SentenceTransformer

QUERIES = ["How can someone reduce exposure to pollen during allergy season?"]
DOCUMENTS = ["People with pollen allergy can reduce exposure by staying indoors "
             "on dry, windy days, avoiding early-morning outdoor activity, and "
             "going outside after rain when pollen levels are lower."]

model = SentenceTransformer(
    "nvidia/Nemotron-3-Embed-8B-BF16",
    device="cuda",
    model_kwargs={"dtype": torch.bfloat16,
                  # use "sdpa" if FlashAttention-2 is unavailable
                  "attn_implementation": "flash_attention_2"},
    processor_kwargs={"padding_side": "left"},
)
model.max_seq_length = 32768

q = model.encode_query(QUERIES, batch_size=1, convert_to_tensor=True)
d = model.encode_document(DOCUMENTS, batch_size=1, convert_to_tensor=True)
print(model.similarity(q, d))  # card's published q[3]/d[3] score: 0.8008

encode_query and encode_document read the saved prompts. So you never add prefixes by hand. For serving, /v2/embed applies them from input_type instead:

vllm serve nvidia/Nemotron-3-Embed-1B-NVFP4 \
  --max-model-len 4096 \
  --max-num-batched-tokens 4096 \
  --max-cudagraph-capture-size 4096
import numpy as np, requests

def embed(input_type: str, texts: list[str]) -> np.ndarray:
    r = requests.post(
        "http://localhost:8000/v2/embed",
        json={"model": "nvidia/Nemotron-3-Embed-1B-NVFP4",
              "input_type": input_type,          # "query" or "document"
              "texts": texts,
              "embedding_types": ["float"],
              "truncate": "END"},
        timeout=120,
    )
    r.raise_for_status()
    return np.array(r.json()["embeddings"]["float"], dtype=np.float32)

scores = embed("query", QUERIES) @ embed("document", DOCUMENTS).T

Use Cases With Examples

  • Multilingual enterprise search: A support team indexes Hindi, Japanese, and English tickets together. Because retrieval is cross-lingual, a German query can surface a Japanese resolution note.
  • Code retrieval: Training included coir_apps, coir_cosqa, synthetic_text2sql, and SWE-bench. Natural-language-to-code lookup is therefore closer to in-distribution.
  • Agent memory: The 32,768-token limit lets an agent embed long conversation summaries without aggressive chunking.
  • Cost-tiered RAG: Serve 1B-NVFP4 for high-volume recall, and route hard queries to the 8B. Because widths differ, this needs two indexes.

Key Takeaways

  • Nemotron-3-Embed-8B-BF16 ranks #1 on RTEB at 78.46 avg NDCG@10.
  • Three open checkpoints span 8B BF16, 1B BF16, and 1B NVFP4.
  • NVFP4 retains 99%+ of BF16 accuracy at up to 2x Blackwell throughput.
  • The 1B came from ModelOpt NAS pruning plus COS+MSE distillation from the 8B.
  • All checkpoints use OpenMDW-1.1 and support 32,768-token inputs.

Check out the NVIDIA launch post on Hugging Face, Nemotron 3 Embed collection, 8B-BF16 card, 1B-BF16 card and 1B-NVFP4 card. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Sony Files Another Lawsuit Against AI Music Generator Udio
AI & Technology

Sony Files Another Lawsuit Against AI Music Generator Udio

July 21, 2026
Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today
AI & Technology

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

July 20, 2026
The First UL 3700-Compliant Plug-In Solar Microinverter Is Now Available In The US
AI & Technology

The First UL 3700-Compliant Plug-In Solar Microinverter Is Now Available In The US

July 20, 2026
Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy
AI & Technology

Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy

July 20, 2026
Next Post
Andy Burnham Becomes Labour Leader and Is Set to Be UK Prime Minister: Live Updates – The New York Times

Andy Burnham Becomes Labour Leader and Is Set to Be UK Prime Minister: Live Updates - The New York Times

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Infrared Tech Is Decades Old – Why Does Almost Every TV Remote Use It?

Infrared Tech Is Decades Old – Why Does Almost Every TV Remote Use It?

July 19, 2026
England vs. Argentina live updates: Score, highlights and analysis from the World Cup semifinal – Yahoo Sports

England vs. Argentina live updates: Score, highlights and analysis from the World Cup semifinal – Yahoo Sports

July 15, 2026
Heavenly Spices garlic powder recalled over possible microbial contamination

Heavenly Spices garlic powder recalled over possible microbial contamination

July 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!