• bitcoinBitcoin(BTC)$81,098.00-0.28%
  • ethereumEthereum(ETH)$2,620.430.19%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.12-0.11%
  • rippleXRP(XRP)$1.40-0.95%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$110.27-2.93%
  • tronTRON(TRX)$0.3405110.66%
  • zcashZcash(ZEC)$1,468.99-5.92%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.50%
  • HyperliquidHyperliquid(HYPE)$91.51-2.43%
  • dogecoinDogecoin(DOGE)$0.087318-1.26%
  • moneroMonero(XMR)$542.02-6.38%
  • whitebitWhiteBIT Coin(WBT)$82.67-0.77%
  • RainRain(RAIN)$0.0137272.51%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.33-0.58%
  • cardanoCardano(ADA)$0.228278-0.90%
  • leo-tokenLEO Token(LEO)$8.910.05%
  • stellarStellar(XLM)$0.1954150.67%
  • uniswapUniswap(UNI)$8.83-3.91%
  • bitcoin-cashBitcoin Cash(BCH)$251.81-2.30%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.58-6.76%
  • daiDai(DAI)$1.00-0.01%
  • litecoinLitecoin(LTC)$58.00-1.40%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6712.23%
  • CantonCanton(CC)$0.108219-3.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.42%
  • hedera-hashgraphHedera(HBAR)$0.0818813.12%
  • suiSui(SUI)$0.864.67%
  • MemeCoreMemeCore(M)$1.5317.43%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000060.33%
  • BittensorBittensor(TAO)$263.784.00%
  • crypto-com-chainCronos(CRO)$0.059877-0.83%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • tether-goldTether Gold(XAUT)$4,370.77-0.05%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$117.12-0.20%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.42%
  • aaveAave(AAVE)$141.53-0.83%
  • EthenaEthena(ENA)$0.21339822.50%
  • mantleMantle(MNT)$0.62-0.34%
  • AsterAster(ASTER)$0.76-1.71%
  • OndoOndo(ONDO)$0.4157101.03%
  • Pump.funPump.fun(PUMP)$0.004139-2.44%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI releases C-RADIOv4 vision backbone unifying SigLIP2, DINOv3, SAM3 for classification, dense prediction, segmentation workloads at scale

February 7, 2026
in AI & Technology
Reading Time: 8 mins read
A A
NVIDIA AI releases C-RADIOv4 vision backbone unifying SigLIP2, DINOv3, SAM3 for classification, dense prediction, segmentation workloads at scale
ShareShareShareShareShare

How do you combine SigLIP2, DINOv3, and SAM3 into a single vision backbone without sacrificing dense or segmentation performance? NVIDIA’s C-RADIOv4 is a new agglomerative vision backbone that distills three strong teacher models, SigLIP2-g-384, DINOv3-7B, and SAM3, into a single student encoder. It extends the AM-RADIO and RADIOv2.5 line, keeping similar computational cost while improving dense prediction quality, resolution robustness, and drop-in compatibility with SAM3.

The key idea is simple. Instead of choosing between a vision language model, a self supervised dense model, and a segmentation model, C-RADIOv4 tries to approximate all three at once with one backbone.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

https://www.arxiv.org/pdf/2601.17237

Agglomerative distillation in RADIO

The RADIO family uses agglomerative distillation. A single ViT style student is trained to match both dense feature maps and summary tokens from several heterogeneous teachers.

Earlier RADIO models combined DFN CLIP, DINOv2, and SAM. They already supported multi resolution training but showed ‘mode switching’, where the representation changed qualitatively as input resolution changed. Later work such as PHI-S, RADIOv2.5, and FeatSharp added better multi resolution distillation and regularization, but the teacher set was still limited.

C-RADIOv4 upgrades the teachers:

  • SigLIP2-g-384 for stronger image text alignment
  • DINOv3-7B for high quality self supervised dense features
  • SAM3 for segmentation oriented features and compatibility with the SAM3 decoder

The student is trained so that its dense features match DINOv3 and SAM3, while its summary tokens match SigLIP2 and DINOv3. This gives one encoder that can support classification, retrieval, dense prediction, and segmentation.

Stochastic multi resolution training

C-RADIOv4 uses stochastic multi resolution training rather than a small fixed set of resolutions.

Training samples input sizes from two partitions:

  • Low resolution: {128, 192, 224, 256, 384, 432}
  • High resolution: {512, 768, 1024, 1152}

SigLIP2 operates natively at 384 pixels. Its features are upsampled by a factor of 3 using FeatSharp to align with 1152 pixel SAM3 features. SAM3 is trained with mosaic augmentation at 1152 × 1152.

This design smooths the performance curve over resolution and improves low resolution behavior. For example, on ADE20k linear probing, C-RADIOv4-H reaches around:

  • 55.20 mIoU at 512 px
  • 57.02 mIoU at 1024 px
  • 57.72 mIoU at 1536 px

The scaling trend is close to DINOv3-7B while using roughly an order of magnitude fewer parameters.

Removing teacher noise with shift equivariant losses and MESA

Distilling from large vision models tends to copy their artifacts, not just their useful structure. SigLIP2 has border noise patterns, and ViTDet style models can show window boundary artifacts. Direct feature regression can force the student to reproduce those patterns.

C-RADIOv4 introduces two shift equivariant mechanisms to suppress such noise:

  1. Shift equivariant dense loss: Each teacher and the student see independently shifted crops of an image. Before computing the squared error, features are aligned via a shift mapping and the loss only uses overlapping spatial positions. Because the student never sees the same absolute positions as the teacher, it cannot simply memorize position fixed noise and is forced to track input dependent structure instead.
  2. Shift equivariant MESA: C-RADIOv4 also uses MESA style regularization between the online network and an EMA copy. Here again, the student and its EMA see different crops, features are aligned by a shift, and the loss is applied after layer normalization. This encourages smooth loss landscapes and robustness, while being invariant to absolute position.

In addition, training uses DAMP, which injects multiplicative noise into weights. This further improves robustness to corruptions and small distribution shifts.

Balancing teachers with an angular dispersion aware summary loss

The summary loss in previous RADIO models used cosine distance between student and teacher embeddings. Cosine distance removes magnitude but not directional dispersion on the sphere. Some teachers, such as SigLIP2, produce embeddings concentrated in a narrow cone, while DINOv3 variants produce more spread out embeddings.

If raw cosine distance is used, teachers with wider angular dispersion contribute larger losses and dominate optimization. In practice, DINOv3 tended to overshadow SigLIP2 in the summary term.

C-RADIOv4 replaces this with an angle normalized loss. The squared angle between student and teacher embeddings is divided by the teacher’s angular dispersion. Measured dispersions show SigLIP2-g-384 around 0.694, while DINOv3-H+ and DINOv3-7B are around 2.12 and 2.19. Normalizing by these values equalizes their influence and preserves both vision language and dense semantics.

Performance: classification, dense prediction, and Probe3d

On ImageNet-1k zero shot classification, C-RADIOv4-H reaches about 83.09 % top-1 accuracy. It matches or improves on RADIOv2.5-H and C-RADIOv3-H across resolutions, with the best performance near 1024 px.

On k-NN classification, C-RADIOv4-H improves over RADIOv2.5 and C-RADIOv3, and matches or surpasses DINOv3 starting around 256 px. DINOv3 peaks near 192–256 px and then degrades, while C-RADIOv4 keeps stable or improving performance at higher resolutions.

Dense and 3D aware metrics show the intended tradeoff. On ADE20k, PASCAL VOC, NAVI, and SPair, C-RADIOv4-H and the SO400M variant outperform earlier RADIO models and are competitive with DINOv3-7B on dense benchmarks. For C-RADIOv4-H, typical scores are:

  • ADE20k: 55.20 mIoU
  • VOC: 87.24 mIoU
  • NAVI: 63.44
  • SPair: 60.57
https://www.arxiv.org/pdf/2601.17237

On Probe3d, which includes Depth Normals, Surface Normals, NAVI, and SPair, C-RADIOv4-H achieves the best NAVI and SPair scores in the RADIO family. Depth and Surface metrics are close to those of C-RADIOv3-H, with small differences in either direction, rather than a uniform improvement.

Integration with SAM3 and ViTDet-mode deployment

C-RADIOv4 is designed to be a drop in replacement for the Perception Encoder backbone in SAM3. The SAM3 decoder and memory components remain unchanged. A reference implementation is provided in a SAM3 fork. Qualitative examples show that segmentation behavior is preserved for both text prompts such as “shoe”, “helmet”, “bike”, “spectator” and box prompts, and in some reported cases C-RADIOv4 based SAM3 resolves failure cases from the original encoder.

For deployment, C-RADIOv4 exposes a ViTDet-mode configuration. Most transformer blocks use windowed attention, while a few use global attention. Supported window sizes range from 6 × 6 to 32 × 32 tokens, subject to divisibility with patch size and image resolution. On an A100, the SO400M model with window size at most 12 is faster than the SAM3 ViT-L+ encoder across a wide range of input sizes, and the Huge model with window size 8 is close in latency.

This makes C-RADIOv4 a practical backbone for high resolution dense tasks where full global attention at all layers is too expensive.

Key Takeaways

  1. Single unified backbone: C-RADIOv4 distills SigLIP2-g-384, DINOv3-7B, and SAM3 into one ViT-style encoder that supports classification, retrieval, dense prediction, and segmentation.
  2. Any-resolution behavior: Stochastic multi resolution training over {128…1152} px, and FeatSharp upsampling for SigLIP2, stabilizes performance across resolutions and tracks DINOv3-7B scaling with far fewer parameters.
  3. Noise suppression via shift equivariance: Shift equivariant dense loss and shift equivariant MESA prevent the student from copying teacher border and window artifacts, focusing learning on input dependent semantics.
  4. Balanced multi-teacher distillation: An angular dispersion normalized summary loss equalizes the contribution of SigLIP2 and DINOv3, preserving both text alignment and dense representation quality.
  5. SAM3 and ViTDet-ready deployment: C-RADIOv4 can directly replace the SAM3 Perception Encoder, offers ViTDet-mode windowed attention for faster high resolution inference, and is released under the NVIDIA Open Model License.

Check out the Paper, Repo, Model-1 and Model-2. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post NVIDIA AI releases C-RADIOv4 vision backbone unifying SigLIP2, DINOv3, SAM3 for classification, dense prediction, segmentation workloads at scale appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
AI & Technology

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI
AI & Technology

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
Next Post
Crypto startup Erebor becomes first bank approved in Trump’s 2nd term

Crypto startup Erebor becomes first bank approved in Trump's 2nd term

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Larry Ellison cancels plan to sell Oracle stock worth up to .5B

Larry Ellison cancels plan to sell Oracle stock worth up to $7.5B

September 13, 2026
Mark Sanchez to plead guilty

Mark Sanchez to plead guilty

September 18, 2026
Building a portfolio from scratch? David Wagner goes rapid-fire on the top names that make the cut

Building a portfolio from scratch? David Wagner goes rapid-fire on the top names that make the cut

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!