• bitcoinBitcoin(BTC)$64,483.000.70%
  • ethereumEthereum(ETH)$1,913.030.70%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$603.61-0.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.000.20%
  • solanaSolana(SOL)$76.941.80%
  • tronTRON(TRX)$0.3325820.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-0.10%
  • HyperliquidHyperliquid(HYPE)$58.57-0.90%
  • dogecoinDogecoin(DOGE)$0.0699720.10%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013095-0.40%
  • leo-tokenLEO Token(LEO)$9.440.30%
  • zcashZcash(ZEC)$504.89-0.60%
  • moneroMonero(XMR)$410.81-0.70%
  • chainlinkChainlink(LINK)$9.490.50%
  • whitebitWhiteBIT Coin(WBT)$55.630.60%
  • cardanoCardano(ADA)$0.1737790.50%
  • stellarStellar(XLM)$0.154519-1.60%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$202.71-0.30%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-0.30%
  • CantonCanton(CC)$0.0909021.60%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • litecoinLitecoin(LTC)$44.39-0.20%
  • Circle USYCCircle USYC(USYC)$1.130.10%
  • hedera-hashgraphHedera(HBAR)$0.0672832.30%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • avalanche-2Avalanche(AVAX)$6.330.30%
  • suiSui(SUI)$0.65-3.10%
  • tether-goldTether Gold(XAUT)$4,327.16-1.50%
  • shiba-inuShiba Inu(SHIB)$0.000004-1.20%
  • crypto-com-chainCronos(CRO)$0.046112-1.70%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.10%
  • okbOKB(OKB)$101.961.70%
  • nearNEAR Protocol(NEAR)$1.58-2.60%
  • uniswapUniswap(UNI)$3.290.80%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0606200.00%
  • pax-goldPAX Gold(PAXG)$4,335.55-1.60%
  • BittensorBittensor(TAO)$190.65-2.40%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • AsterAster(ASTER)$0.600.30%
  • OndoOndo(ONDO)$0.323893-2.00%
  • HTX DAOHTX DAO(HTX)$0.000002-3.60%
  • MemeCoreMemeCore(M)$1.15-0.20%
  • usddUSDD(USDD)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands

August 18, 2026
in AI & Technology
Reading Time: 15 mins read
A A
NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands
ShareShareShareShareShare

NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an open-source project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands. There is no intermediate ONNX export step. The build produces a versioned .bundle artifact that runs through native C++ task APIs, so inference can execute in a C++ service, embedded application, or robotics stack without PyTorch in the runtime path. The project is Apache-2.0 licensed and ships as a collection of family-owned reference implementations rather than a single generic converter. NVIDIA also states that the entire project — model implementations, performance tuning, tests, integrations, and docs — was built using OpenAI Codex agents under human direction and review.

Is it deployable?

Yes, for evaluation and native integration work, with real conditions. The code is open and installable. Release wheels currently target Linux aarch64 only, with Python 3.10 or 3.12, glibc 2.39 or newer, and TensorRT 11.1.0.106. x86_64 wheels are not published; x86_64 users must take the Docker source-build path.

YOU MAY ALSO LIKE

ICE Agents Can’t Wear Meta Glasses While They Work, Official Memo Warns

WhiteFiber Proposes $250M Convertible Senior Notes to Fund Data Center Expansion – Unite.AI

  • Company level: Best fit today is teams that already own their inference stack: NVIDIA-shop startups, robotics and device companies, and platform or inference teams inside mid-size and large enterprises. Small teams shipping a Python service get less from it. Regulated enterprises should wait for a tagged release before standardizing on it.
  • Industries: Robotics and autonomous machines, industrial inspection and manufacturing, automotive in-vehicle compute, medical devices, defense and aerospace edge systems, and media processing — anywhere inference has to live inside a C++ binary rather than a Python server.
  • Applications: On-device text generation, speech recognition and synthesis, OCR and document parsing, embeddings and reranking for a retrieval service written in C++, diffusion image and video generation, segmentation, and time-series forecasting.

The two commands

The quick start builds and runs Qwen3-0.6B:

trtmc build Qwen/Qwen3-0.6B --precision bf16 --max-cache-length 16384 --output qwen3-0.6b.bundle
trtmc run ./qwen3-0.6b.bundle --prompt "What is the capital of France? Answer in one word." --chat-template --no-thinking

The same .bundle loads from C++ with trtmc::load("./qwen3-0.6b.bundle").

The bundle is the actual design decision

TRTMC splits build and runtime at a versioned artifact. Python owns checkpoint resolution and TensorRT engine construction. Native profiles then execute inference in C++ without PyTorch. A small number of hybrid profiles invoke a helper Python executable, and their manifests declare that dependency explicitly.

Applications call task APIs — generate(), transcribe(), generate_image(), embed(), solve() — instead of maintaining conversion stages and per-model application glue. trtmc inspect exposes bundle kind, model family, precision, runtime identity, and engines, which makes the artifact auditable rather than opaque.

NVIDIA frames the conventional route as PyTorch → ONNX or TorchScript → TensorRT → model-specific C++ integration, and names the failure modes it removes: export gaps, repeated per-model integration, and validation spread across several conversion artifacts.

Key Takeaways

  • Two commands take a supported Hugging Face checkpoint to native C++ TensorRT inference, with no ONNX step.
  • A versioned .bundle is the handoff between the Python build and a PyTorch-free C++ runtime.
  • The July 29, 2026 GB300 snapshot covers 105 profiles across 76 families; 102 beat their declared reference by more than 5%.
  • Wheels are Linux aarch64 only today; x86_64 requires the Docker source build.

Check out the GitHub Repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

ICE Agents Can’t Wear Meta Glasses While They Work, Official Memo Warns
AI & Technology

ICE Agents Can’t Wear Meta Glasses While They Work, Official Memo Warns

August 18, 2026
WhiteFiber Proposes 0M Convertible Senior Notes to Fund Data Center Expansion – Unite.AI
AI & Technology

WhiteFiber Proposes $250M Convertible Senior Notes to Fund Data Center Expansion – Unite.AI

August 18, 2026
Spotify’s Running Mode Is Now Available On Android
AI & Technology

Spotify’s Running Mode Is Now Available On Android

August 18, 2026
Saulius Lazaravičius, VP of Product at Hostinger – Interview Series – Unite.AI
AI & Technology

Saulius Lazaravičius, VP of Product at Hostinger – Interview Series – Unite.AI

August 18, 2026
Next Post
You Guys Make Good Money But Can’t Say No

You Guys Make Good Money But Can't Say No

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE NOW: SUPERMICRO & COREWEAVE EARNINGS REPORT

LIVE NOW: SUPERMICRO & COREWEAVE EARNINGS REPORT

August 13, 2026
MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption

MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption

August 17, 2026
You Can Now Watch Classic Movies Like The Martian, E.T. And Zodiac On Apple TV

You Can Now Watch Classic Movies Like The Martian, E.T. And Zodiac On Apple TV

August 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!