• bitcoinBitcoin(BTC)$86,671.001.27%
  • ethereumEthereum(ETH)$2,769.770.94%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$792.000.01%
  • rippleXRP(XRP)$1.605.76%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$119.061.24%
  • tronTRON(TRX)$0.342862-1.09%
  • zcashZcash(ZEC)$1,608.299.39%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.80%
  • HyperliquidHyperliquid(HYPE)$97.354.28%
  • dogecoinDogecoin(DOGE)$0.1019363.32%
  • moneroMonero(XMR)$566.92-2.16%
  • whitebitWhiteBIT Coin(WBT)$87.071.04%
  • chainlinkChainlink(LINK)$13.101.23%
  • cardanoCardano(ADA)$0.2559434.70%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013134-5.30%
  • leo-tokenLEO Token(LEO)$8.980.17%
  • stellarStellar(XLM)$0.2178212.94%
  • bitcoin-cashBitcoin Cash(BCH)$344.0130.22%
  • uniswapUniswap(UNI)$10.7519.80%
  • nearNEAR Protocol(NEAR)$4.380.77%
  • avalanche-2Avalanche(AVAX)$11.120.21%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$62.932.74%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.115481-1.43%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0994468.21%
  • suiSui(SUI)$1.03-0.71%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.470.98%
  • shiba-inuShiba Inu(SHIB)$0.0000063.43%
  • BittensorBittensor(TAO)$313.751.74%
  • crypto-com-chainCronos(CRO)$0.0672962.40%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.31-7.57%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,348.67-0.07%
  • okbOKB(OKB)$124.781.10%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9017.44%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • aaveAave(AAVE)$149.624.19%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.28%
  • mantleMantle(MNT)$0.686.26%
  • EthenaEthena(ENA)$0.2194644.15%
  • OndoOndo(ONDO)$0.439970-0.05%
  • Pump.funPump.fun(PUMP)$0.0044590.55%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Nous Research Team Releases Hermes 4: A Family of Open-Weight AI Models with Hybrid Reasoning

August 28, 2025
in AI & Technology
Reading Time: 11 mins read
A A
Nous Research Team Releases Hermes 4: A Family of Open-Weight AI Models with Hybrid Reasoning
ShareShareShareShareShare

Nous Research has released Hermes 4, a family of open-weight models (14B, 70B, and 405B parameter sizes based on Llama 3.1 checkpoints) that achieves frontier-level performance through pure post-training techniques. Hermes 4 introduces hybrid reasoning – models can toggle between standard responses and explicit reasoning using <think>...</think> tags when complex problems require deeper deliberation.

What makes Hermes 4 particularly significant is its achievement of state-of-the-art performance among open-weight models while maintaining complete transparency and neutral alignment philosophy, demonstrating that sophisticated reasoning capabilities can be developed entirely through open-source methodologies.

YOU MAY ALSO LIKE

The Pros And Cons Of Using A Password Manager Over An Authenticator App

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)

DataForge: Graph-Based Synthetic Data Generation

DataForge is the main component behind Hermes 4’s core structure. But what is DataForge? DataForge is a revolutionary graph-based synthetic data generation system that transforms how training data is created. Unlike traditional curation approaches, DataForge operates through a directed acyclic graph (DAG) where each node implements a PDDL (Planning Domain Definition Language) action interface.

Each node specifies preconditions, postconditions, and transformations, facilitating the automatic creation of complex data pipelines. By using pre-training seed data from DCLM and FineWeb, the system can transform a Wikipedia article into a rap song, and then generate instruction-answer pairs based on that transformation.

This approach generates approximately 5 million samples totaling 19 billion tokens, with reasoning samples being intentionally token-heavy – averaging five times more tokens than non-reasoning counterparts to accommodate thinking traces up to 16,000 tokens long.

https://arxiv.org/pdf/2508.18255

Rejection Sampling at Unprecedented Scale

Hermes 4 uses Atropos, Nous Research’s open-source reinforcement learning environment, to implement rejection sampling across approximately 1,000 different task-specific verifiers. This massive verification infrastructure filters for high-quality reasoning trajectories across diverse domains.

Key verification environments include Answer Format Training (rewarding correct formatting across 150+ output formats), Instruction Following (using RLVR-IFEval tasks with complex constraints), Schema Adherence (for JSON generation using Pydantic models), and Tool Use training for agentic behavior.

The rejection sampling process creates a large corpus of verified reasoning trajectories, with multiple unique solution paths to the same verified result. This approach ensures the model learns robust reasoning patterns rather than memorizing specific solution templates.

Length Control: Solving Overlong Generation

One of Hermes 4’s most innovative contributions addresses the overlong reasoning problem – where reasoning models generate excessively long chains of thought without termination. The research team discovered their 14B model reached maximum context length 60% of the time on LiveCodeBench when in reasoning mode.

Their super effective solution involves a second supervised fine-tuning stage teaching models to stop reasoning at exactly 30,000 tokens:

  1. Generate reasoning traces from the current policy
  2. Insert </think> tokens at exactly 30,000 tokens
  3. Train only on the termination decision, not the reasoning chain
  4. Apply gradient updates solely to </think> and <eos> tokens

This approach achieves remarkable results: 78.4% reduction in overlong generation on AIME’24, 65.3% on AIME’25, and 79.8% on LiveCodeBench, with only 4.7% to 12.7% relative accuracy cost. By focusing learning signals entirely on the termination decision, the method avoids model collapse risks while teaching effective “counting behavior.”

https://hermes4.nousresearch.com/
https://hermes4.nousresearch.com/

Benchmark Performance and Neutral Alignment

Hermes 4 demonstrates state-of-the-art performance among open-weight models. The 405B model achieves 96.3% on MATH-500 (reasoning mode), 81.9% on AIME’24, 78.1% on AIME’25, 70.5% on GPQA Diamond, and 61.3% on LiveCodeBench.

Particularly notable is its performance on RefusalBench, achieving 57.1% in reasoning mode – the highest score among evaluated models, significantly outperforming GPT-4o (17.67%) and Claude Sonnet 4 (17%). This demonstrates the model’s willingness to engage with controversial topics while maintaining appropriate boundaries, reflecting Nous Research’s neutral alignment philosophy.

https://arxiv.org/pdf/2508.18255

Technical Architecture and Training

Hermes 4 training leverages a modified TorchTitan across 192 NVIDIA B200 GPUs. The system handles highly heterogeneous sample length distribution through efficient packing (achieving >99.9% batch efficiency), flex attention, and sophisticated loss masking where only assistant-role tokens contribute to cross-entropy loss.

Training follows a cosine learning rate schedule with 300 warmup steps and 9,000 total steps at 16,384 token context length with global batch size of 384 samples, combining Data Parallelism, Tensor Parallelism, and Fully Sharded Data Parallelism.

Summary

Hermes 4 marks a significant advancement in open-source AI development, proving that frontier-level reasoning capabilities can be achieved through transparent, reproducible methodologies without relying on proprietary training data or closed development processes. By combining innovative graph-based synthetic data generation, massive-scale rejection sampling, and elegant length control mechanisms, Nous Research has created models that not only match the performance of leading proprietary systems but also maintain the neutral alignment and steerability that make them genuinely useful tools rather than restrictive assistants


Check out the Paper, Technical details, Model on Hugging Face and Chat. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)
AI & Technology

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)

September 22, 2026
Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor
AI & Technology

Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor

September 22, 2026
Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5
AI & Technology

Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

September 22, 2026
Next Post
Sharapova introduced by Serena Williams at hall of fame

Sharapova introduced by Serena Williams at hall of fame

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
J.P. Morgan’s Secret Trading Plan For The FOMC Decision Tomorrow!

J.P. Morgan’s Secret Trading Plan For The FOMC Decision Tomorrow!

September 16, 2026
Vanderbilt recovers fumble with 1 second left to stun NC State – ESPN

Vanderbilt recovers fumble with 1 second left to stun NC State – ESPN

September 19, 2026
Taking Stock After Six Months of Iran War; How Sports Betting Is Influencing the Midterms | Aug. 28

Taking Stock After Six Months of Iran War; How Sports Betting Is Influencing the Midterms | Aug. 28

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!