• bitcoinBitcoin(BTC)$78,053.000.64%
  • ethereumEthereum(ETH)$2,446.400.61%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$690.970.25%
  • rippleXRP(XRP)$1.391.14%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$104.891.36%
  • tronTRON(TRX)$0.338651-0.32%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.51%
  • HyperliquidHyperliquid(HYPE)$82.763.49%
  • zcashZcash(ZEC)$833.646.40%
  • dogecoinDogecoin(DOGE)$0.0851010.51%
  • RainRain(RAIN)$0.017615-0.21%
  • USDSUSDS(USDS)$1.000.03%
  • leo-tokenLEO Token(LEO)$9.642.36%
  • moneroMonero(XMR)$463.12-0.42%
  • chainlinkChainlink(LINK)$11.390.17%
  • whitebitWhiteBIT Coin(WBT)$71.940.64%
  • cardanoCardano(ADA)$0.200757-0.10%
  • stellarStellar(XLM)$0.1786230.64%
  • bitcoin-cashBitcoin Cash(BCH)$246.080.02%
  • CantonCanton(CC)$0.1174107.18%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.02%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.81%
  • litecoinLitecoin(LTC)$48.67-0.16%
  • Global DollarGlobal Dollar(USDG)$1.000.04%
  • hedera-hashgraphHedera(HBAR)$0.075387-0.34%
  • avalanche-2Avalanche(AVAX)$7.300.79%
  • suiSui(SUI)$0.740.72%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.41%
  • uniswapUniswap(UNI)$4.635.85%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • crypto-com-chainCronos(CRO)$0.0573040.82%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • tether-goldTether Gold(XAUT)$4,456.34-0.08%
  • nearNEAR Protocol(NEAR)$1.852.13%
  • MemeCoreMemeCore(M)$1.062.83%
  • okbOKB(OKB)$112.843.70%
  • Ripple USDRipple USD(RLUSD)$1.000.06%
  • BittensorBittensor(TAO)$236.921.14%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • pax-goldPAX Gold(PAXG)$4,462.080.01%
  • aaveAave(AAVE)$123.421.44%
  • Pump.funPump.fun(PUMP)$0.0047892.39%
  • AsterAster(ASTER)$0.702.44%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0578040.64%
  • OndoOndo(ONDO)$0.351347-0.44%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Soofi Consortium Releases Soofi S 30B-A3B: An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English

July 15, 2026
in AI & Technology
Reading Time: 24 mins read
A A
Soofi Consortium Releases Soofi S 30B-A3B: An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English
ShareShareShareShareShare

A German research consortium has published the pretraining report for Soofi S 30B-A3B. It is an open base model for German and English. Training ran end to end on Deutsche Telekom’s Industrial AI Cloud in Munich. Preview weights are on Hugging Face. It is worth noting that among some of the fully open base models tested, Soofi S records the highest English and German aggregate scores.

What is Soofi S 30B-A3B?

Soofi S is a Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model. It totals ~31.6B parameters and activates ~3.2B per token. As a base model, it has no instruction tuning, alignment, or safety tuning. The KI Bundesverband coordinates the consortium, funded by the German Federal Ministry for Economic Affairs and Energy. Participants include Fraunhofer IAIS, DFKI, TU Darmstadt, ellamind, and Merantix Momentum.

How the architecture works?

The efficiency claim starts with the layer stack. The network holds 52 layers. That is 23 Mamba-2 sequence-mixing layers, 23 granular MoE layers, and 6 Grouped-Query Attention (GQA) layers. Only those 6 GQA layers maintain a KV cache. Each MoE layer holds 128 routed experts, activates 6 per token, and adds 2 shared experts. Other details: model dimension 2688, squared ReLU, RMSNorm, and no positional embeddings.

Soofi S adopts the Nemotron 3 Nano reference design without modification. The research team gives three reasons for that choice. Those are deployability on stacks such as vLLM, serving efficiency, and scientific control. Because the backbone is fixed, Nemotron 3 Nano becomes an architecture-identical baseline. The data recipe is the only moving part.

The training recipe: ~26.68T consumed tokens in three phases

That recipe follows a Warmup–Stable–Decay (WSD) schedule with a minus_sqrt decay segment. Phase 1 consumed ~20T tokens on a diverse, quality-tiered mixture at a 1e-3 plateau. Phase 2 consumed ~6.58T tokens of high-quality annealing data. It decays 1e-3 to 1e-5, then continues at a constant 1e-5. Phase 3 consumed ~0.10T tokens at a 1,048,576-token sequence length. It extends the usable context window up to 1M tokens.

German is the deliberate variable. It rises from 7.2% of Phase 1 effective tokens to 15.32% in Phase 2. The reference Nemotron 3 Nano mixture allocates about 5% to all non-English languages combined. German sources include HPLT v3 and v4, German Commons, German FinePDFs, and FineWiki. Genios adds 193M articles from 916 newspaper and trade-press archives, commercially licensed.

YOU MAY ALSO LIKE

Why 1080p Movies Can Look Better Than 4K Movies On Your TV

Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI

Infrastructure follows the same sovereignty logic. The run used up to 512 NVIDIA B200 GPUs, from 24 March to 13 May 2026. It consumed ~253,000 B200 GPU-hours.

Performance

Those choices show up in the evaluation. Soofi S ran against 16 other open base models. All used the same lm-evaluation-harness pipeline, prompts, and few-shot settings.

Benchmark (%) Soofi S 30B-A3B Olmo 3 32B Apertus 70B EuroLLM 22B Alia 40B
English aggregate 70.1 67.3 62.4 61.2 59.0
German aggregate 79.1 69.2 72.8 70.6 68.4
Held-out (EN / DE) 41.4 / 41.8 33.1 / 36.2 27.6 / 33.5 30.8 / 33.9 28.0 / 29.4
HumanEval (pass@1) 73.8 63.0 30.2 39.3 23.8
MBPP-DE (pass@1) 84.2 70.8 50.9 59.4 45.6
LBPP (pass@1) 31.0 32.1 6.4 10.7 8.6
GSM8K 86.1 80.7 65.4 25.1 65.4
Minerva MATH-DE 56.0 48.5 29.0 28.4 12.9
INCLUDE-DE 61.2 48.2 50.4 51.1 43.9
GPQA-Diamond 43.4 33.3 27.3 30.3 29.8
GLP-DE 88.8 73.0 81.2 78.2 65.4

Against its architecture-identical reference, Soofi S gains 1.8 points on the English aggregate. German gains 4.2, and held-out English 6.7. That isolates the data recipe from the backbone.

The picture changes against larger open-weight models. Qwen3.5 35B-A3B holds the highest English, German, and held-out means. Soofi S scores 70.1 English against 70.3 for Gemma 3 27B and Ministral 3 14B. On German it leads both, 79.1 to 78.4 and 78.3.

Running the base model

Reproducing any of this starts with the weights. The base repo is a gated preview, and it ships custom modeling code.

# pip install -U transformers accelerate torch
# hf auth login   # base repo is gated: accept the terms on the model page first
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Soofi-Project/Soofi-S-Base"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, dtype="auto", device_map="auto"
)

# Base model: plain text completion. No chat template, no system prompt.
prompt = "AI sovereignty is the idea that"
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

The same repo serves through vLLM:

vllm serve "Soofi-Project/Soofi-S-Base"

Where it fits?

Together, the numbers suggest three deployment shapes. First, German document work: GLP-DE 88.8 and INCLUDE-DE 61.2 suit an insurer fine-tuning on policy PDFs. Second, bilingual code assistance: MBPP-DE 84.2 suits teams prompting in German against Python tasks. Third, high-concurrency long-context serving: a support-ticket RAG system at batch 32 and 40K context matches the measured regime. For that case, test retrieval against the RULER and NaturalQuestions gaps.

Key Takeaways

  • Soofi S activates 3.2B of 31.6B parameters; only 6 of 52 layers hold a KV cache.
  • It leads fully open base models: 70.1 English aggregate, 79.1 German aggregate.
  • German hits 15.32% of the Phase 2 mixture, versus ~5% multilingual in Nemotron.
  • Decode measures 8–9× dense 14–24B models at 40K context, flat from 4K to 256K.
  • Open gaps: RULER extraction at long inputs, factual recall, gated preview weights, unfinalized license.

Check out the Pretraining report, Project page and Hugging Face . Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why 1080p Movies Can Look Better Than 4K Movies On Your TV
AI & Technology

Why 1080p Movies Can Look Better Than 4K Movies On Your TV

August 29, 2026
Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI
AI & Technology

Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI

August 29, 2026
Don’t Fall For This Digital TV Antenna Myth
AI & Technology

Don’t Fall For This Digital TV Antenna Myth

August 29, 2026
Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling
AI & Technology

Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling

August 29, 2026
Next Post
TOP 5 STOCKS TO WATCH AS IRAN WAR CONTINUES…

TOP 5 STOCKS TO WATCH AS IRAN WAR CONTINUES...

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Sen. Mitch McConnell says he’s out of rehab and recovering at home

Sen. Mitch McConnell says he’s out of rehab and recovering at home

August 28, 2026
Stock up, stock down after Steelers’ preseason loss to Jets: Mason Rudolph looks good by comparison – TribLIVE.com

Stock up, stock down after Steelers’ preseason loss to Jets: Mason Rudolph looks good by comparison – TribLIVE.com

August 22, 2026
This changed everything for me…

This changed everything for me…

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!