• bitcoinBitcoin(BTC)$79,699.00-0.19%
  • ethereumEthereum(ETH)$2,494.940.46%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$748.99-2.82%
  • rippleXRP(XRP)$1.41-0.30%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$105.221.66%
  • tronTRON(TRX)$0.3353990.36%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,223.6920.21%
  • HyperliquidHyperliquid(HYPE)$86.811.52%
  • dogecoinDogecoin(DOGE)$0.089861-0.81%
  • RainRain(RAIN)$0.016722-1.85%
  • moneroMonero(XMR)$538.42-2.43%
  • USDSUSDS(USDS)$1.000.03%
  • chainlinkChainlink(LINK)$12.856.77%
  • whitebitWhiteBIT Coin(WBT)$73.50-0.01%
  • leo-tokenLEO Token(LEO)$9.370.91%
  • cardanoCardano(ADA)$0.219669-0.06%
  • stellarStellar(XLM)$0.1851390.16%
  • bitcoin-cashBitcoin Cash(BCH)$257.18-0.83%
  • daiDai(DAI)$1.00-0.01%
  • uniswapUniswap(UNI)$7.170.89%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • CantonCanton(CC)$0.1091530.14%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$54.690.10%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-0.72%
  • hedera-hashgraphHedera(HBAR)$0.0808070.46%
  • avalanche-2Avalanche(AVAX)$7.751.72%
  • suiSui(SUI)$0.800.22%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.29%
  • nearNEAR Protocol(NEAR)$2.4311.14%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0572741.01%
  • tether-goldTether Gold(XAUT)$4,420.00-0.16%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.13-0.14%
  • BittensorBittensor(TAO)$262.3011.14%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$112.82-0.36%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.25%
  • AsterAster(ASTER)$0.78-0.10%
  • aaveAave(AAVE)$134.17-0.12%
  • mantleMantle(MNT)$0.602.74%
  • pax-goldPAX Gold(PAXG)$4,424.17-0.19%
  • OndoOndo(ONDO)$0.3784131.76%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056535-1.39%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

September 6, 2026
in AI & Technology
Reading Time: 19 mins read
A A
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
ShareShareShareShareShare

Most visual document retrievers in production today are hand-me-downs. ColPali and the models that followed it take a generative vision-language model and repurpose it as an encoder. The result still carries a separately pretrained vision tower and a causal decoder that never generates a token. That is parameter and compute overhead for a task that only needs representations.

H Company has released NeoMME, a family of 260M and 800M bidirectional encoders that drops both components. One Transformer processes multilingual text tokens and raw 32×32 RGB image patches through the same layers, trained from random initialization. The retrieval fine-tune, NeoMME-Retriever, reaches 0.523 nDCG@10 on ViDoRe v3 at 260M parameters.

Is it deployable? Yes. Every checkpoint ships under Apache 2.0 with day-zero support in Hugging Face Transformers. The 260M model indexes 51.3 pages per second on a single NVIDIA L40S and encodes a query in 78.3 ms on a CPU-only host.

One tower, two modalities

Text enters through an ALBERT-style factorized embedding: a 256-dimensional lookup projected to model width. Images are split into non-overlapping 32×32 patches and projected by a 2-layer MLP trained from scratch. No patch-merging module, no SigLIP2 tower.

Both models support a 16,384-token context, enough for two standard 3,840×2,160 4K UHD images after patching. Most layers use symmetric sliding-window attention; every sixth layer and the final layer attend globally. The stack uses grouped-query attention, query-key normalization, gated attention, 2D rotary position embeddings, and squared-ReLU MLPs. Exact parameter counts are 262,937,906 and 793,715,032.

The tokenizer is a whitespace-unconstrained BPE with a 131,072-entry vocabulary, trained from scratch. Across 14 target languages in FLORES-200 devtest, it emits 44.4% fewer tokens than ModernBERT.

Trained as a masked diffusion denoiser

Pretraining is discrete masked diffusion over text, optionally conditioned on visible image patches. Text-only segments draw a corruption rate uniformly from 0 to 1. Multimodal segments draw from 0.30 to 1, which removes the language-only shortcut and forces the model to read the page.

A cross-modal ablation probe confirms this works. At 90% masking, visible page patches raise masked-token accuracy by 38.4 points for the 260M model and 40.5 points for the 800M model. Each run processes about 524 billion packed input tokens, roughly 290 billion of them text-only, on 16 and 32 H100 accelerators respectively.

Retrieval results

NeoMME-Retriever adds two jointly trained heads on the shared backbone: a mean-pooled dense head with Matryoshka widths, and a late-interaction head projecting every token and patch to 128 dimensions. One forward pass returns both.

On ViDoRe v3, the 260M model scores 0.523 nDCG@10 and the 800M model 0.556. The 260M result sits within 0.002 of ColQwen2.5-v0.2 at 3.75B parameters, and 26.1 points above the best other sub-300M model. The 800M model lands 0.9 points behind the similarly sized Vultron Retriever Flash. On ViDoRe v1 and v2 the models reach 0.860/0.522 and 0.874/0.559 nDCG@5.

Text retrieval is weaker. On BEIR-15, late interaction reaches 0.4881 and 0.5126, against 0.5722 for LateOn at 149M parameters. The authors attribute this partly to supervision scale: NeoMME saw roughly 430K text query examples, against roughly 660M contrastive examples for mLateOn.

Storage and throughput

Late-interaction indexes are expensive. A 2048×2048 page yields 4,162 vectors, about 1.5 MB per ViDoRe v3 document in float32. Two methods bring that down. Hierarchical token pooling at factor 10 with int8 queries and documents gives 39.0 kB per page, a 39.4× reduction retaining 99.16% of baseline nDCG@10. Pool factor 8 with int8 queries and binary documents gives 6.0 kB, a 255.5× reduction retaining 95.19%.

Indexing is fast for the vector count. At a matched 2048×2048 input on one L40S, NeoMME-260M encodes 51.3 pages per second against ColModernVBERT’s 26.0, a 1.97× gap.

Interactive explainer

YOU MAY ALSO LIKE

How To Send High-Quality Images And Videos From Android To iPhone

What Is Vibe Coding And Why Does It Get So Much Hate?

Key Takeaways

  • One bidirectional Transformer handles text and raw image patches, with no vision tower and no decoder.
  • NeoMME-Retriever-260M scores 0.523 nDCG@10 on ViDoRe v3, beating every evaluated model below 800M.
  • It matches 3.75B-parameter ColQwen2.5 on ViDoRe v3 while being 14.4× smaller.
  • Token pooling plus asymmetric quantization cut the index from roughly 1.5 MB to 6 kB per page.
  • Text-only retrieval and frozen natural-image transfer remain clear weak spots.

Check out the Paper, Model Collection and Demo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Send High-Quality Images And Videos From Android To iPhone
AI & Technology

How To Send High-Quality Images And Videos From Android To iPhone

September 6, 2026
What Is Vibe Coding And Why Does It Get So Much Hate?
AI & Technology

What Is Vibe Coding And Why Does It Get So Much Hate?

September 6, 2026
My Content Tracker Idea Became a Real App – Unite.AI
AI & Technology

My Content Tracker Idea Became a Real App – Unite.AI

September 6, 2026
Is 256GB Enough For An iPhone? Here’s When You Should Go Bigger
AI & Technology

Is 256GB Enough For An iPhone? Here’s When You Should Go Bigger

September 6, 2026
Next Post
Is It Too Late For Nike? NKE Stock Deep Dive

Is It Too Late For Nike? NKE Stock Deep Dive

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Truck bursts into flames after being struck by lightning

Truck bursts into flames after being struck by lightning

September 1, 2026
Anthropic back on ‘right side’ with Trump admin: Howard Lutnick

Anthropic back on ‘right side’ with Trump admin: Howard Lutnick

September 2, 2026
Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device

Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device

September 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!