• bitcoinBitcoin(BTC)$84,016.000.57%
  • ethereumEthereum(ETH)$2,696.721.54%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$776.09-0.05%
  • rippleXRP(XRP)$1.596.41%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$120.735.39%
  • tronTRON(TRX)$0.336475-0.88%
  • zcashZcash(ZEC)$1,572.855.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.91%
  • HyperliquidHyperliquid(HYPE)$92.380.62%
  • dogecoinDogecoin(DOGE)$0.0979844.00%
  • chainlinkChainlink(LINK)$14.0512.73%
  • moneroMonero(XMR)$552.651.02%
  • whitebitWhiteBIT Coin(WBT)$83.910.29%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2561494.30%
  • RainRain(RAIN)$0.011855-1.55%
  • leo-tokenLEO Token(LEO)$8.82-0.89%
  • stellarStellar(XLM)$0.2174145.91%
  • bitcoin-cashBitcoin Cash(BCH)$334.450.02%
  • nearNEAR Protocol(NEAR)$5.1114.19%
  • uniswapUniswap(UNI)$9.746.44%
  • litecoinLitecoin(LTC)$70.04-3.56%
  • CantonCanton(CC)$0.12496014.63%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.1314.62%
  • avalanche-2Avalanche(AVAX)$10.411.82%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.02%
  • hedera-hashgraphHedera(HBAR)$0.0942051.50%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.420.20%
  • BittensorBittensor(TAO)$305.016.73%
  • shiba-inuShiba Inu(SHIB)$0.0000062.21%
  • BitwayBitway(BTW)$1.2216.87%
  • crypto-com-chainCronos(CRO)$0.0659686.06%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-3.94%
  • tether-goldTether Gold(XAUT)$4,277.640.58%
  • OndoOndo(ONDO)$0.549.34%
  • EthenaEthena(ENA)$0.25163916.91%
  • okbOKB(OKB)$120.401.17%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • aaveAave(AAVE)$149.315.37%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • mantleMantle(MNT)$0.67-0.35%
  • polkadotPolkadot(DOT)$1.193.32%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google Introduces T5Gemma 2: Encoder Decoder Models with Multimodal Inputs via SigLIP and 128K Context

December 19, 2025
in AI & Technology
Reading Time: 8 mins read
A A
Google Introduces T5Gemma 2: Encoder Decoder Models with Multimodal Inputs via SigLIP and 128K Context
ShareShareShareShareShare

Google has published T5Gemma 2, a family of open encoder-decoder Transformer checkpoints built by adapting Gemma 3 pretrained weights into an encoder-decoder layout, then continuing pretraining with the UL2 objective. The release is pretrained only, intended for developers to post-train for specific tasks, and Google explicitly notes it is not releasing post-trained or IT checkpoints for this drop.

T5Gemma 2 is positioned as an encoder-decoder counterpart to Gemma 3 that keeps the same low level building blocks, then adds 2 structural changes aimed at small model efficiency. The models inherit Gemma 3 features that matter for deployment, notably multimodality, long context up to 128K tokens, and broad multilingual coverage, with the blog stating over 140 languages.

YOU MAY ALSO LIKE

Microsoft’s Copilot App Adds Office, Natural Coding And Automation

Google Adds Creepy Avatars To Gemini 3.8 Live’s Agents

https://arxiv.org/pdf/2512.14856

What Google actually released?

The release includes 3 pretrained sizes, 270M-270M, 1B-1B, and 4B-4B, where the notation means the encoder and decoder are the same size. The research team reports approximate totals excluding the vision encoder, about 370M, 1.7B, and 7B parameters. The multimodal accounting lists a 417M parameter vision encoder, along with encoder and decoder parameters broken into embedding and non embedding components.

The adaptation, encoder-decoder without training from scratch

T5Gemma 2 follows the same adaptation idea introduced in T5Gemma, initialize an encoder-decoder model from a decoder-only checkpoint, then adapt with UL2. In the above figure the research team show encoder and decoder parameters initialized from the pretrained decoder-only model, then pretrained with UL2, with images first converted by SigLIP into 256 tokens.

This matters because encoder-decoder splits the workload, the encoder can read the full input bidirectionally, while the decoder focuses on autoregressive generation. The research team argues this separation can help long context tasks where the model must retrieve relevant evidence from a large input before generating.

Two efficiency changes that are easy to miss but affect small models

First, T5Gemma 2 uses tied word embeddings across encoder input embedding, decoder input embedding, and decoder output or softmax embedding. This reduces parameter redundancy, and references an ablation showing little quality change while reducing embedding parameters.

Second, it introduces merged attention in the decoder. Instead of separate self-attention and cross-attention sublayers, the decoder performs a single attention operation where K and V are formed by concatenating encoder outputs and decoder states, and masking preserves causal visibility for decoder tokens. This ties to easier initialization, because it narrows differences between the adapted decoder and the original Gemma style decoder stack, and it reports parameter savings with a small average quality drop in their ablations.

https://arxiv.org/pdf/2512.14856
https://arxiv.org/pdf/2512.14856

Multimodality, image understanding is encoder side, not decoder side

T5Gemma 2 is multimodal by reusing Gemma 3’s vision encoder and keeping it frozen during training. Vision tokens are always fed to the encoder and encoder tokens have full visibility to each other in self attention. This is a pragmatic encoder-decoder design, the encoder fuses image tokens with text tokens into contextual representations, and the decoder can then attend to those representations while generating text.

On the tooling side, T5Gemma 2 is placed under an image-text-to-text pipeline, which matches the research’s design, image in, text prompt in, text out. That pipeline example is the fastest way to validate the end to end multimodal path, including dtype choices like bfloat16 and automatic device mapping.

Long context to 128K, what enables it

Google researchers attributes the 128K context window to Gemma 3’s alternating local and global attention mechanism. The Gemma 3 team describes a repeating 5 to 1 pattern, 5 local sliding window attention layers followed by 1 global attention layer, with a local window size of 1024. This design reduces KV cache growth relative to making every layer global, which is one reason long context becomes feasible at smaller footprints.

In the T5Gemma 2, the research team also mention adopting positional interpolation methods for long context, and they pretrain on sequences up to 16K input paired with 16K target outputs, then evaluate long context performance up to 128K on benchmarks including RULER and MRCR. The detailed pretraining results table includes 32K and 128K evaluations, showing the long context deltas they claim over Gemma 3 at the same scale.

https://arxiv.org/pdf/2512.14856

Training setup and what “pretrained only” implies for users

The research team states the models are pretrained on 2T tokens and describes a training setup that includes a batch size of 4.2M tokens, cosine learning rate decay with 100 warmup steps, global gradient clipping at 1.0, and checkpoint averaging over the last 5 checkpoints.

Key Takeaways

  1. T5Gemma 2 is an encoder decoder family adapted from Gemma 3 and continued with UL2, it reuses Gemma 3 pretrained weights, then applies the same UL2 based adaptation recipe used in T5Gemma.
  2. Google released pretrained checkpoints only, no post trained or instruction tuned variants are included in this drop, so downstream use requires your own post training and evaluation.
  3. Multimodal input is handled by a SigLIP vision encoder that outputs 256 image tokens and stays frozen, those vision tokens go into the encoder, the decoder generates text.
  4. Two parameter efficiency changes are central, tied word embeddings share encoder, decoder, and output embeddings, merged attention unifies decoder self attention and cross attention into a single module.
  5. Long context up to 128K is enabled by Gemma 3’s interleaved attention design, a repeating 5 local sliding window layers with window size 1024 followed by 1 global layer, and T5Gemma 2 inherits this mechanism.

Check out the Paper, Technical details and Model on Hugging Face. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Google Introduces T5Gemma 2: Encoder Decoder Models with Multimodal Inputs via SigLIP and 128K Context appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Microsoft’s Copilot App Adds Office, Natural Coding And Automation
AI & Technology

Microsoft’s Copilot App Adds Office, Natural Coding And Automation

September 25, 2026
Google Adds Creepy Avatars To Gemini 3.8 Live’s Agents
AI & Technology

Google Adds Creepy Avatars To Gemini 3.8 Live’s Agents

September 25, 2026
Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
AI & Technology

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

September 25, 2026
Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120
AI & Technology

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

September 25, 2026
Next Post
LIVE: Supreme Court hears oral arguments in the FTC firing case | NBC News

LIVE: Supreme Court hears oral arguments in the FTC firing case | NBC News

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Bain Capital Ventures Bets .6 Billion on AI’s Next Act

Bain Capital Ventures Bets $1.6 Billion on AI’s Next Act

September 20, 2026
LIVE: Trump holds back-to-school event at the White House | NBC News

LIVE: Trump holds back-to-school event at the White House | NBC News

September 25, 2026
New York Times Cooking Is Coming To Meta’s AI And Display Glasses

New York Times Cooking Is Coming To Meta’s AI And Display Glasses

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!