• bitcoinBitcoin(BTC)$85,695.000.17%
  • ethereumEthereum(ETH)$2,697.20-0.39%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$781.78-0.60%
  • rippleXRP(XRP)$1.510.58%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$120.900.74%
  • tronTRON(TRX)$0.336190-0.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.04-2.55%
  • zcashZcash(ZEC)$1,359.282.25%
  • HyperliquidHyperliquid(HYPE)$91.99-1.57%
  • dogecoinDogecoin(DOGE)$0.094518-0.23%
  • moneroMonero(XMR)$555.16-0.19%
  • chainlinkChainlink(LINK)$13.950.49%
  • cardanoCardano(ADA)$0.2713472.61%
  • whitebitWhiteBIT Coin(WBT)$85.550.08%
  • USDSUSDS(USDS)$1.000.00%
  • leo-tokenLEO Token(LEO)$8.91-0.42%
  • RainRain(RAIN)$0.011010-4.67%
  • stellarStellar(XLM)$0.2138890.60%
  • nearNEAR Protocol(NEAR)$5.05-1.13%
  • bitcoin-cashBitcoin Cash(BCH)$315.230.15%
  • uniswapUniswap(UNI)$8.64-4.54%
  • litecoinLitecoin(LTC)$69.13-1.93%
  • CantonCanton(CC)$0.1297743.59%
  • avalanche-2Avalanche(AVAX)$11.465.31%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • suiSui(SUI)$1.19-0.22%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.100654-0.31%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.541.67%
  • quant-networkQuant(QNT)$257.020.81%
  • BittensorBittensor(TAO)$303.671.68%
  • shiba-inuShiba Inu(SHIB)$0.000006-1.26%
  • tether-goldTether Gold(XAUT)$4,167.610.64%
  • crypto-com-chainCronos(CRO)$0.066657-2.33%
  • BitwayBitway(BTW)$1.20-4.34%
  • EthenaEthena(ENA)$0.238477-4.31%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • okbOKB(OKB)$138.828.21%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • Pump.funPump.fun(PUMP)$0.006218-1.14%
  • aaveAave(AAVE)$181.91-0.06%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • OndoOndo(ONDO)$0.4976131.77%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.052.07%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

DeepMind Debuts EmbeddingGemma 2, Mapping Five Modalities Into One Space – Unite.AI

October 6, 2026
in AI & Technology
Reading Time: 4 mins read
A A
DeepMind Debuts EmbeddingGemma 2, Mapping Five Modalities Into One Space – Unite.AI
ShareShareShareShareShare

Google DeepMind launched EmbeddingGemma 2 on October 6, 2026, a 740-million-parameter open model that maps text, code, images, video, and audio into a single 768-dimensional embedding space, released under the Apache 2.0 license and designed to run on consumer hardware.

YOU MAY ALSO LIKE

Don’t Throw Away Your Old Digital Camera — Do This Instead

Affordable Alternatives To Nintendo’s Switch 2 Pro Controller

The model was unveiled in a post on Google’s official blog by Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera. Google describes EmbeddingGemma 2 as the most capable model for on-device multimodal embeddings and says it is built from the same technology as its Gemini Embedding models. The company lists capability examples such as retrieving a specific video clip with a voice memo or querying hours of audio recordings with text.

The release follows the original EmbeddingGemma, which Google introduced in 2025 as a lightweight option for text embeddings on consumer hardware. Google reports that model has passed 20 million downloads, with developers using it for on-device search tools and privacy-first retrieval-augmented generation pipelines.

Modular Encoders on a Gemma 4 Base

EmbeddingGemma 2 is built on the Gemma 4 architecture. According to the EmbeddingGemma 2 model card, its 740 million parameters combine a 270-million-parameter text model (a 130-million transformer backbone plus a 140-million embedder) with a 170-million-parameter vision encoder and a 300-million-parameter audio encoder, all projecting into the shared 768-dimensional space. Because the encoders are independent, developers can selectively load 270 million parameters for text and code, 440 million for text and vision, 570 million for text and audio, or the full multimodal configuration from the same checkpoint, and every configuration shares one vector space.

The model card lists a 24-layer architecture with grouped-query and multi-query attention, a 262,144-token vocabulary, mean pooling, and a projection layer from 512 to 768 dimensions. All modalities share an 8,192-token context window, which Google says is four times larger than EmbeddingGemma 1’s. Each modality consumes that budget at fixed rates: 280 tokens per image, 140 tokens per video frame, and 25 tokens per second of audio, so a single input can hold up to 29 images, 58 video frames, or 5.5 minutes of audio. Interleaved inputs mix text and media, with placeholder tokens marking each media item’s position.

Benchmark Results and Vector Truncation

Google reports that EmbeddingGemma 2 matches its predecessor’s multilingual text performance while raising its MTEB Code score by 9.92 points, from 68.76 to 78.68. The model card’s full-precision results list 61.36 on MTEB multilingual (v2), 64.64 on MIEB lite, 57.28 on MMEB v2 image retrieval, 67.84 on visual-document retrieval, 50.67 on video retrieval, 69.54 on MSEB sound retrieval, and 49.39 on MAEB audio tasks. Google says the model achieves leading scores among multimodal embedders under one billion parameters and outperforms some specialist models more than twice its size.

Matryoshka Representation Learning lets developers truncate output vectors from the native 768 dimensions down to 512, 256, or 128, which Google says cuts vector storage requirements by up to 6x. According to the model card, quality holds with minimal impact down to 256 dimensions, while 128 dimensions is best suited to text-only workloads. A companion developer guide reports that 256-dimensional vectors retain about 95 percent of full quality on image, video, and speech retrieval, while 128 dimensions retains around 90 percent on text and code but drops to roughly 75 percent on multimodal retrieval. Storing one million 768-dimensional vectors in bfloat16 precision takes roughly 1.5 gigabytes of memory, the guide states, versus about 250 megabytes at 128 dimensions.

Safety Posture and Stated Limitations

The model card describes EmbeddingGemma 2 as a pre-trained embedding model that has not undergone post-training alignment, safety tuning, or output-level moderation, with safety mitigations concentrated on filtering the pre-training data. That preparation included filtering for child sexual abuse material at multiple stages and automated filtering of personal information and other sensitive data. Training data spanned web documents, code, images, video, audio, and paired cross-modality samples, with web text in more than 140 languages and a cutoff of January 2025.

The card notes that while the model supports more than 100 languages, performance may not be equal across them, and that omitting the recommended task instruction prefixes on text inputs reduces embedding precision. It also warns that the model’s activation range exceeds the dynamic range of float16, which can yield NaN values or silently degraded embeddings, and directs inference to bfloat16 or float32. Deployments must adhere to the Gemma Prohibited Use Policy, and the card assigns developers responsibility for application-level safeguards such as retrieval filtering and fairness testing.

On-Device Performance and Availability

Google reports that, with quantization on a Pixel 11 Pro, EmbeddingGemma 2 requires as little as roughly 191 megabytes of active RAM for text-only weights and about 567 megabytes for the full multimodal model. Weights are available on Hugging Face and Kaggle, with on-device-optimized versions through the LiteRT Community on Hugging Face. The model runs through transformers, sentence-transformers 6.1.0 or later, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, with fine-tuning guidance from Unsloth and browser deployment through transformers.js and WebGPU.

Google AI Edge’s MediaPipe and LiteRT handle cross-platform deployment, and a MediaPipe Decision Task API supports real-time classification and routing on multimodal context. Demos including Instant Media Search and Video Moments Finder ship in the Google AI Edge Gallery app, and the Google AI Edge Foresight app pairs EmbeddingGemma 2 for local file retrieval with Gemma 4 for contextual reasoning. Because the two models share a text tokenizer and audio encoder architecture, Google says running them together lowers their combined memory footprint. Availability in the Gemini Enterprise Agent Platform Model Garden is coming soon, according to the company.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Don’t Throw Away Your Old Digital Camera — Do This Instead
AI & Technology

Don’t Throw Away Your Old Digital Camera — Do This Instead

October 6, 2026
Affordable Alternatives To Nintendo’s Switch 2 Pro Controller
AI & Technology

Affordable Alternatives To Nintendo’s Switch 2 Pro Controller

October 6, 2026
Sushil Kumar, CEO of Cyara – Interview Series – Unite.AI
AI & Technology

Sushil Kumar, CEO of Cyara – Interview Series – Unite.AI

October 6, 2026
Spotify Expands Its Music Quiz Trivia Feature
AI & Technology

Spotify Expands Its Music Quiz Trivia Feature

October 6, 2026
Next Post
Here are the top 3 AI stocks Kenny Polcari would own right now 👀💰

Here are the top 3 AI stocks Kenny Polcari would own right now 👀💰

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Red Sox vs. Yankees live updates, news, starting pitchers for Game 2: Sonny Gray, Max Fried take the mound Wednesday – Yahoo Sports

Red Sox vs. Yankees live updates, news, starting pitchers for Game 2: Sonny Gray, Max Fried take the mound Wednesday – Yahoo Sports

October 1, 2026
OpenAI Fires Three Employees Who Allegedly Shared Info With An External AI Safety Group

OpenAI Fires Three Employees Who Allegedly Shared Info With An External AI Safety Group

October 1, 2026
Two Iranians charged over alleged plot targeting Jewish community in UK – Al Jazeera

Two Iranians charged over alleged plot targeting Jewish community in UK – Al Jazeera

October 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!