• bitcoinBitcoin(BTC)$85,888.006.41%
  • ethereumEthereum(ETH)$2,736.765.34%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$802.646.37%
  • rippleXRP(XRP)$1.497.61%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$117.648.74%
  • tronTRON(TRX)$0.3450770.20%
  • zcashZcash(ZEC)$1,499.254.35%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • HyperliquidHyperliquid(HYPE)$93.342.09%
  • dogecoinDogecoin(DOGE)$0.09586612.12%
  • moneroMonero(XMR)$571.465.36%
  • whitebitWhiteBIT Coin(WBT)$86.405.18%
  • RainRain(RAIN)$0.0141318.07%
  • chainlinkChainlink(LINK)$12.985.85%
  • USDSUSDS(USDS)$1.000.02%
  • cardanoCardano(ADA)$0.2434149.09%
  • leo-tokenLEO Token(LEO)$8.980.46%
  • stellarStellar(XLM)$0.2093288.24%
  • uniswapUniswap(UNI)$8.832.43%
  • bitcoin-cashBitcoin Cash(BCH)$263.136.50%
  • nearNEAR Protocol(NEAR)$4.048.92%
  • avalanche-2Avalanche(AVAX)$11.190.48%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • litecoinLitecoin(LTC)$62.298.88%
  • CantonCanton(CC)$0.1160189.20%
  • daiDai(DAI)$1.000.03%
  • USD1USD1(USD1)$1.000.03%
  • suiSui(SUI)$1.0221.58%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.434.44%
  • hedera-hashgraphHedera(HBAR)$0.0908694.72%
  • shiba-inuShiba Inu(SHIB)$0.0000068.87%
  • MemeCoreMemeCore(M)$1.49-3.24%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • BittensorBittensor(TAO)$285.1012.87%
  • crypto-com-chainCronos(CRO)$0.0636228.63%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,354.66-0.34%
  • okbOKB(OKB)$123.295.97%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9124.72%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.41%
  • aaveAave(AAVE)$143.746.23%
  • EthenaEthena(ENA)$0.2161027.57%
  • OndoOndo(ONDO)$0.4396096.49%
  • mantleMantle(MNT)$0.647.20%
  • pepePepe(PEPE)$0.00000521.92%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Researchers Present Magma: A Multimodal AI Model Integrating Vision, Language, and Action for Advanced Robotics, UI Navigation, and Intelligent Decision-Making

February 20, 2025
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Researchers Present Magma: A Multimodal AI Model Integrating Vision, Language, and Action for Advanced Robotics, UI Navigation, and Intelligent Decision-Making
ShareShareShareShareShare

Multimodal AI agents are designed to process and integrate various data types, such as images, text, and videos, to perform tasks in digital and physical environments. They are used in robotics, virtual assistants, and user interface automation, where they need to understand and act based on complex multimodal inputs. These systems aim to bridge verbal and spatial intelligence by leveraging deep learning techniques, enabling interactions across multiple domains.

AI systems often specialize in vision-language understanding or robotic manipulation but struggle to combine these capabilities into a single model. Many AI models are designed for domain-specific tasks, such as UI navigation in digital environments or physical manipulation in robotics, limiting their generalization across different applications. The challenge lies in developing a unified model to understand and act across multiple modalities, ensuring effective decision-making in structured and unstructured environments.

YOU MAY ALSO LIKE

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

How To Choose The Right USB To USB-C Adapter

Existing Vision-Language-Action (VLA) models attempt to address multimodal tasks by pretraining on large datasets of vision-language pairs followed by action trajectory data. However, these models typically lack adaptability across different environments. Examples include Pix2Act and WebGUM, which excel in UI navigation, and OpenVLA and RT-2, which are optimized for robotic manipulation. These models often require separate training processes and fail to generalize across both digital and physical environments. Also, conventional multimodal models struggle with integrating spatial and temporal intelligence, limiting their ability to perform complex tasks autonomously.

Researchers from Microsoft Research, the University of Maryland, the University of Wisconsin-Madison KAIST, and the University of Washington introduced Magma, a foundation model designed to unify multimodal understanding with action execution, enabling AI agents to function seamlessly in digital and physical environments. Magma is designed to overcome the shortcomings of existing VLA models by incorporating a robust training methodology that integrates multimodal understanding, action grounding, and planning. Magma is trained using a diverse dataset comprising 39 million samples, including images, videos, and robotic action trajectories. It incorporates two novel techniques, 

  1. Set-of-Mark (SoM): SoM enables the model to label actionable visual objects, such as buttons in UI environments
  2. Trace-of-Mark (ToM): ToM allows it to track object movements over time and plan future actions accordingly

Magma employs a combination of deep learning architectures and large-scale pretraining to optimize its performance across multiple domains. The model uses a ConvNeXt-XXL vision backbone to process images and videos, while an LLaMA-3-8B language model handles textual inputs. This architecture enables Magma to integrate vision-language understanding with action execution seamlessly. It is trained on a curated dataset that includes UI navigation tasks from SeeClick and Vision2UI, robotic manipulation datasets from Open-X-Embodiment, and instructional videos from sources like Ego4D, Something-Something V2, and Epic-Kitchen. By leveraging SoM and ToM, Magma can effectively learn action grounding from UI screenshots and robotics data while enhancing its ability to predict future actions based on observed visual sequences. During training, the model processes up to 2.7 million UI screenshots, 970,000 robotic trajectories, and over 25 million video samples to ensure robust multimodal learning.

In zero-shot UI navigation tasks, Magma achieved an element selection accuracy of 57.2%, outperforming models like GPT-4V-OmniParser and SeeClick. In robotic manipulation tasks, Magma attained a success rate of 52.3% in Google Robot tasks and 35.4% in Bridge simulations, significantly surpassing OpenVLA, which only achieved 31.7% and 15.9% in the same benchmarks. The model also performed exceptionally well in multimodal understanding tasks, reaching 80.0% accuracy in VQA v2, 66.5% in TextVQA, and 87.4% in POPE evaluations. Magma also demonstrated strong spatial reasoning capabilities, scoring 74.8% on the BLINK dataset and 80.1% on the Visual Spatial Reasoning (VSR) benchmark. In video question-answering tasks, Magma achieved an accuracy of 88.6% on IntentQA and 72.9% on NextQA, further highlighting its ability to process temporal information effectively.

Several Key Takeaways emerge from the Research on Magma:

  1. Magma was trained on 39 million multimodal samples, including 2.7 million UI screenshots, 970,000 robotic trajectories, and 25 million video samples.
  2. The model combines vision, language, and action in a unified framework, overcoming the limitations of domain-specific AI models.
  3. SoM enables accurate labeling of clickable objects, while ToM allows tracking object movement over time, improving long-term planning capabilities.
  4. Magma achieved a 57.2% accuracy rate in element selection in UI tasks, a 52.3% success rate in robotic manipulation, and an 80.0% accuracy rate in VQA tasks.
  5. Magma outperformed existing AI models by over 19.6% in spatial reasoning benchmarks and improved by 28% over previous models in video-based reasoning.
  6. Magma demonstrated superior generalization across multiple tasks without requiring additional fine-tuning, making it a highly adaptable AI agent.
  7. Magma’s capabilities can enhance decision-making and execution in robotics, autonomous systems, UI automation, digital assistants, and industrial AI.

Check out the Paper and Project Page. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 75k+ ML SubReddit.

🚨 Recommended Read- LG AI Research Releases NEXUS: An Advanced System Integrating Agent AI System and Data Compliance Standards to Address Legal Concerns in AI Datasets


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI
AI & Technology

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

September 21, 2026
How To Choose The Right USB To USB-C Adapter
AI & Technology

How To Choose The Right USB To USB-C Adapter

September 21, 2026
A Laptop That Works Better With Your Android Phone
AI & Technology

A Laptop That Works Better With Your Android Phone

September 21, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI
AI & Technology

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
Next Post
Growing feud between tech billionaires Musk and Altman

Growing feud between tech billionaires Musk and Altman

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
JD Vance says U.S. is not at ‘war’ with Iran

JD Vance says U.S. is not at ‘war’ with Iran

September 18, 2026
LIVE NOW: FOMC RATE DECISION & PRESS CONFERENCE 2026

LIVE NOW: FOMC RATE DECISION & PRESS CONFERENCE 2026

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!