• bitcoinBitcoin(BTC)$84,384.000.30%
  • ethereumEthereum(ETH)$2,682.480.85%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$780.142.36%
  • rippleXRP(XRP)$1.520.37%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.301.90%
  • tronTRON(TRX)$0.338955-0.17%
  • zcashZcash(ZEC)$1,518.05-2.78%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.85%
  • HyperliquidHyperliquid(HYPE)$93.700.81%
  • dogecoinDogecoin(DOGE)$0.0956762.98%
  • moneroMonero(XMR)$548.89-0.16%
  • whitebitWhiteBIT Coin(WBT)$84.33-0.14%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.754.44%
  • cardanoCardano(ADA)$0.2498085.03%
  • RainRain(RAIN)$0.012051-3.45%
  • leo-tokenLEO Token(LEO)$8.90-0.95%
  • stellarStellar(XLM)$0.2086632.81%
  • bitcoin-cashBitcoin Cash(BCH)$339.13-1.59%
  • nearNEAR Protocol(NEAR)$4.606.54%
  • litecoinLitecoin(LTC)$74.4625.00%
  • uniswapUniswap(UNI)$9.292.24%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.400.67%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.1109863.36%
  • suiSui(SUI)$1.014.94%
  • hedera-hashgraphHedera(HBAR)$0.0920962.33%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.431.70%
  • shiba-inuShiba Inu(SHIB)$0.0000062.90%
  • BittensorBittensor(TAO)$291.470.20%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0625932.19%
  • BitwayBitway(BTW)$1.0510.95%
  • MemeCoreMemeCore(M)$1.230.25%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,258.81-0.59%
  • okbOKB(OKB)$119.631.32%
  • OndoOndo(ONDO)$0.5225.08%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.13%
  • mantleMantle(MNT)$0.684.56%
  • aaveAave(AAVE)$144.273.65%
  • EthenaEthena(ENA)$0.2175987.36%
  • polkadotPolkadot(DOT)$1.176.62%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet M3-Agent: A Multimodal Agent with Long-Term Memory and Enhanced Reasoning Capabilities

August 20, 2025
in AI & Technology
Reading Time: 3 mins read
A A
Meet M3-Agent: A Multimodal Agent with Long-Term Memory and Enhanced Reasoning Capabilities
ShareShareShareShareShare

In the future, a home robot could manage daily chores itself and learn household patterns from ongoing experience. It may serve coffee in the morning without asking, having remembered your habits over time. For a multimodal agent, this intelligence depends on (a) observing the world through multimodal sensors continuously, (b) storing its experience in long-term memories, and (c) reasoning over this memory to guide its actions. Current research is focused on LLM-based agents, but multimodal agents process diverse inputs and store richer, multimodal content. This poses new challenges in maintaining consistency in long-term memory. Instead of simply storing descriptive experiences, multimodal agents must build internal world knowledge similar to how humans learn.

Existing attempts include appending raw agent trajectories, such as dialogues or execution histories, directly to memory. Some methods enhance this by combining summaries, latent embeddings, or structured knowledge representations. In multimodal agents, memory formation is closely tied to online video understanding, where early methods like extending context windows or compressing visual tokens often fail to scale for long video streams. Memory-based methods, which store encoded visual features, improve scalability but struggle with maintaining long-term consistency. The Socratic Models framework generates language-based memory to describe videos, offering scalability, but faces challenges in tracking evolving events and entities over time.

YOU MAY ALSO LIKE

Apple Explores Screenless Fitness Tracker to Rival Whoop

Razer’s Kiyo V2 Pro Webcam Can Capture 4K Video At 60 Fps

Researchers from ByteDance Seed, Zhejiang University, and Shanghai Jiao Tong University have proposed M3-Agent, a multimodal agent framework with long-term memory. M3-Agent processes real-time visual and auditory inputs to build and update its memory, just like humans. Unlike standard episodic memory, it also develops semantic memory, allowing the accumulation of world knowledge over time. Its memory is organized in an entity-centric, multimodal structure, ensuring a deeper and more coherent understanding of the environment. When given instructions, M3-Agent engages in multi-turn reasoning and autonomously retrieves relevant information. Moreover, M3-Bench is developed for long-video question answering to evaluate the effectiveness of M3-Agent.

M3-Agent contains a multimodal LLM and a long-term memory module, operating through two parallel processes: memorization and control. Long-term memory is an external database that stores structured, multimodal data in a memory graph, where nodes represent distinct memory items with unique IDs, modalities, raw content, embeddings, and metadata. During memorization, M3-Agent processes video streams clip by clip, generating episodic memory for raw content and semantic memory for abstract knowledge, such as identities and relationships. For control, the agent conducts multi-turn reasoning, using search functions to fetch relevant memory in up to H rounds. RL optimizes the framework, with separate models trained for memorization and control to achieve peak performance.

M3-Agent and all baselines are evaluated on both M3-Bench-robot and M3-Bench-web. On M3-Bench-robot, M3-agent achieves a 6.3% accuracy improvement over the strongest baseline, MA-LLM, while on M3-Bench-web and VideoMME-long, it outperforms GeminiGPT4o-Hybrid by 7.7% and 5.3%, respectively. Moreover, M3-Agent outperforms MA-LMM by 4.2% in human understanding and 8.5% in cross-modal reasoning on M3-Bench-robot. On M3-Bench-web, it outperforms Gemini-GPT4o-Hybrid with 15.5% gain and 6.7% in these categories. These results underscore M3-Agent’s ability to maintain character consistency, enhance human understanding, and effectively integrate multimodal information.

In conclusion, researchers introduced M3-Agent, a multimodal framework with long-term memory, capable of processing real-time video and audio streams to build episodic and semantic memories. This enables the agent to accumulate world knowledge and maintain consistent, context-rich memory over time. Experimental results show that M3-Agent outperforms all baselines across multiple benchmarks. Detailed case studies highlight current limitations and suggest future directions, such as improving attention mechanisms for semantic memory and developing more efficient visual memory systems. These advancements pave the way for more human-like AI agents in practical applications.


Check out the Paper and GitHub Page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

📥 Sponsorship Media Kit

The post Meet M3-Agent: A Multimodal Agent with Long-Term Memory and Enhanced Reasoning Capabilities appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple Explores Screenless Fitness Tracker to Rival Whoop
AI & Technology

Apple Explores Screenless Fitness Tracker to Rival Whoop

September 24, 2026
Razer’s Kiyo V2 Pro Webcam Can Capture 4K Video At 60 Fps
AI & Technology

Razer’s Kiyo V2 Pro Webcam Can Capture 4K Video At 60 Fps

September 24, 2026
How To Get Your Cut Of Apple’s 0 Million Siri Settlement
AI & Technology

How To Get Your Cut Of Apple’s $250 Million Siri Settlement

September 24, 2026
Revolut Is Piloting Facial Recognition At Store Checkouts In The UK
AI & Technology

Revolut Is Piloting Facial Recognition At Store Checkouts In The UK

September 24, 2026
Next Post
A Coding Implementation to Build a Complete Self-Hosted LLM Workflow with Ollama, REST API, and Gradio Chat Interface

A Coding Implementation to Build a Complete Self-Hosted LLM Workflow with Ollama, REST API, and Gradio Chat Interface

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Good Samaritans lift car off woman after e-bike accident

Good Samaritans lift car off woman after e-bike accident

September 18, 2026
Lindsay Clancy Trial: An inside look at the evidence at the center of the case as jurors deliberate

Lindsay Clancy Trial: An inside look at the evidence at the center of the case as jurors deliberate

September 22, 2026
Graham wins South Carolina Republican Senate primary runoff, NBC News projects

Graham wins South Carolina Republican Senate primary runoff, NBC News projects

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!