• bitcoinBitcoin(BTC)$84,517.00-2.04%
  • ethereumEthereum(ETH)$2,675.29-2.56%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$766.46-2.65%
  • rippleXRP(XRP)$1.53-2.67%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$115.08-1.63%
  • tronTRON(TRX)$0.340030-0.47%
  • zcashZcash(ZEC)$1,599.303.49%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.61%
  • HyperliquidHyperliquid(HYPE)$94.930.14%
  • dogecoinDogecoin(DOGE)$0.094798-5.20%
  • moneroMonero(XMR)$553.86-3.56%
  • whitebitWhiteBIT Coin(WBT)$84.80-2.15%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.31-5.43%
  • cardanoCardano(ADA)$0.240426-4.41%
  • RainRain(RAIN)$0.012522-6.84%
  • leo-tokenLEO Token(LEO)$8.980.01%
  • stellarStellar(XLM)$0.206054-3.66%
  • bitcoin-cashBitcoin Cash(BCH)$342.765.80%
  • nearNEAR Protocol(NEAR)$4.665.50%
  • uniswapUniswap(UNI)$9.484.41%
  • Ethena USDeEthena USDe(USDE)$1.000.08%
  • litecoinLitecoin(LTC)$60.76-1.32%
  • avalanche-2Avalanche(AVAX)$10.45-6.12%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.108733-6.10%
  • suiSui(SUI)$0.98-2.53%
  • hedera-hashgraphHedera(HBAR)$0.091755-5.71%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-1.61%
  • BittensorBittensor(TAO)$302.13-4.88%
  • shiba-inuShiba Inu(SHIB)$0.000006-5.05%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.062198-6.98%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.21-7.64%
  • tether-goldTether Gold(XAUT)$4,291.90-1.02%
  • BitwayBitway(BTW)$0.969.82%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$119.00-2.19%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.07%
  • aaveAave(AAVE)$141.80-0.89%
  • mantleMantle(MNT)$0.65-0.61%
  • EthenaEthena(ENA)$0.206358-1.59%
  • OndoOndo(ONDO)$0.419297-2.42%
  • Pump.funPump.fun(PUMP)$0.004094-7.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Splunk Researchers Introduce MAG-V: A Multi-Agent Framework For Synthetic Data Generation and Reliable AI Trajectory Verification

December 11, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Splunk Researchers Introduce MAG-V: A Multi-Agent Framework For Synthetic Data Generation and Reliable AI Trajectory Verification
ShareShareShareShareShare

These days, large language models (LLMs) are getting integrated with multi-agent systems, where multiple intelligent agents collaborate to achieve a unified objective. Multi-agent frameworks are designed to improve problem-solving, enhance decision-making, and optimize the ability of AI systems to address diverse user needs. By distributing responsibilities among agents, these systems ensure better task execution and offer scalable solutions. They are valuable in applications like customer support, where accurate responses and adaptability are paramount.

However, to deploy these multi-agent systems, realistic and scalable datasets need to be created for testing and training. The scarcity of domain-specific data and privacy concerns surrounding proprietary information limits the ability to train AI systems effectively. Also, customer-facing AI agents must maintain logical reasoning and correctness when navigating through sequences of actions or trajectories to arrive at solutions. This process often involves external tool calls, resulting in errors if the wrong sequence or parameters are used. These inaccuracies lead to diminished user trust and reduced system reliability, creating a critical need for more robust methods to verify agent trajectories and generate realistic test datasets.

YOU MAY ALSO LIKE

Never Use ChatGPT For These Five Tasks

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

Traditionally, addressing these challenges involved relying on human-labeled data or leveraging LLMs as judges to verify trajectories. While LLM-based solutions have shown promise, they face significant limitations, including sensitivity to input prompts, inconsistent outputs from API-based models, and high operational costs. Also, these approaches are time-intensive and need to scale more effectively, especially when applied to complex domains that demand precise and context-aware responses. Consequently, there is an urgent need for a cost-effective and deterministic solution to validate AI agent behaviors and ensure reliable outcomes.

Researchers at Splunk Inc. have proposed an innovative framework called MAG-V (Multi-Agent Framework for Synthetic Data Generation and Verification), which aims to overcome these limitations. MAG-V is a multi-agent system designed to generate synthetic datasets and verify the trajectories of AI agents. The framework introduces a novel approach combining classical machine-learning techniques with advanced LLM capabilities. Unlike traditional systems, MAG-V does not rely on LLMs as feedback mechanisms. Instead, it utilizes deterministic methods and machine-learning models to ensure accuracy and scalability in trajectory verification.

MAG-V uses three specialized agents: 

  1. An investigator: The investigator generates questions that mimic realistic customer queries
  2. An assistant: The assistant responds based on predefined trajectories
  3. A reverse engineer: The reverse engineer creates alternative questions from the assistant’s responses

This process allows the framework to generate synthetic datasets that stress-test the assistant’s capabilities. The team began with a seed dataset of 19 questions and expanded to 190 synthetic questions through an iterative process. After rigorous filtering, 45 high-quality questions were selected for testing. Each question was run five times to identify the most common trajectory, ensuring reliability in the dataset.

MAG-V employs semantic similarity, graph edit distance, and argument overlap to verify trajectories. These features train machine learning models like k-Nearest Neighbors (k-NN), Support Vector Machines (SVM), and Random Forests. The framework succeeded in its evaluation, outperforming GPT-4o judge baselines by 11% accuracy and matching GPT-4’s performance in several metrics. For example, MAG-V’s k-NN model achieved an accuracy of 82.33% and demonstrated an F1 score of 71.73. The approach also showed cost-efficiency by coupling cheaper models like GPT-4o-mini with in-context learning samples, guiding them to perform at levels comparable to more expensive LLMs.

The MAG-V framework delivers results by addressing critical challenges in trajectory verification. Its deterministic nature ensures consistent outcomes, eliminating the variability associated with LLM-based approaches. By generating synthetic datasets, MAG-V reduces dependence on real customer data, addressing privacy concerns and data scarcity. The framework’s ability to verify trajectories using statistical and embedding-based features represents progress in AI system reliability. Also, MAG-V’s reliance on alternative questions for trajectory verification offers a robust method to test and validate the reasoning pathways of AI agents.

Several key takeaways from the research on MAG-V are as follows:

  1. MAG-V generated 190 synthetic questions from a seed dataset of 19, filtering them down to 45 high-quality queries. This process demonstrated the potential for scalable data creation to support AI testing and training.
  2. The framework’s deterministic methodology eliminates reliance on LLM-as-a-judge approaches, offering consistent and reproducible outcomes.
  3. Machine learning models trained using MAG-V’s features achieved accuracy improvements of up to 11% over GPT-4o baselines, showcasing the approach’s efficacy.
  4. By integrating in-context learning with cheaper LLMs like GPT-4o-mini, MAG-V provided a cost-effective alternative to high-end models without compromising performance.
  5. The framework is adaptable to various domains and demonstrates scalability by leveraging alternative questions to validate trajectories.

In conclusion, the MAG-V framework effectively addresses critical challenges in synthetic data generation and trajectory verification for AI systems. The framework offers a scalable, cost-effective, and deterministic solution by integrating multi-agent systems with classical machine learning models like k-NN, SVM, and Random Forests. MAG-V’s ability to generate high-quality synthetic datasets and verify trajectories with precision makes it deemed for deploying reliable AI applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 60k+ ML SubReddit.

🚨 [Must Subscribe]: Subscribe to our newsletter to get trending AI research and dev updates


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🧵🧵 [Download] Evaluation of Large Language Model Vulnerabilities Report (Promoted)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes
AI & Technology

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
AI & Technology

Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning

September 23, 2026
Next Post
Beam Ventures launches 0M fund to make Abu Dhabi into a global gaming hub

Beam Ventures launches $150M fund to make Abu Dhabi into a global gaming hub

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Frozen food giant shuts California factory, eliminating 260 jobs

Frozen food giant shuts California factory, eliminating 260 jobs

September 19, 2026
Robots cause sidewalk traffic jam after tech error

Robots cause sidewalk traffic jam after tech error

September 17, 2026
Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

Larry Ellison’s about-face on an Oracle stock sale sparks chatter in Silicon Valley, Hollywood 

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!