• bitcoinBitcoin(BTC)$85,373.004.87%
  • ethereumEthereum(ETH)$2,730.902.44%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$789.061.43%
  • rippleXRP(XRP)$1.516.30%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$116.484.28%
  • tronTRON(TRX)$0.3473311.28%
  • zcashZcash(ZEC)$1,452.77-4.66%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.27%
  • HyperliquidHyperliquid(HYPE)$92.27-0.88%
  • dogecoinDogecoin(DOGE)$0.09924411.88%
  • moneroMonero(XMR)$573.99-1.02%
  • whitebitWhiteBIT Coin(WBT)$85.913.24%
  • RainRain(RAIN)$0.013824-2.59%
  • chainlinkChainlink(LINK)$12.881.77%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2454285.01%
  • leo-tokenLEO Token(LEO)$8.960.21%
  • stellarStellar(XLM)$0.2123026.39%
  • nearNEAR Protocol(NEAR)$4.401.33%
  • uniswapUniswap(UNI)$9.022.85%
  • bitcoin-cashBitcoin Cash(BCH)$263.893.46%
  • avalanche-2Avalanche(AVAX)$11.14-0.57%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • litecoinLitecoin(LTC)$60.341.95%
  • CantonCanton(CC)$0.1169484.71%
  • daiDai(DAI)$1.000.03%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.049.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.443.67%
  • hedera-hashgraphHedera(HBAR)$0.0916405.33%
  • BittensorBittensor(TAO)$313.3316.58%
  • shiba-inuShiba Inu(SHIB)$0.0000067.34%
  • crypto-com-chainCronos(CRO)$0.0662647.46%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.43-5.53%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,345.50-0.39%
  • okbOKB(OKB)$121.901.59%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BitwayBitway(BTW)$0.8311.51%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
  • aaveAave(AAVE)$141.702.11%
  • pepePepe(PEPE)$0.00000527.94%
  • mantleMantle(MNT)$0.644.76%
  • OndoOndo(ONDO)$0.4352110.85%
  • EthenaEthena(ENA)$0.209540-2.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

AgentA/B: A Scalable AI System Using LLM Agents that Simulate Real User Behavior to Transform Traditional A/B Testing on Live Web Platforms

April 26, 2025
in AI & Technology
Reading Time: 6 mins read
A A
AgentA/B: A Scalable AI System Using LLM Agents that Simulate Real User Behavior to Transform Traditional A/B Testing on Live Web Platforms
ShareShareShareShareShare

Designing and evaluating web interfaces is one of the most critical tasks in today’s digital-first world. Every change in layout, element positioning, or navigation logic can influence how users interact with websites. This becomes even more crucial for platforms that rely on extensive user engagement, such as e-commerce or content streaming services. One of the most trusted methods for assessing the impact of design changes is A/B testing. In A/B testing, two or more versions of a webpage are shown to different user groups to measure their behavior and determine which variant performs better. It’s not just about aesthetics but also functional usability. This method enables product teams to gather user-centered evidence before fully rolling out a feature, allowing businesses to optimize user interfaces systematically based on observed interactions.

Despite being a widely accepted tool, the traditional A/B testing process brings several inefficiencies that have proven problematic for many teams. The most significant challenge is the volume of real-user traffic needed to yield statistically valid results. In some scenarios, hundreds of thousands of users must interact with webpage variants to identify meaningful patterns. For smaller websites or early-stage features, securing this level of user interaction can be nearly impossible. The feedback cycle is also notably slow. Even after launching an experiment, it might take weeks to months before results can be confidently assessed due to the requirement of long observation periods. Also, these tests are resource-heavy; only a few variants can be evaluated due to the time and manpower required. Consequently, numerous promising ideas go untested because there’s simply no capacity to explore them all.

YOU MAY ALSO LIKE

Why It’s Important To Unplug Your PC During A Power Outage

Why Is Your Laptop Fan So Loud?

Several methods have been explored to overcome these limitations; however, each has its shortcomings. For example, offline A/B testing techniques depend on rich historical interaction logs, which are not always available or reliable. Tools that enable prototyping and experimentation, such as Apparition and Fuse, have accelerated early design exploration but are primarily useful for prototyping physical interfaces. Algorithms that reframe A/B testing as a search problem through evolutionary models help automate some aspects but still depend on historical or real-user deployment data. Other strategies, like cognitive modeling with GOMS or ACT-R frameworks, require high levels of manual configuration and do not easily adapt to the complexities of dynamic web behavior. These tools, although innovative, have not provided the scalability and automation necessary to address the deeper structural limitations in A/B testing workflows.

Researchers from Northeastern University, Pennsylvania State University, and Amazon introduced a new automated system named AgentA/B. This system offers an alternative approach to traditional user testing, utilizing Large Language Model (LLM)-based agents. Rather than depending on live user interaction, AgentA/B simulates human behavior using thousands of AI agents. These agents are assigned detailed personas that mimic characteristics such as age, educational background, technical proficiency, and shopping preferences. These personas enable agents to simulate a wide range of user interactions on real websites. The goal is to provide researchers and product managers with an efficient and scalable method for testing multiple design variants without relying on live user feedback or extensive traffic coordination.

The system architecture of AgentA/B is structured into four main components. First, it generates agent personas based on the input demographics and behavioral diversity specified by the user. These personas are fed into the second stage, where testing scenarios are defined—this includes assigning agents to control and treatment groups and specifying which two webpage versions should be tested. The third component executes the interactions: agents are deployed into real browser environments, where they process the content through structured web data (converted into JSON observations) and take action like real users. They can search, filter, click, and even simulate purchases. The fourth and final component involves analyzing the results, where the system provides metrics like the number of clicks, purchases, or interaction durations to assess design effectiveness.

During their testing phase, researchers used Amazon.com to demonstrate the tool’s practical value. A total of 100,000 virtual customer personas were generated, and 1,000 were randomly selected from this pool to act as LLM agents in the simulation. The experiment compared two different webpage layouts: one with all product filter options shown in a left-hand panel and another with only a reduced set of filters. The outcome was compelling. The agents interacting with the reduced-filter version performed more purchases and filter-based actions than those with the full list. Also, these virtual agents were significantly more efficient. Compared with one million real user interactions, LLM agents took fewer actions on average to complete tasks, indicating more goal-oriented behavior. These results mirrored the behavioral direction observed in human A/B tests, strengthening the case for AgentA/B as a valid complement to traditional testing.

This research demonstrates a compelling advancement in interface evaluation. It doesn’t aim to replace live user A/B testing but instead proposes a supplementary method that offers rapid feedback, cost efficiency, and broader experimental coverage. By using AI agents instead of live participants, the system enables product teams to test numerous interface variations that would otherwise be infeasible. This model can significantly compress the design cycle, allowing ideas to be validated or rejected at a much earlier stage. It addresses the practical concerns of long wait times, traffic limitations, and testing resource constraints, making the web design process more data-informed and less prone to bottlenecks.

Some Key Takeaways from the Research on AgentA/B include:

  • AgentA/B uses LLM-based agents to simulate realistic user behavior on live webpages.  
  • The system allows automated A/B testing with no need for live user deployment.  
  • 100,000 user personas were generated, and 1,000 were selected for live testing simulation.  
  • The system compared two webpage variants on Amazon.com: full filter panel vs. reduced filters.  
  • LLM agents in the reduced-filter group made more purchases and performed more filtering actions.  
  • Compared to 1 million human users, LLM agents showed shorter action sequences and more goal-directed behavior.  
  • AgentA/B can help evaluate interface changes before real user testing, saving months of development time.  
  • The system is modular and extensible, allowing it to be adaptable to various web platforms and testing goals.  
  • It directly addresses three core A/B testing challenges: long cycles, high user traffic needs, and experiment failure rates.

Check out the Paper. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 90k+ ML SubReddit.

🔥 [Register Now] miniCON Virtual Conference on AGENTIC AI: FREE REGISTRATION + Certificate of Attendance + 4 Hour Short Event (May 21, 9 am- 1 pm PST) + Hands on Workshop


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Next Post
Meta AI Introduces Token-Shuffle: A Simple AI Approach to Reducing Image Tokens in Transformers

Meta AI Introduces Token-Shuffle: A Simple AI Approach to Reducing Image Tokens in Transformers

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Woman narrowly escapes Russian drone strike near Kyiv

Woman narrowly escapes Russian drone strike near Kyiv

September 21, 2026
No charges for woman in fatal Walmart parking shooting

No charges for woman in fatal Walmart parking shooting

September 17, 2026
Hundreds of teens swarm Colorado shopping center for so-called ‘takeover’

Hundreds of teens swarm Colorado shopping center for so-called ‘takeover’

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!