• bitcoinBitcoin(BTC)$81,219.004.35%
  • ethereumEthereum(ETH)$2,634.845.74%
  • tetherTether(USDT)$1.000.05%
  • binancecoinBNB(BNB)$762.661.36%
  • rippleXRP(XRP)$1.426.46%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$111.965.82%
  • tronTRON(TRX)$0.3378940.44%
  • zcashZcash(ZEC)$1,565.755.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.22%
  • HyperliquidHyperliquid(HYPE)$92.895.42%
  • dogecoinDogecoin(DOGE)$0.0872193.42%
  • moneroMonero(XMR)$575.096.85%
  • RainRain(RAIN)$0.0139847.95%
  • whitebitWhiteBIT Coin(WBT)$83.283.88%
  • USDSUSDS(USDS)$1.000.02%
  • chainlinkChainlink(LINK)$12.414.92%
  • cardanoCardano(ADA)$0.2233694.33%
  • leo-tokenLEO Token(LEO)$8.89-0.15%
  • stellarStellar(XLM)$0.1931863.07%
  • uniswapUniswap(UNI)$9.204.65%
  • bitcoin-cashBitcoin Cash(BCH)$247.88-0.28%
  • nearNEAR Protocol(NEAR)$3.696.45%
  • Ethena USDeEthena USDe(USDE)$1.000.05%
  • daiDai(DAI)$1.000.02%
  • litecoinLitecoin(LTC)$57.163.23%
  • CantonCanton(CC)$0.110102-1.54%
  • USD1USD1(USD1)$1.000.06%
  • avalanche-2Avalanche(AVAX)$8.679.51%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.48%
  • hedera-hashgraphHedera(HBAR)$0.0789443.28%
  • suiSui(SUI)$0.824.64%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000051.44%
  • MemeCoreMemeCore(M)$1.311.70%
  • crypto-com-chainCronos(CRO)$0.0593240.51%
  • BittensorBittensor(TAO)$256.824.19%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,370.96-0.35%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.862.71%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.23%
  • aaveAave(AAVE)$144.056.93%
  • AsterAster(ASTER)$0.760.91%
  • mantleMantle(MNT)$0.612.43%
  • OndoOndo(ONDO)$0.3987353.91%
  • Pump.funPump.fun(PUMP)$0.004140-2.81%
  • MorphoMorpho(MORPHO)$2.7314.97%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Zhipu AI Unveils ComputerRL: An AI Framework Scaling End-to-End Reinforcement Learning for Computer Use Agents

August 22, 2025
in AI & Technology
Reading Time: 11 mins read
A A
Zhipu AI Unveils ComputerRL: An AI Framework Scaling End-to-End Reinforcement Learning for Computer Use Agents
ShareShareShareShareShare

In the rapidly evolving landscape of AI-driven automation, Zhipu AI has introduced ComputerRL, a groundbreaking framework designed to empower agents with the ability to navigate and manipulate complex digital workspaces. This innovation addresses a core challenge in AI agent development: the disconnect between computer agents and human-designed graphical user interfaces (GUIs). By integrating programmatic API calls with direct GUI interactions, ComputerRL enables more efficient and versatile desktop operations, marking a significant step toward autonomous computer use agents.

Image source: https://arxiv.org/abs/2508.14040

The API-GUI Paradigm: Bridging Human and Machine Interactions

Traditional GUI agents often struggle with environments optimized for human users, leading to inefficient simulations of actions like clicking or scrolling. ComputerRL introduces the API-GUI paradigm, which combines the precision of API invocations with the flexibility of GUI-based operations. This hybrid approach allows agents to leverage machine-friendly APIs for tasks that benefit from programmatic control, while falling back on GUI actions for broader adaptability.

YOU MAY ALSO LIKE

GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI

The framework automates API construction using large language models (LLMs). Users provide example tasks, and the system analyzes requirements, implements APIs using relevant Python libraries, and generates test cases. This process ensures APIs encapsulate general-purpose functionalities, reducing complexity and enhancing agent performance. For instance, APIs for Ubuntu applications like GIMP and LibreOffice are integrated, enabling tasks such as image processing or document formatting with fewer steps than GUI-only methods.

Scalable Infrastructure for Large-Scale RL Training

A major hurdle in training desktop agents is the inefficiency of virtual environments. ComputerRL overcomes this with a distributed reinforcement learning (RL) infrastructure built on Docker and gRPC, supporting thousands of parallel Ubuntu virtual machines. This setup is compatible with benchmarks like AgentBench and addresses issues in prior systems, such as resource intensiveness and network bottlenecks.

Key features include lightweight VM deployment via qemu-in-docker, multi-node clustering for scalability, and a web-based monitoring interface. Paired with the AgentRL framework, it enables fully asynchronous training, decoupling data collection from parameter updates to boost efficiency. This infrastructure allows for high-throughput RL, with dynamic batch sizing and off-policy bias mitigation, facilitating extended training runs without stagnation.

Image source: https://arxiv.org/abs/2508.14040

Entropulse: Enhancing RL with Alternating Training Phases

To tackle entropy collapse—a common issue where agents lose exploratory behavior during prolonged RL—ComputerRL incorporates Entropulse. This method alternates RL phases with supervised fine-tuning (SFT) on successful rollout trajectories, restoring entropy and enabling sustained performance gains.

The training pipeline begins with behavior cloning (BC) using trajectories from multiple LLMs for diversity. It then applies step-level Group Relative Policy Optimization (GRPO) with rule-based rewards, assigning positive scores only to correct, contributing actions in successful trajectories. Entropulse intervenes by curating diverse, high-quality data from prior rollouts for SFT, preventing premature convergence and scaling effective training steps.

Image source: https://arxiv.org/abs/2508.14040

Experimental Validation on OSWorld Benchmark

The research team applied ComputerRL to open-source models like GLM-4-9B-0414 and Qwen2.5-14B, resulting in AutoGLM-OS variants. On the OSWorld benchmark, which evaluates agents in interactive Ubuntu environments, AutoGLM-OS-9B achieved a success rate of 48.1%, surpassing proprietary models like OpenAI’s CUA o3 (42.9%) and Claude 4.0 (30.7%). It also excelled on OSWorld-Verified, scoring 47.3%.

Ablation studies highlight the framework’s strengths. The API-GUI paradigm improved success rates by 134% over GUI-only baselines, particularly in office and professional domains. Training ablations showed BC providing a 31.9% baseline, with RL phases adding up to 45.8% through Entropulse-enabled exploration. Entropy curves confirmed Entropulse’s role in maintaining learning momentum.

Case studies demonstrate practical efficacy, such as creating sales summary tables in LibreOffice Calc or generating system reports via Terminal commands. However, error analysis revealed challenges like visual perception issues (25.8% of failures) and multi-app coordination (34.4%), pointing to areas for refinement.

Image source: https://arxiv.org/abs/2508.14040

Future Directions in Desktop Autonomy

Looking ahead, ComputerRL sets the stage for more robust agents capable of handling dynamic environments and long-horizon tasks. Potential advancements include expanding training diversity, integrating multimodal perception, and developing hierarchical planning. Safety features like permission frameworks and action validation will be crucial for real-world deployment, ensuring aligned and trustworthy automation.

ComputerRL represents a pivotal advancement in AI agents, blending scalable RL with innovative interaction paradigms to transform desktop intelligence. As open models like AutoGLM-OS push boundaries, this framework paves the way for more capable, general-purpose agents in everyday computing.


Check out the Technical paper here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
AI & Technology

GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)

September 19, 2026
Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI
AI & Technology

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI

September 19, 2026
How Focus Mode Has Changed In iOS 27
AI & Technology

How Focus Mode Has Changed In iOS 27

September 18, 2026
AI Almost Led The US Military To Start A War With China, Report Says
AI & Technology

AI Almost Led The US Military To Start A War With China, Report Says

September 18, 2026
Next Post
Gaza father says he sought treatment in hospital after not eating for a week

Gaza father says he sought treatment in hospital after not eating for a week

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Virginia mom convicted of misdemeanor after son walks alone

Virginia mom convicted of misdemeanor after son walks alone

September 15, 2026
At least 21 killed after war-damaged Gaza building collapses – BBC

At least 21 killed after war-damaged Gaza building collapses – BBC

September 16, 2026
Retiring mailman gets heartfelt farewell from residents

Retiring mailman gets heartfelt farewell from residents

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!