• bitcoinBitcoin(BTC)$85,949.002.29%
  • ethereumEthereum(ETH)$2,739.751.49%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$785.320.53%
  • rippleXRP(XRP)$1.533.93%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.981.47%
  • tronTRON(TRX)$0.3485041.25%
  • zcashZcash(ZEC)$1,496.09-1.13%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • HyperliquidHyperliquid(HYPE)$95.13-0.68%
  • dogecoinDogecoin(DOGE)$0.0985777.16%
  • moneroMonero(XMR)$567.46-2.14%
  • whitebitWhiteBIT Coin(WBT)$86.421.32%
  • chainlinkChainlink(LINK)$12.900.45%
  • RainRain(RAIN)$0.013597-3.47%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2476494.49%
  • leo-tokenLEO Token(LEO)$8.980.52%
  • stellarStellar(XLM)$0.2117502.17%
  • nearNEAR Protocol(NEAR)$4.505.42%
  • uniswapUniswap(UNI)$8.76-3.00%
  • bitcoin-cashBitcoin Cash(BCH)$267.432.43%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.84-2.77%
  • litecoinLitecoin(LTC)$60.421.84%
  • CantonCanton(CC)$0.1173161.05%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0961578.71%
  • suiSui(SUI)$1.021.08%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.61%
  • BittensorBittensor(TAO)$319.8214.87%
  • shiba-inuShiba Inu(SHIB)$0.0000066.22%
  • crypto-com-chainCronos(CRO)$0.0659084.60%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.33-13.70%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,322.55-0.75%
  • okbOKB(OKB)$122.911.26%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.31%
  • BitwayBitway(BTW)$0.821.99%
  • aaveAave(AAVE)$141.44-2.21%
  • mantleMantle(MNT)$0.653.09%
  • EthenaEthena(ENA)$0.210002-5.15%
  • Pump.funPump.fun(PUMP)$0.0045211.91%
  • OndoOndo(ONDO)$0.432723-1.68%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

DigiRL: A Novel Autonomous Reinforcement Learning RL Method to Train Device-Control Agents

June 23, 2024
in AI & Technology
Reading Time: 4 mins read
A A
DigiRL: A Novel Autonomous Reinforcement Learning RL Method to Train Device-Control Agents
ShareShareShareShareShare

Advances in vision-language models (VLMs) have shown impressive common sense, reasoning, and generalization abilities. This means that developing a fully independent digital AI assistant, that can perform daily computer tasks through natural language is possible. However, better reasoning and common-sense abilities don’t automatically lead to intelligent assistant behavior. AI assistants are used to complete tasks, behave rationally, and recover from mistakes, not just provide plausible responses based on pre-training data. So, a method is required to turn pre-training abilities into practical AI “agents.” Even the best VLMs, like GPT-4V and Gemini 1.5 Pro, still struggle to perform the right actions when completing device tasks.  

This paper discusses three existing methods. The first method is training multi-modal digital agents, which face challenges like device control being done directly at the pixel level in a coordinate-based action space, and the stochastic and unpredictable nature of device ecosystems and the internet. The second method is Environments for device control agents. These environments are designed for evaluation, and offer a limited range of tasks in fully deterministic and stationary settings. The last method is Reinforcement learning (RL) for LLM/VLMs, where research with RL for foundation models focuses on single-turn tasks like preference optimization, but optimizing for single-turn interaction from expert demonstrations can lead to sub-optimal strategies for multi-step problems.

YOU MAY ALSO LIKE

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

Researchers from UC Berkeley, UIUC, and Google DeepMind have introduced DigiRL (RL for Digital Agents), a novel autonomous RL method for training device control agents. The resulting agent attains state-of-the-art performance on several Android device-control tasks. The training process involves two phases: first, an initial offline RL phase to initialize the agent using existing data, followed by an offline-to-online RL phase, that is used for fine-tuning the model obtained from offline RL on online data. To train online RL a scalable and parallelizable Android learning environment was developed that includes a robust general-purpose evaluator (average error rate 2.8% against human judgment) based on VLM.

Researchers carried out experiments to evaluate the performance of DigiRL on challenging Android device control problems. It is important to understand if DigiRL has the potential to produce agents that can learn effectively through autonomous interaction, while still being able to utilize offline data for learning. So, a comparative analysis was performed on DigiRL against the following:

  • State-of-the-art agents built around proprietary VLMs using several prompting and retrieval-style techniques. 
  • Running imitation learning on static human demonstrations with the same instruction distribution
  • A filtered Behavior Cloning approach.

An agent trained using DigiRL was tested on various tasks from the Android in the Wild dataset (AitW) with real Android device emulators. The agent achieved a 28.7% improvement over the existing state-of-the-art agents (raising the success rate from 38.5% to 67.2%) 18B CogAgent. It also outperformed the previous top autonomous learning method based on Filtered Behavior Cloning by more than 9%. Moreover, despite having only 1.3B parameters, the agent performed better than advanced models like GPT-4V and Gemini 1.5 Pro (17.7% success rate). This makes it the first agent to achieve state-of-the-art performance in device control using an autonomous offline-to-online RL approach.

In summary, researchers proposed DigiRL, a novel autonomous RL approach for training device-control agents that sets a new state-of-the-art performance on several Android control tasks from AitW. A scalable and parallelizable Android environment was developed to achieve this with a robust VLM-based general-purpose evaluator for quick online data collection. The agent trained on DigiRL achieved a 28.7% improvement over the existing state-of-the-art agents 18B CogAgent. However, the training was limited to tasks from the AitW dataset instead of all possible device tasks. So, future work includes building algorithmic research and expanding the task space, making DigiRL the base algorithm. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.

[Announcing Gretel Navigator] Create, edit, and augment tabular data with the first compound AI system trusted by EY, Databricks, Google, and Microsoft


Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Next Post
SoftBank’s Son Plans to Create ‘Super’ AI

SoftBank’s Son Plans to Create ‘Super’ AI

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
What Is Prompt Injection? The Security Flaw Every AI User Should Understand – Unite.AI

What Is Prompt Injection? The Security Flaw Every AI User Should Understand – Unite.AI

September 20, 2026
Dutch Pension Funds No Longer A Shock Absorber Of Higher Rates

Dutch Pension Funds No Longer A Shock Absorber Of Higher Rates

September 21, 2026
Driver speeds by school bus as children start to cross

Driver speeds by school bus as children start to cross

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!