• bitcoinBitcoin(BTC)$83,458.00-1.37%
  • ethereumEthereum(ETH)$2,678.90-0.38%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$764.40-1.69%
  • rippleXRP(XRP)$1.49-2.47%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$118.72-3.37%
  • tronTRON(TRX)$0.3358490.62%
  • zcashZcash(ZEC)$1,472.75-8.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$88.03-4.00%
  • dogecoinDogecoin(DOGE)$0.094009-3.02%
  • chainlinkChainlink(LINK)$15.288.71%
  • moneroMonero(XMR)$540.96-1.21%
  • whitebitWhiteBIT Coin(WBT)$83.42-1.14%
  • USDSUSDS(USDS)$1.00-0.02%
  • cardanoCardano(ADA)$0.245583-3.63%
  • RainRain(RAIN)$0.012507-0.53%
  • leo-tokenLEO Token(LEO)$9.060.52%
  • stellarStellar(XLM)$0.2282355.57%
  • nearNEAR Protocol(NEAR)$4.81-11.52%
  • bitcoin-cashBitcoin Cash(BCH)$308.48-7.61%
  • uniswapUniswap(UNI)$8.82-9.54%
  • litecoinLitecoin(LTC)$69.34-2.47%
  • hedera-hashgraphHedera(HBAR)$0.11998826.71%
  • CantonCanton(CC)$0.129138-6.20%
  • avalanche-2Avalanche(AVAX)$10.44-4.88%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • suiSui(SUI)$1.16-8.49%
  • daiDai(DAI)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.61-1.22%
  • USD1USD1(USD1)$1.000.00%
  • quant-networkQuant(QNT)$242.4029.02%
  • BittensorBittensor(TAO)$307.34-5.69%
  • crypto-com-chainCronos(CRO)$0.0696163.05%
  • shiba-inuShiba Inu(SHIB)$0.000006-4.02%
  • tether-goldTether Gold(XAUT)$4,123.40-3.63%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • BitwayBitway(BTW)$1.02-15.58%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.16-2.91%
  • EthenaEthena(ENA)$0.260621-7.82%
  • OndoOndo(ONDO)$0.52-6.83%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$118.26-2.57%
  • Pump.funPump.fun(PUMP)$0.0051804.66%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • aaveAave(AAVE)$148.07-4.30%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

DeepMind Researchers Introduce AlphaStar Unplugged: A Leap Forward in Large-Scale Offline Reinforcement Learning by Mastering the Real-Time Strategy Game StarCraft II

August 15, 2023
in AI & Technology
Reading Time: 5 mins read
A A
DeepMind Researchers Introduce AlphaStar Unplugged: A Leap Forward in Large-Scale Offline Reinforcement Learning by Mastering the Real-Time Strategy Game StarCraft II
ShareShareShareShareShare

Games have long served as crucial testing grounds for evaluating the capabilities of artificial intelligence (AI) systems. As AI technologies have evolved, researchers have sought more complex games to assess various intelligence facets relevant to real-world challenges. StarCraft, a Real-Time Strategy (RTS) game, has emerged as a “grand challenge” for AI research due to its intricate gameplay, pushing the boundaries of AI techniques to navigate its complexity.

In contrast to earlier AI achievements in video games like Atari, Mario, Quake III Arena Capture the Flag, and Dota 2, which were based on online reinforcement learning (RL), often involved constraining game rules, providing superhuman abilities, or utilizing simplified maps, StarCraft’s complexity has proven a formidable obstacle for AI methods. However, these online reinforcement learning (RL) algorithms have succeeded significantly in this domain. Yet, their interactive nature poses challenges for real-world applications, demanding high interaction and exploration.

This research introduces a transformative shift towards offline RL, allowing agents to learn from fixed datasets – a more practical and safer approach. While online RL excels in interactive domains, offline RL harnesses existing data to create deployment-ready policies. The introduction of the AlphaStar program by DeepMind researchers marked a significant milestone by becoming the first AI to defeat a top professional StarCraft player. AlphaStar has mastered StarCraft II’s gameplay, using a deep neural network trained through supervised learning and reinforcement learning on raw game data.

Leveraging an expansive dataset of human player replays from StarCraft II; this framework enables agent training and evaluation without requiring direct environment interaction. StarCraft II, with its distinctive challenges such as partial observability, stochasticity, and multi-agent dynamics, makes for an ideal testing ground to push the boundaries of offline RL algorithm capabilities. “AlphaStar Unplugged” establishes a benchmark tailored to intricate, partially observable games like StarCraft II by bridging the gap between traditional online RL methods and offline RL.

The core methodology of “AlphaStar Unplugged” revolves around several key contributions that establish this challenging offline RL benchmark:

  1. The training setup employed a fixed dataset and defined rules to ensure fair comparisons between methods.
  2. A novel set of evaluation metrics is introduced to measure agent performance accurately.
  3. A range of well-tuned baseline agents is provided as starting points for experimentation.
  4. Recognizing the considerable engineering effort required to build effective agents for StarCraft II, the researchers furnish a well-tuned behavior cloning agent that forms the foundation for all agents detailed in the paper.

The “AlphaStar Unplugged” architecture involves several reference agents for baseline comparisons and metric evaluations. Inputs to the StarCraft II API are structured around three modalities: vectors, units, and feature planes. The actions consist of seven modalities: function, delay, queued, repeat, unit tags, target unit tag, and world action. Multi-layer perceptrons (MLP) encode and process vector inputs, transformers handle unit inputs, and residual convolutional networks manage feature planes. Modalities are interconnected through unit scattering, vector embedding, convolutional reshaping, and memory usage. Memory is incorporated into the vector modality, and a value function is employed alongside action sampling.

The experimental results underscore the remarkable achievement of offline RL algorithms, demonstrating a 90% win rate against the previously leading AlphaStar Supervised agent. Notably, this performance is achieved solely through the utilization of offline data. The researchers envision their work will significantly advance large-scale offline reinforcement learning research.

The matrix shows normalized win rates of reference agents, scaled between 0 and 100. Note that draws can affect totals, and AS-SUP represents the original AlphaStar Supervised agent.

In conclusion, DeepMind’s “AlphaStar Unplugged” introduces an unprecedented benchmark that pushes the boundaries of offline reinforcement learning. By harnessing the intricate game dynamics of StarCraft II, this benchmark sets the stage for improved training methodologies and performance metrics in the realm of RL research. Furthermore, it highlights the promise of offline RL in bridging the gap between simulated and real-world applications, presenting a safer and more practical approach to training RL agents for complex environments.


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 28k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Meta Bets on AI, Devices for Its Next Chapter

AI Emerges as a Flashpoint in Trump-Xi Talks

Madhur Garg is a consulting intern at MarktechPost. He is currently pursuing his B.Tech in Civil and Environmental Engineering from the Indian Institute of Technology (IIT), Patna. He shares a strong passion for Machine Learning and enjoys exploring the latest advancements in technologies and their practical applications. With a keen interest in artificial intelligence and its diverse applications, Madhur is determined to contribute to the field of Data Science and leverage its potential impact in various industries.


🔥 Use SQL to predict the future (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Bets on AI, Devices for Its Next Chapter
AI & Technology

Meta Bets on AI, Devices for Its Next Chapter

September 28, 2026
AI Emerges as a Flashpoint in Trump-Xi Talks
AI & Technology

AI Emerges as a Flashpoint in Trump-Xi Talks

September 28, 2026
Darktrace CEO: AI Agents Are the New ‘Insider Threat’
AI & Technology

Darktrace CEO: AI Agents Are the New ‘Insider Threat’

September 28, 2026
Meta Is Starting An Enterprise Business To Justify Its Massive AI Spending
AI & Technology

Meta Is Starting An Enterprise Business To Justify Its Massive AI Spending

September 28, 2026
Next Post
Salesforce: Price Momentum Likely To Top Out On Valuation Concerns (NYSE:CRM)

Salesforce: Price Momentum Likely To Top Out On Valuation Concerns (NYSE:CRM)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
‘Nothing has gotten better’: Focus group shares affordability concerns as US debt tops  trillion

‘Nothing has gotten better’: Focus group shares affordability concerns as US debt tops $40 trillion

September 26, 2026
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

September 28, 2026
Apple customers can now submit claims for 0M settlement in deceptive marketing suit

Apple customers can now submit claims for $250M settlement in deceptive marketing suit

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!