• bitcoinBitcoin(BTC)$84,497.000.63%
  • ethereumEthereum(ETH)$2,688.611.04%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$782.612.39%
  • rippleXRP(XRP)$1.531.65%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.092.39%
  • tronTRON(TRX)$0.3406700.54%
  • zcashZcash(ZEC)$1,535.36-0.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.041.31%
  • HyperliquidHyperliquid(HYPE)$93.690.56%
  • dogecoinDogecoin(DOGE)$0.0963214.25%
  • moneroMonero(XMR)$549.460.16%
  • whitebitWhiteBIT Coin(WBT)$84.580.32%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.935.46%
  • cardanoCardano(ADA)$0.2488874.43%
  • RainRain(RAIN)$0.012075-3.01%
  • leo-tokenLEO Token(LEO)$8.91-0.74%
  • stellarStellar(XLM)$0.2122654.44%
  • bitcoin-cashBitcoin Cash(BCH)$336.63-3.33%
  • nearNEAR Protocol(NEAR)$4.669.12%
  • uniswapUniswap(UNI)$9.281.74%
  • litecoinLitecoin(LTC)$74.1123.05%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.431.45%
  • daiDai(DAI)$1.000.02%
  • CantonCanton(CC)$0.1127594.57%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.015.64%
  • hedera-hashgraphHedera(HBAR)$0.0928452.85%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.27%
  • shiba-inuShiba Inu(SHIB)$0.0000063.50%
  • BittensorBittensor(TAO)$294.240.69%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0627801.43%
  • BitwayBitway(BTW)$1.089.76%
  • MemeCoreMemeCore(M)$1.241.87%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,275.32-0.28%
  • OndoOndo(ONDO)$0.5224.11%
  • okbOKB(OKB)$119.981.17%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.21%
  • aaveAave(AAVE)$145.063.77%
  • mantleMantle(MNT)$0.673.89%
  • EthenaEthena(ENA)$0.2190176.21%
  • polkadotPolkadot(DOT)$1.165.24%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Policy Learning with Large World Models: Advancing Multi-Task Reinforcement Learning Efficiency and Performance

July 7, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Policy Learning with Large World Models: Advancing Multi-Task Reinforcement Learning Efficiency and Performance
ShareShareShareShareShare

Reinforcement Learning (RL) excels at tackling individual tasks but struggles with multitasking, especially across different robotic forms. World models, which simulate environments, offer scalable solutions but often rely on inefficient, high-variance optimization methods. While large models trained on vast datasets have advanced generalizability in robotics, they typically need near-expert data and fail to adapt across diverse morphologies. RL can learn from suboptimal data, making it promising for multitask settings. However, methods like zeroth-order planning in world models face scalability issues and become less effective as model size increases, particularly in massive models like GAIA-1 and UniSim.

Researchers from Georgia Tech and UC San Diego have introduced Policy learning with large World Models (PWM), an innovative model-based reinforcement learning (MBRL) algorithm. PWM pretrains world models on offline data and uses them for first-order gradient policy learning, enabling it to solve tasks with up to 152 action dimensions. This approach outperforms existing methods by achieving up to 27% higher rewards without costly online planning. PWM emphasizes the utility of smooth, stable gradients over long horizons rather than mere accuracy. It demonstrates that efficient first-order optimization leads to better policies and faster training than traditional zeroth-order methods.

YOU MAY ALSO LIKE

Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP

AI Agents Fuel a New Cybersecurity Boom

RL splits into model-based and model-free approaches. Model-free methods like PPO and SAC dominate real-world applications and employ actor-critic architectures. SAC uses First-order Gradients (FoG) for policy learning, offering low variance but facing issues with objective discontinuities. Conversely, PPO relies on zeroth-order gradients, which are robust to discontinuities but prone to high variance and slower optimization. Recently, the focus in robotics has shifted to large multi-task models trained via behavior cloning. Examples include RT-1 and RT-2 for object manipulation. However, the potential of large models in RL still needs to be explored. MBRL methods like DreamerV3 and TD-MPC2 leverage large world models, but their scalability could be improved, particularly with the growing size of models like GAIA-1 and UniSim.

The study focuses on discrete-time, infinite-horizon RL scenarios represented by a Markov Decision Process (MDP) involving states, actions, dynamics, and rewards. RL aims to maximize cumulative discounted rewards through a policy. Commonly, this is tackled using actor-critic architectures, which approximate state values and optimize policies. In MBRL, additional components such as learned dynamics and reward models, often called world models, are used. These models can encode true states into latent representations. Leveraging these world models, PWM efficiently optimizes policies using FoG, reducing variance and improving sample efficiency even in complex environments.

In evaluating the proposed method, complex control tasks were tackled using the flex simulator, focusing on environments like Hopper, Ant, Anymal, Humanoid, and muscle-actuated Humanoid. Comparisons were made against SHAC, which uses ground truth models, and TD-MPC2, a model-free method that actively plans at inference time. Results showed that PWM achieved higher rewards and smoother optimization landscapes than SHAC and TD-MPC2. Further tests on 30 and 80 multi-task environments revealed PWM’s superior reward performance and faster inference time than TD-MPC2. Ablation studies highlighted PWM’s robustness to stiff contact models and higher sample efficiency, especially with better-trained world models.

The study introduced PWM as an approach in MBRL. PWM utilizes large multi-task world models as differentiable physics simulators, leveraging first-order gradients for efficient policy training. The evaluations highlighted PWM’s ability to outperform existing methods, including those with access to ground-truth simulation models like TD-MPC2. Despite its strengths, PWM relies heavily on extensive pre-existing data for world model training, limiting its applicability in low-data scenarios. Additionally, while PWM offers efficient policy training, it requires re-training for each new task, posing challenges for rapid adaptation. Future research could explore enhancements in world model training and extend PWM to image-based environments and real-world applications.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP
AI & Technology

Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP

September 24, 2026
AI Agents Fuel a New Cybersecurity Boom
AI & Technology

AI Agents Fuel a New Cybersecurity Boom

September 24, 2026
Bessemer: Anthropic Has Been Consistent on AI Safety
AI & Technology

Bessemer: Anthropic Has Been Consistent on AI Safety

September 24, 2026
Apple Explores Screenless Fitness Tracker to Rival Whoop
AI & Technology

Apple Explores Screenless Fitness Tracker to Rival Whoop

September 24, 2026
Next Post
Yara International – Down In The Interesting Doldrums (OTCMKTS:YARIY)

Yara International - Down In The Interesting Doldrums (OTCMKTS:YARIY)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Starbucks resolves Florida DEI lawsuit with blockbuster companywide agreement

Starbucks resolves Florida DEI lawsuit with blockbuster companywide agreement

September 17, 2026
UK to refuel Saudi jets to help counter Houthi attacks – Al Jazeera

UK to refuel Saudi jets to help counter Houthi attacks – Al Jazeera

September 22, 2026
Harbor Diversified International All Cap Fund Q2 2026 Commentary

Harbor Diversified International All Cap Fund Q2 2026 Commentary

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!