• bitcoinBitcoin(BTC)$76,695.00-0.80%
  • ethereumEthereum(ETH)$2,475.77-2.34%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$715.14-2.87%
  • rippleXRP(XRP)$1.34-2.38%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.56-2.59%
  • tronTRON(TRX)$0.340691-0.09%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,088.90-6.20%
  • HyperliquidHyperliquid(HYPE)$77.41-3.60%
  • dogecoinDogecoin(DOGE)$0.083300-2.08%
  • RainRain(RAIN)$0.0152610.83%
  • moneroMonero(XMR)$530.510.38%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$79.53-1.05%
  • chainlinkChainlink(LINK)$11.25-2.65%
  • leo-tokenLEO Token(LEO)$9.05-0.64%
  • cardanoCardano(ADA)$0.204309-2.10%
  • stellarStellar(XLM)$0.178157-1.96%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.48-3.34%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.49-0.94%
  • uniswapUniswap(UNI)$6.26-2.96%
  • CantonCanton(CC)$0.095115-3.40%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-2.37%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0751820.64%
  • avalanche-2Avalanche(AVAX)$7.30-1.97%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.60%
  • nearNEAR Protocol(NEAR)$2.31-2.53%
  • suiSui(SUI)$0.71-2.70%
  • crypto-com-chainCronos(CRO)$0.0585840.45%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.32-0.02%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-2.65%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.72-1.24%
  • BittensorBittensor(TAO)$233.78-1.13%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • aaveAave(AAVE)$125.36-1.35%
  • pax-goldPAX Gold(PAXG)$4,350.50-0.04%
  • AsterAster(ASTER)$0.691.00%
  • mantleMantle(MNT)$0.56-1.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575070.97%
  • polkadotPolkadot(DOT)$1.01-3.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from CMU and Peking Introduces ‘DiffTOP’ that Uses Differentiable Trajectory Optimization to Generate the Policy Actions for Deep Reinforcement Learning and Imitation Learning

February 24, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from CMU and Peking Introduces ‘DiffTOP’ that Uses Differentiable Trajectory Optimization to Generate the Policy Actions for Deep Reinforcement Learning and Imitation Learning
ShareShareShareShareShare

According to recent studies, a policy’s depiction can significantly affect learning performance. Policy representations such as feed-forward neural networks, energy-based models, and diffusion have all been investigated in earlier research.

A recent study by Carnegie Mellon University and Peking University researchers proposes producing actions for deep reinforcement and imitation learning using high-dimensional sensory data (images/point clouds) and differentiable trajectory optimization as the policy representation. A cost function and a dynamics function are typically used to define trajectory optimization, a popular and successful control approach. Consider it a policy whose parameters define the cost function and the dynamics function, in this case represented by neural networks.

After receiving the input state (such as pictures, point clouds, or robot joint states) and the learned cost and dynamics functions, the policy will solve the trajectory optimization problem to determine the actions to take. It is also possible to make trajectory optimization differentiable, which opens the door to back-propagation inside the optimization process. Problems with low-dimensional states in robotics, imitation learning, system identification, and inverse optimal control have all been addressed in earlier work using differentiable trajectory optimization. 

This is the first demonstration of a hybrid approach that combines deep model-based RL algorithms with differentiable trajectory optimization. The team learns the dynamics and cost functions to optimize the reward by computing the policy gradient loss on the generated actions, which is made possible by using differentiable trajectory optimization for action generation. 

Models that perform better during training (e.g., with a lower mean squared error) when learning a dynamics model are not always better when it comes to control, and this is the “objective mismatch” problem that this method seeks to solve in present model-based RL algorithms. In order to solve this problem, they developed DiffTOP, which stands for “Differentiable Trajectory Optimization.” By optimizing the trajectory, they maximize task performance by back-propagating the policy gradient loss, which optimizes both the latent dynamics and the reward models. 

The comprehensive experiments demonstrate that DiffTOP outperforms previous state-of-the-art methods in both model-based RL (15 tasks) and imitation learning (13 tasks) using standard benchmarking with high-dimensional sensory observations. These tasks included  5 Robomimic tasks using images as inputs and 9 Maniskill1 and Maniskill2 challenges using point clouds as inputs.

The team also compares their approach to feed-forward policy classes, Energy-Based Models (EBM), and Diffusion and evaluates DiffTOP for imitation learning on common robotic manipulation task suites using high-dimensional sensory data. Compared to the EBM approach utilized in previous work, which can experience training instability because it requires sampling high-quality negative samples, their training procedure using differentiable trajectory optimization leads to improved performance. The proposed method of learning and optimizing a cost function during testing allows us to outperform diffusion-based alternatives as well. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 37k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Why Do Routers Have So Many Antennas?
AI & Technology

Why Do Routers Have So Many Antennas?

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
Next Post
Prologis Stock: Too Expensive At Current Interest Rates (NYSE:PLD)

Prologis Stock: Too Expensive At Current Interest Rates (NYSE:PLD)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Column | Schools should stop banning peanuts. Here’s what really keeps kids safe. – The Washington Post

Column | Schools should stop banning peanuts. Here’s what really keeps kids safe. – The Washington Post

September 8, 2026
L.L. Bean CEO on how he manages his day to day

L.L. Bean CEO on how he manages his day to day

September 12, 2026
How Long Should A Laptop Battery Last On A Single Charge?

How Long Should A Laptop Battery Last On A Single Charge?

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!