• bitcoinBitcoin(BTC)$76,065.000.39%
  • ethereumEthereum(ETH)$2,415.270.78%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$722.161.57%
  • rippleXRP(XRP)$1.290.54%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$98.841.90%
  • tronTRON(TRX)$0.3357110.91%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.41%
  • zcashZcash(ZEC)$1,362.6922.43%
  • HyperliquidHyperliquid(HYPE)$78.571.69%
  • dogecoinDogecoin(DOGE)$0.0805510.80%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$500.87-1.24%
  • whitebitWhiteBIT Coin(WBT)$78.180.34%
  • RainRain(RAIN)$0.012903-8.35%
  • chainlinkChainlink(LINK)$11.041.80%
  • leo-tokenLEO Token(LEO)$8.951.29%
  • cardanoCardano(ADA)$0.1944040.42%
  • stellarStellar(XLM)$0.1829254.32%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$219.610.68%
  • USD1USD1(USD1)$1.000.00%
  • uniswapUniswap(UNI)$6.625.14%
  • litecoinLitecoin(LTC)$51.832.00%
  • CantonCanton(CC)$0.0962315.67%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-0.11%
  • nearNEAR Protocol(NEAR)$2.5911.75%
  • avalanche-2Avalanche(AVAX)$7.483.23%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.073595-1.02%
  • suiSui(SUI)$0.714.14%
  • shiba-inuShiba Inu(SHIB)$0.0000050.95%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0564511.86%
  • tether-goldTether Gold(XAUT)$4,291.96-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.11-1.27%
  • BittensorBittensor(TAO)$221.922.28%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$110.73-0.55%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.29%
  • BitwayBitway(BTW)$0.736.42%
  • AsterAster(ASTER)$0.715.27%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0590483.51%
  • pax-goldPAX Gold(PAXG)$4,292.07-0.07%
  • aaveAave(AAVE)$119.93-0.57%
  • mantleMantle(MNT)$0.551.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

TD3-BST: A Machine Learning Algorithm to Adjust the Strength of Regularization Dynamically Using Uncertainty Model

April 28, 2024
in AI & Technology
Reading Time: 4 mins read
A A
TD3-BST: A Machine Learning Algorithm to Adjust the Strength of Regularization Dynamically Using Uncertainty Model
ShareShareShareShareShare

Reinforcement learning (RL) is a type of learning approach where an agent interacts with an environment to collect experiences and aims to maximize the reward received from the environment. This usually involves a looping process of experience collecting and enhancement, and due to the requirement of policy rollouts, it is called online RL. Both on-policy and off-policy RL need online interaction, which can be impractical in certain domains due to experimental or environmental constraints. Offline RL algorithms are framed so that they can extract optimal policies from static datasets.

Offline RL algorithms are used to learn effective and well-applicable policies with the help of static datasets. Many approaches to this algorithm have achieved major success recently. However, they demand significant hyperparameter tuning specific to each dataset to achieve reported performance, which needs policy rollouts in the environment to evaluate. This can create a major problem because the need for significant tuning can affect the adoption of these algorithms in practical domains. Offline RL faces challenges during the evaluation of out-of-distribution (OOD) actions.

Researchers from Imperial College London introduced TD3-BST (TD3 with Behavioral Supervisor Tuning), an algorithm that uses an uncertainty model to adjust the strength of regularization dynamically. The trained uncertainty model is incorporated into the regularized policy yield TD3 with behavioral supervisor tuning (TD3-BST). TD3-BST helps adjust regularization dynamically using an uncertainty network, helping the learned policy optimize Q-values around dataset modes. TD3-BST outperforms other methods, showcasing state-of-the-art performance when tested on D4RL datasets. 

Tuning TD3-BST is simple and straight, which involves selecting the choice and scale of the kernel (λ), along with the temperature, using primary hyperparameters of the Morse network. For high-dimensional actions, increasing λ helps hold the region around modes tight. Training with Morse-weighted behavioral cloning (BC) reduces the impact of BC loss for distant modes, allowing the policy to focus on selecting and optimizing errors for a single mode. Moreover, the study has proven the importance of letting some OOD actions in the TD3-BST framework, and it depends on λ. 

Simple versions of RL, called One-step algorithms, have the potential to learn a policy from an offline dataset. They depend on weighted BC, which has some limitations, and to improve the performance, relaxing the policy objective will play a major role. A BST objective is integrated into an existing IQL algorithm to overcome this issue and learn an optimal policy while retaining an in-sample policy evaluation. This new approach, IQL-BST, is tested using the same setup as the original IQL, and the results obtained match closely with the original IQL with a very slight drop in performance on larger datasets. However, relaxing weighted BC with a BST objective performs well, especially on difficult-medium and large datasets.

In conclusion, researchers from Imperial College London introduced TD3-BST, an algorithm that uses an uncertainty model to adjust the strength of regularization dynamically. On comparing with previous methods in Gym Locomotion tasks, TD3-BST achieves the best score resulting in strong performance when learning from suboptimal data. In addition, integrating policy regularization with an ensemble-based source of uncertainty enhances the performance. Future work includes: working on different methods to estimate uncertainty, alternative uncertainty measures, and the best way to combine multiple sources of uncertainty.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


YOU MAY ALSO LIKE

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

Snap Introduces A Standalone AI Assistant, Specs Intelligence

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI
AI & Technology

House Passes Ratepayer Protection Act on Data Center Power Costs – Unite.AI

September 16, 2026
Snap Introduces A Standalone AI Assistant, Specs Intelligence
AI & Technology

Snap Introduces A Standalone AI Assistant, Specs Intelligence

September 16, 2026
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
AI & Technology

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

September 16, 2026
Apple’s Redesigned Health App Is Available Now In The iOS 27.2 Developer Beta
AI & Technology

Apple’s Redesigned Health App Is Available Now In The iOS 27.2 Developer Beta

September 16, 2026
Next Post
Meta Earnings Preview | Bloomberg Technology 04/24/2024

Meta Earnings Preview | Bloomberg Technology 04/24/2024

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trying all the viral foods at the U.S. Open

Trying all the viral foods at the U.S. Open

September 14, 2026
Bunching Deductions Into One Year Can Beat the Standard Deduction

Bunching Deductions Into One Year Can Beat the Standard Deduction

September 12, 2026
Wall Street Lunch: Treasury Yields Hit Financial Crisis-Era Highs

Wall Street Lunch: Treasury Yields Hit Financial Crisis-Era Highs

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!