• bitcoinBitcoin(BTC)$86,200.00-0.24%
  • ethereumEthereum(ETH)$2,756.85-0.49%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$788.72-1.24%
  • rippleXRP(XRP)$1.582.48%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$118.57-0.10%
  • tronTRON(TRX)$0.341987-0.72%
  • zcashZcash(ZEC)$1,625.7010.70%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.78%
  • HyperliquidHyperliquid(HYPE)$96.983.53%
  • dogecoinDogecoin(DOGE)$0.1004830.68%
  • moneroMonero(XMR)$571.52-3.60%
  • whitebitWhiteBIT Coin(WBT)$86.71-0.27%
  • chainlinkChainlink(LINK)$13.06-0.52%
  • cardanoCardano(ADA)$0.2559955.17%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013091-6.07%
  • leo-tokenLEO Token(LEO)$8.980.05%
  • stellarStellar(XLM)$0.2163790.44%
  • bitcoin-cashBitcoin Cash(BCH)$343.1528.90%
  • uniswapUniswap(UNI)$10.0812.35%
  • nearNEAR Protocol(NEAR)$4.423.23%
  • avalanche-2Avalanche(AVAX)$11.24-0.04%
  • litecoinLitecoin(LTC)$63.462.57%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.02%
  • CantonCanton(CC)$0.115187-2.33%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0992697.41%
  • suiSui(SUI)$1.03-0.46%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.471.96%
  • shiba-inuShiba Inu(SHIB)$0.0000061.64%
  • BittensorBittensor(TAO)$316.721.27%
  • crypto-com-chainCronos(CRO)$0.0669811.50%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.32-9.87%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,360.520.06%
  • okbOKB(OKB)$123.02-0.03%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BitwayBitway(BTW)$0.899.78%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • aaveAave(AAVE)$147.711.22%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.38%
  • EthenaEthena(ENA)$0.2193523.24%
  • mantleMantle(MNT)$0.672.92%
  • OndoOndo(ONDO)$0.442533-2.55%
  • pepePepe(PEPE)$0.0000054.81%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Unraveling Human Reward Learning: A Hybrid Approach Combining Reinforcement Learning with Advanced Memory Architectures

August 10, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Unraveling Human Reward Learning: A Hybrid Approach Combining Reinforcement Learning with Advanced Memory Architectures
ShareShareShareShareShare

Human reward-guided learning is often modeled using simple RL algorithms that summarize past experiences into key variables like Q-values, representing expected rewards. However, recent findings suggest that these models oversimplify the complexity of human memory and decision-making. For instance, individual events and global reward statistics can significantly influence behavior, indicating that memory involves more than just summary statistics. ANNs, particularly RNNs, offer a more complex model by capturing long-term dependencies and intricate learning mechanisms, though they often need to be more interpretable than traditional RL models.

Researchers from institutions including Google DeepMind, University of Oxford, Princeton University, and University College London studied human reward-learning behavior using a hybrid approach combining RL models with ANNs. Their findings suggest that human behavior needs to be adequately explained by algorithms that incrementally update choice variables. Instead, human reward learning relies on a flexible memory system that forms complex representations of past events over multiple timescales. By iteratively replacing components of a classic RL model with ANNs, they uncovered insights into how experiences shape memory and guide decision-making.

A dataset was gathered from a reward-learning task involving 880 participants. In this task, participants repeatedly chose between four actions, each rewarded based on noisy, drifting reward magnitudes. After filtering, the study included 862 participants and 617,871 valid trials. Most participants learned the task by consistently choosing actions with higher rewards. This extensive dataset enabled significant behavioral variance extraction using RNNs and hybrid models, outperforming basic RL models in capturing human decision-making patterns.

The data was initially modeled using a traditional RL model (Best RL) and a flexible Vanilla RNN. Best RL, identified as the most effective among incremental-update models, employed a reward module to update Q-values and an action module for action perseverance. However, its simplicity limited its expressivity. The Vanilla RNN, which processes actions, rewards, and latent states together, predicted choices more accurately (68.3% vs. 58.9%). Further hybrid models like RL-ANN and Context-ANN, while improving upon Best RL, still fell short of Vanilla RNN. Memory-ANN, incorporating recurrent memory representations, matched Vanilla RNN’s performance, suggesting that detailed memory use was key to participants’ learning in the task.

The study reveals that traditional RL models, which rely solely on incrementally updated decision variables, need to catch up in predicting human choices compared to a novel model incorporating memory-sensitive decision-making. This new model distinguishes between decision variables that drive choices and memory variables that modulate how these decision variables are updated based on past rewards. Unlike RL models, where decision and learning variables are intertwined, this approach separates them, providing a clearer understanding of how learning influences choices. The model suggests that human knowledge is influenced by compressed memories of task history, reflecting both short- and long-term reward and action histories, which modulate learning independently of how they are implemented.

Memory-ANN, the proposed modular cognitive architecture, separates reward-based learning from action-based learning, supported by evidence from computational models and neuroscience. The architecture comprises a “surface” level of decision rules that process observable data and a “deep” level that handles complex, context-rich representations. This dual-layer system allows for flexible, context-driven decision-making, suggesting that human reward learning involves simple surface-level processes and deeper memory-based mechanisms. These findings agree that complex models with rich representations must capture the full spectrum of human behavior, particularly in learning tasks. The insights gained here could have broader applications, extending to various learning tasks and cognitive science.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 48k+ ML SubReddit

Find Upcoming AI Webinars here



Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

YOU MAY ALSO LIKE

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)

Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor


Credit: Source link

ShareTweetSendSharePin

Related Posts

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)
AI & Technology

Why Are Some Songs Grayed Out On Apple Music (And How To Fix It)

September 22, 2026
Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor
AI & Technology

Motorola’s New Signature 27 Is Among The First Smartphone To Use The Snapdragon 8 Elite Extreme Gen 6 Processor

September 22, 2026
Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5
AI & Technology

Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

September 22, 2026
The Latest PlayStation Update Made PSSR 2.0 The Default For PS5 Pro Owners
AI & Technology

The Latest PlayStation Update Made PSSR 2.0 The Default For PS5 Pro Owners

September 22, 2026
Next Post
News Corporation (NWS) Q4 2024 Earnings Call Transcript

News Corporation (NWS) Q4 2024 Earnings Call Transcript

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Synthesia Opens 50,000-Square-Foot US Headquarters in New York – Unite.AI

Synthesia Opens 50,000-Square-Foot US Headquarters in New York – Unite.AI

September 18, 2026
Russia strikes Kyiv after pause for U.S. envoys’ visit

Russia strikes Kyiv after pause for U.S. envoys’ visit

September 16, 2026
Flash Flooding Overwhelms Nepal, and Jury Deliberates in Lindsay Clancy’s Murder Trial | Aug. 27

Flash Flooding Overwhelms Nepal, and Jury Deliberates in Lindsay Clancy’s Murder Trial | Aug. 27

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!