Citation

Jaripour, E. (2026). Trial-by-Trial Behavioral Adaptation in a Restless Bandit Task: A Mixed-Effects Modeling Approach. Research Square Preprint. DOI: 10.21203/rs.3.rs-10740751/v1

Abstract

Adaptive decision-making requires individuals to modify behavior as reward environments change over time. Restless bandit tasks provide a framework for studying behavioral adaptation under dynamic conditions, although observable changes do not necessarily reveal latent cognitive mechanisms. The present study examined trial-by-trial behavioral adaptation in a four-arm restless bandit task using a publicly available dataset of 965 participants and 139,816 analyzed trials. Rather than estimating latent reinforcement-learning parameters, the study examined observable outcomes, trial-related change, differences across payoff environments, and individual differences in baseline behavior and adaptation. Mixed-effects models were fitted to payoff-maximizing choice, obtained reward, log-transformed reaction time, and choice switching. The primary analysis used a quadratic trial trajectory with payoff-environment interactions and participant-specific random intercepts, linear slopes, and quadratic slopes. The quadratic model provided substantially better AIC fit than the corresponding linear interaction model. Payoff-maximizing choice showed distinct linear and quadratic trajectories across payoff environments, while participants varied in baseline performance and trial-related change. Secondary analyses showed different reward trajectories across payoff environments, decreasing reaction times across trials without clear payoff-group differences, and decreasing choice switching, with evidence that this trajectory differed for Group 3 but not Group 4 relative to the reference environment. These findings indicate nonlinear changes in observable decision behavior during repeated decisions in a dynamic reward environment. Because latent values, beliefs, and learning parameters were not estimated, the findings do not identify the mechanisms underlying these changes. Instead, they provide a quantitative characterization of observable behavioral adaptation and a basis for future computational investigations of decision-making in dynamic environments.

Key Contributions

  • Quantified trial-by-trial changes in payoff-maximizing choice using a publicly available four-arm restless bandit dataset.
  • Estimated nonlinear behavioral trajectories across different payoff environments using quadratic mixed-effects models.
  • Quantified individual differences in baseline behavior and trial-related adaptation through participant-specific random effects.
  • Examined complementary behavioral outcomes, including obtained reward, reaction time, and choice switching.
  • Evaluated alternative fixed-effects, random-effects, and trial-trajectory specifications using model comparison and robustness analyses within a fully reproducible research workflow.

Resources