Evidence map›Paper›PMID 38552190›Full record

ArticlePLoS computational biology2024

Active reinforcement learning versus action bias and hysteresis: control with a mixture of experts and nonexperts.

Jaron T Colas, John P O'Doherty, Scott T Grafton

Abstract read
In one paragraph

Article in PLoS computational biology, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Social information creates self-fulfilling prophecies in judgments of pain, vicarious pain, and cognitive effort.Proceedings of the National Academy of Sciences of the United States of America · 2026
    Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Jaron T ColasDepartment of Psychological and Brain Sciences, University of California, Santa Barbara, California, United States of America.ORCID 0000-0003-1872-7614
John P O'DohertyDivision of the Humanities and Social Sciences, California Institute of Technology, Pasadena, California, United States of America.ORCID 0000-0003-0016-3531
Scott T GraftonDepartment of Psychological and Brain Sciences, University of California, Santa Barbara, California, United States of America.ORCID 0000-0003-4015-3151

Funding

The Neurobiology of Social Decision-Making: Social Inference and ContextP50MH094258 · NIMH · CALIFORNIA INSTITUTE OF TECHNOLOGY · PI ANDERSEN, RICHARD A · 2012 to 2021
$18.6M
Determining the neural substrates of model-based and model-free reinforcement-learning during Pavlovian conditioning (Minority Supplement)R01DA040011 · NIDA · CALIFORNIA INSTITUTE OF TECHNOLOGY · PI O'DOHERTY, JOHN P · 2016 to 2020
$2.6M
NIDA NIH HHS R01 DA040011NIMH NIH HHS P50 MH094258
6 · The paper itself

Abstract

Active reinforcement learning enables dynamic prediction and control, where one should not only maximize rewards but also minimize costs such as of inference, decisions, actions, and time. For an embodied agent such as a human, decisions are also shaped by physical aspects of actions. Beyond the effects of reward outcomes on learning processes, to what extent can modeling of behavior in a reinforcement-learning task be complicated by other sources of variance in sequential action choices? What of the effects of action bias (for actions per se) and action hysteresis determined by the history of actions chosen previously? The present study addressed these questions with incremental assembly of models for the sequential choice data from a task with hierarchical structure for additional complexity in learning. With systematic comparison and falsification of computational models, human choices were tested for signatures of parallel modules representing not only an enhanced form of generalized reinforcement learning but also action bias and hysteresis. We found evidence for substantial differences in bias and hysteresis across participants-even comparable in magnitude to the individual differences in learning. Individuals who did not learn well revealed the greatest biases, but those who did learn accurately were also significantly biased. The direction of hysteresis varied among individuals as repetition or, more commonly, alternation biases persisting from multiple previous actions. Considering that these actions were button presses with trivial motor demands, the idiosyncratic forces biasing sequences of action choices were robust enough to suggest ubiquity across individuals and across tasks requiring various actions. In light of how bias and hysteresis function as a heuristic for efficient control that adapts to uncertainty or low motivation by minimizing the cost of effort, these phenomena broaden the consilient theory of a mixture of experts to encompass a mixture of expert and nonexpert controllers of behavior.

Indexed as

LearningReinforcement, PsychologyBiasHumansProblem-Based LearningReward

Identifiers

PMID38552190
PMCPMC10980507

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.