Publication: Causal Directed Acyclic Graph-informed Reward Design

Authors: Luton Zou, Ziping Xu, Daiqi Gao, Susan Murphy Abstract It is well known that in reinforcement learning (RL) different reward functions may lead to the same optimal policy, while some reward functions can be substantially easier to learn. In this paper, we propose a framework for reward design by constructing surrogate rewards with mediators informed by causal directed acyclic graphs (DAGs), which are often available in real-world applications through domain knowledge. We show that under the surrogacy assumption, the proposed reward is unbiased and has lower variance than the primary reward. Specifically, we use an online reward design agent that adaptively learns the target surrogate reward in an unknown environment. Feeding the surrogate rewards to standard online learning oracles, we show that the regret bound can be improved. Our framework provides a theoretical improvement…

Continue ReadingPublication: Causal Directed Acyclic Graph-informed Reward Design

Publication: Digital Twins for Just-in-Time Adaptive Interventions (JITAI-Twins): A Framework for Optimizing and Continually Improving JITAIs

Authors: Asim H. Gazi, Daiqi Gao, Susobhan Ghosh, Ziping Xu, Anna Trella, Predrag Klasnja, Susan A. Murphy Abstract Just-in-time adaptive interventions (JITAIs) are nascent precision medicine systems that extend personalized healthcare support to everyday life. A challenge in designing JITAIs is that personalized support often involves sophisticated decision-making algorithms. These decision-making algorithms can require numerous non-trivial design decisions that must be made between successive JITAI deployments (e.g., hyperparameter selection for an artificial intelligence algorithm). Making design decisions between deployments–rather than during deployment–ensures intervention fidelity and enhances the ability to replicate results. Yet, each deployment can be costly, precluding the use of A/B testing for every design decision. How should design decisions be made strategically between JITAI deployments? This paper introduces digital twins for just-in-time adaptive interventions (JITAI-Twins) to address this question. JITAI-Twins are “digital twins…

Continue ReadingPublication: Digital Twins for Just-in-Time Adaptive Interventions (JITAI-Twins): A Framework for Optimizing and Continually Improving JITAIs

Oral Presentation: Inference for Longitudinal Data After Adaptive Sampling

This presentation was given on November 20, 2024 by Dr. Susan Murphy as part of the L. Brown Distinguished Lecture series at University of Pennsylvania Abstract: Adaptive sampling methods, such as reinforcement learning (RL) and bandit algorithms, are increasingly used for the real-time personalization of interventions in digital applications like mobile health and education. As a result, there is a need to be able to use the resulting adaptively collected user data to address a variety of inferential questions, including questions about time-varying causal effects. However, current methods for statistical inference on such data (a) make strong assumptions regarding the environment dynamics, e.g., assume the longitudinal data follows a Markovian process, or (b) require data to be collected with one adaptive sampling algorithm per user, which excludes algorithms that learn to select actions using data…

Continue ReadingOral Presentation: Inference for Longitudinal Data After Adaptive Sampling

Oral Presentation: NIA workshop – Leveraging Adaptive Technology (“Just-in-Time”) Interventions for Aging and Alzheimer’s Disease and Alzheimer’s Disease-related Dementias

Dr. Inbal Billie Nahum-Shani (PI) presented work from this pilot project during a talk titled "Adaptive interventions and JITAIs as decision policies: What and why?" as part of Session 1 - Digital adaptive interventions: decision-focused evidence production held on October 16, 2024. Source: NIA Event Page

Continue ReadingOral Presentation: NIA workshop – Leveraging Adaptive Technology (“Just-in-Time”) Interventions for Aging and Alzheimer’s Disease and Alzheimer’s Disease-related Dementias

Oral Presentation: Replicable Bandits for Digital Health

Abstract: Adaptive treatment assignment algorithms, such as bandit and reinforcement learning algorithms, are increasingly used in digital health interventions. Between implementation of the digital health intervention, data analyses are critical for producing generalizable knowledge and deciding how to update the intervention for the next implementation. However the replicability of these between-implementation data analyses has received relatively little attention. This work investigates the replicability of statistical analyses from data collected by adaptive treatment assignment algorithms. We demonstrate that many standard statistical estimators can be inconsistent and fail to be replicable across repetitions of the clinical trial, even as the sample size grows large. We show that this non-replicability is intimately related to properties of the adaptive algorithm itself. We introduce a formal definition of a 'replicable bandit algorithm' and prove that under such algorithms, a wide variety of common statistical analyses are guaranteed to be consistent.…

Continue ReadingOral Presentation: Replicable Bandits for Digital Health