Oral Presentation: Replicable Bandits for Digital Health
Abstract: Adaptive treatment assignment algorithms, such as bandit and reinforcement learning algorithms, are increasingly used in digital health interventions. Between implementation of the digital health intervention, data analyses are critical for producing generalizable knowledge and deciding how to update the intervention for the next implementation. However the replicability of these between-implementation data analyses has received relatively little attention. This work investigates the replicability of statistical analyses from data collected by adaptive treatment assignment algorithms. We demonstrate that many standard statistical estimators can be inconsistent and fail to be replicable across repetitions of the clinical trial, even as the sample size grows large. We show that this non-replicability is intimately related to properties of the adaptive algorithm itself. We introduce a formal definition of a 'replicable bandit algorithm' and prove that under such algorithms, a wide variety of common statistical analyses are guaranteed to be consistent.…
