Tracking and optimizing human performance using deep reinforcement learning in closed-loop behavioral- and neurofeedback: a proof of concept
Reinforcement learning (RL) is a general-purpose powerful machine learning framework within which we can model various deterministic, non-deterministic and complex environments. We applied RL to the problem of tracking and improving human sustained attention during a simple sustained attention to response task (SART) in a proof of concept study with two subjects, using state-of-the-art deep neural network-based RL in the form of Deep Q Networks (DQNs). While others have used RL in EEG settings previously, none have applied it in a neurofeedback (NFB) setting, which seems a natural problem within Brain Computer Interfaces (BCIs) to tackle using end-to-end RL in the form of DQNs, due to both the problems non-stationarity and the ability of RL to learn in a continuous setting. Furthermore, while many have explored phasic alerting previously, learning optimal alerting in a personalized way in real time is a less explored field, which we believe RL to be an most suitable solution for. First, we used empirically-derived simulated data of EEG and reaction times and subsequent parameter/algorithmic exploration within this simulated model to pick parameters for the DQN that are more likely to be optimal for the experimental setup and to explore the behavior of DQNs in this task setting. We then applied the method on two subjects and show that we get different but plausible results for both subjects, suggesting something about the behavior of DQNs in this setting. For this experimental part, we used parameters suggested to us by the simulation results. This RL-based behavioral- and neuro-feedback BCI method we have developed here is input feature agnostic and allows for complex continuous actions to be learned in other more complex closed-loop behavioral or neuro-feedback approaches.