When does system speed influence what gets learned?
In asynchronous reinforcement learning from human feedback, rollouts may complete at different times. If the training pipeline selects or consumes experiences as they become available, response latency can become entangled with which samples are used for learning.
This project studies latency-conditioned selection bias in RLHF, with a focus on asynchronous partial rollouts and length-related selection effects. The project page will collect the paper, conference information, and accompanying research materials in one place.