Research project · 2026

Latency-Conditioned Selection Bias in RLHF

How asynchronous partial rollouts and response latency can shape which experiences enter the reinforcement learning from human feedback pipeline.

Prakul Sunil Hiremath NeurIPS 2026 · Main Conference · Sydney, Australia
NeurIPS 2026 RLHF Selection bias

When does system speed influence what gets learned?

In asynchronous reinforcement learning from human feedback, rollouts may complete at different times. If the training pipeline selects or consumes experiences as they become available, response latency can become entangled with which samples are used for learning.

This project studies latency-conditioned selection bias in RLHF, with a focus on asynchronous partial rollouts and length-related selection effects. The project page will collect the paper, conference information, and accompanying research materials in one place.

Project page in development Additional technical details, figures, and research materials will be added here.

Paper and project resources

Use the official conference record, paper discussion, and code repository for the current materials.