Benchmarking Virtual Cell Models for In-the-Wild Perturbation Response
Under revision at a Nature Portfolio journal · paper · code · page
Abstract
A standardized benchmarking framework for single-cell perturbation prediction, harmonizing datasets, model interfaces, and evaluation protocols. It covers three out-of-distribution scenarios — unseen cell contexts, unseen perturbations, and cross-dataset generalization — revealing that task design strongly influences model ranking and that cross-dataset generalization remains the key challenge.

. My research focuses on reliable learning and decision-making, especially
long-horizon reasoning, agentic systems, and self-improving learning,
for scientific discovery and real-world applications. Previously at
,
,
,
.