Reasoning & post-training
RLVR, rollout efficiency, reward signals, and the internal dynamics of model reasoning.
Research workbench
I study model behavior through controlled experiments, with particular interest in reasoning, post-training, evaluation, and scientific machine learning.
Independent research · Ongoing · 2026
Linear probes can predict eventual correctness above chance, but separating trajectory signal from problem difficulty is the central open issue.
at 80% kill precision; 0.93 advantage correlation
Can hidden states reveal that a reasoning trajectory is unlikely to succeed early enough to save rollout compute?
Linear probes over residual-stream states from 16K verifier-graded MATH rollouts, followed by simulated probe-guided truncation.
Problem difficulty can masquerade as trajectory-level failure signal. Calibration and out-of-distribution generalization remain unresolved.
Working questions
RLVR, rollout efficiency, reward signals, and the internal dynamics of model reasoning.
Measurements that distinguish real model behavior from dataset artifacts and confounding.
Models whose design reflects complex biological, clinical, and temporal data-generating processes.
Selected writing
Peer-reviewed work across genomics, deep learning, and cellular trajectory analysis.
Belfort B, Jia J*
Frontiers in NeuroscienceCo-first author
Xu H, Jia J*, Jeong H, Zhao Z
Cell PatternsCo-first author
Jeong H, Jia J, Dai Y, Simon L
Genes
* Co-first authorship where indicated.
Find my work on Google Scholar