과거 정책의 rollout으로 현재 정책에 유용한 prompt를 고를 때, 핵심은 평균 score 오차가 아니라 top-k 경계 margin에 비해 오차가 충분히 작은가이다. 이 논문은 exact recovery 조건, partial correction의 식별 불가능성, 그리고 실제 비교가 가능한지를 먼저 판정하는 split-half protocol을 분리한다.
현재 정책 $\pi$의 prompt utility를 과거 정책 $\beta$의 고정 rollout으로 추정하면 score error가 생긴다. 현재 top-$k$ 경계의 margin을 $\Delta_k$, 모든 prompt의 최대 score error를 $e$라 하면 $\Delta_k>2e$일 때 stale ranking과 on-policy ranking의 top-$k$ 집합이 정확히 같다. 반대로 history distribution이나 action value 중 한쪽만 보정한 정보로는 일반적으로 현재 ranking을 식별할 수 없다. 작은 KL과 높은 ESS만으로 이 빈 정보를 복원할 수는 없다.
오래된 답안으로 지금 학생에게 맞는 문제를 고르는 상황이다. 문제 사이의 실제 차이가 추정 오차보다 충분히 크면 선택은 유지된다. 경계의 두 문제가 거의 동점이면, 전체 점수 상관이 높아도 마지막 한 자리가 바뀔 수 있다.
RLVR pipeline은 rollout 비용을 줄이기 위해 behavior-policy 응답을 재사용한다. 하지만 stale group으로 정책을 안정적으로 업데이트할 수 있다는 사실과, 그 응답이 현재 정책 기준의 prompt 순서를 보존한다는 사실은 다르다. 이 연구의 estimand는 downstream benchmark 점수가 아니라 held-out validation gradient와의 projected cosine alignment다. 실제로 소비되는 출력은 scalar score가 아니라 그 score로 선택한 top-$k$ 집합이다.
stale rollout으로 학습하는 안정성과 prompt ranking 보존은 같은 주장이 아니다
gradient alignment는 selection utility proxy이며 실제 reward 향상은 matched update가 필요하다
reference reliability gate 실패는 estimator 실패가 아니라 미판정이다
응답 $Y=(A_1,\dots,A_T)$, terminal reward $R(Y)$, token score $z_t=\nabla_\theta\log\pi(A_t\mid H_t)$를 둔다. Token ratio $r_j=\pi(A_j\mid H_j)/\beta(A_j\mid H_j)$, prefix product $P_t=\prod_{j<t}r_j$, suffix product $S_t=\prod_{j>t}r_j$를 쓰면 네 population estimator가 생긴다.
| 추정량 | history distribution | action value | 해석 |
|---|---|---|---|
| $g_{00}$ | $\beta$ | $\beta$ | current-token ratio만 사용 |
| $g_{10}$ | $\pi$ | $\beta$ | history/prefix만 보정 |
| $g_{01}$ | $\beta$ | $\pi$ | continuation/action value만 보정 |
| $g_{11}$ | $\pi$ | $\pi$ | exact unclipped population quantity |
$g_{10}$과 $g_{01}$은 서로 다른 bias를 제거한다. 한쪽 보정이 다른 쪽의 대체재라는 일반 보장은 없다. 이 등식은 finite horizon, support, exact unclipped ratio를 전제로 하므로 실제 capped estimator를 자동으로 certify하지 않는다.
현재 scalar utilities를 $s_{(1)}\ge\cdots\ge s_{(n)}$, stale estimates를 $\widetilde s_i$라 하자. Selection margin과 uniform error는 다음과 같다.
각 true top-$k$ item과 각 unselected item 사이의 간격은 perturbation 뒤에도 최소 $\Delta_k-2e>0$이다. 이 결과는 임의의 scalar score에 적용되지만, empirical cosine proxy의 오차 $e$를 관측 없이 알고 있다는 뜻은 아니다.
History-only $g_{10}$ 또는 action-value-only $g_{01}$ 중 하나만 관측하는 어떤 selection rule에도, 관측 분포는 같지만 true top-$k$ 집합이 다른 두 binary-reward prompt pool이 존재한다. 이 구성은 임의로 작은 $D_{KL}(\pi\Vert\beta)$, unit normalized ESS, 모든 $K\ge2$의 GRPO group normalization 아래에서도 성립한다. 따라서 minimax exact-selection error는 적어도 $1/2$이다.
이 정리는 모든 실제 stale estimator가 항상 실패한다는 말이 아니다. 한쪽 정보만 보고도 항상 안전하다고 선언하는 보편적 certificate가 없다는 말이다. 실제 빈도는 아래의 gated experiment가 측정해야 한다.
Candidate와 validation의 on-policy samples를 ranking split R과 서로 독립인 reference halves A/B로 분리한다. R은 fresh comparator를 만들고, A/B scalar score 평균은 held-out utility reference가 된다. A와 B가 만든 top-$k$ overlap이 split-half reliability다. Tie stream도 독립적으로 고정해 accidental agreement가 reliability로 들어가지 않게 한다.
stale selectors와 R 기반 fresh selector가 prompt를 각각 ranking한다
A/B top-k agreement와 independent-reference interval로 reference 재현성을 검사한다
reliability와 positive fresh-gain gate 통과 후에만 retention과 regime label을 해석한다
Gaussian latent-score model에서는 population split-half reliability가 independent selector가 noisy reference에 대해 보일 수 있는 expected precision의 ceiling을 결정한다. 이 ceiling은 Gaussianity, equal noise variance, conditional independence 가정 아래의 결과이며 모든 데이터에 대한 무가정 상한이 아니다. Exact certification의 표본 비용은 boundary gap $\Delta$에 대해 $\Omega(\sigma^2\Delta^{-2}\log(1/\delta))$로 증가한다.
| 구성 | 고정 값 |
|---|---|
| Base model | raw non-SFT allenai/Olmo-3-1025-7B |
| Domains | MATH-500 symbolic/numeric math, MBPP executable Python |
| Seeds · checkpoints | 5 seeds · RLVR steps 0/25/100/400 |
| Matrix size | 10 dataset-seed families · 40 checkpoint points |
| Candidate · validation | MATH-500 400/100, MBPP 512/100 |
| Rollouts | behavior/current/validation 8/32/8; current R/A/B=16/8/8 |
| Selection | top 10%; fixed 4096-d CountSketch over final four layers + final norm |
| Training | 4 H100 DDP ranks; 8 responses/group; one-epoch verifier-reward GRPO |
| Parameterization | q/v LoRA rank 16, alpha 32, dropout 0 |
Positive checkpoints form one continuous adapter-and-optimizer lineage. One fresh group is used for one optimizer epoch, so the PPO-form ratio is one at the only loss evaluation up to numerical error; clip_epsilon=0.2 is recorded but not claimed as an effective trust region. Homogeneous-reward groups are valid zero-advantage groups and are counted rather than retried.
비교는 split-half reliability의 one-sided 95% lower bound가 $2k/n$ 이상이고 fresh selection gain의 lower bound가 0보다 클 때만 admissible하다. Effective는 등록 seed의 80% 이상이 gate를 통과하고 stale gain lower bound가 양수이며 fresh gain의 50% 이상을 보존할 때다. Ineffective는 80% 이상이 gate를 통과하지만 stale gain upper bound가 0 이하일 때다. 그 밖은 inconclusive다.
5개 main seed 중 최소 4개가 모든 조건을 충족
gate는 통과했으나 stale gain의 상한이 0 이하
reference, fresh gain, seed support 중 하나라도 부족
세 개의 4×H100 node는 동일 command와 물리적으로 공유된 volume을 사용한다. Shared queue는 dataset-seed family 전체를 한 worker에 할당하고 0→25→100→400 lineage를 유지한다. Rollout은 exact policy/adaptor hash와 RNG domain에 묶이며, 중단된 generation은 마지막 complete prompt group 뒤에서, training은 마지막 5-step checkpoint에서 재개한다. Compute mode는 Hugging Face credential discovery와 network를 끄고, 준비된 local snapshot의 revision·content hash·schema·verifier self-test를 GPU 작업 전에 검사한다.
``bash git pull --ff-only bash scripts/run_olmo3_rlzero.sh run ``
모든 40 points가 objective, lineage, generation, analysis, manifest contract를 통과하고 10,000회 bootstrap이 끝나기 전에는 aggregate table을 동결하지 않는다. Active cluster의 partial output은 논문 수치가 아니다.
| 축 | 등록 값 |
|---|---|
| Models | raw sparse OLMoE-1B-7B-0125; raw dense Qwen2.5-14B |
| Domains | Knights-and-Knaves logic; ARC-Challenge science; non-math MMLU-Pro knowledge |
| Objective | 모든 cell에서 동일한 one-epoch verifier-reward GRPO |
| Scale | 2 models × 3 domains × 3 seeds × 4 checkpoints = 72 points |
메인 OLMo-3/math/code와 다른 architecture, capacity, domain, verifier에서 observed-stratum replication을 본다. 단, 두 확장 모델은 architecture와 capacity가 함께 달라지므로 각각의 인과 효과를 분리하지 못한다. Single-domain run이므로 cross-domain transfer, mixed-domain curriculum, alternative RLVR objective, full-parameter training, multilingual generalization은 주장하지 않는다.
이전 positive-only response cross-entropy sweep는 RLVR이 아니라 SFT였다. 그 20개 run은 appendix의 supervised policy-shift ablation으로만 보존한다. GRPO checkpoint로 resume하거나, main regime table을 채우거나, RLVR 결과로 인용할 수 없다. LoRA는 trainable parameterization일 뿐 학습 objective가 아니다.
stored trajectory나 gradient alignment로 training data의 영향도를 추정한다
history와 continuation을 서로 다른 방식으로 보정한다
prompt utility, difficulty, uncertainty, curriculum을 사용한다
stale group 학습, prompt identity replay, rollout 배분은 ranking 보존과 다른 intervention이다
global correlation보다 boundary gap과 reference noise가 결정을 지배한다
이론은 exact unclipped population ratios와 unnormalized directional gradient를 다룬다. 실험은 capped ratios, CountSketch, cosine normalization을 쓰므로 정리를 그대로 구현한 것은 아니다. Main study는 한 raw 7B model, 두 domain, LoRA, one-epoch GRPO에 한정된다. Bootstrap interval은 fixed prompt pool과 fixed stale selector에 조건부이며, 실제 selected-subset training 뒤의 downstream reward는 아직 측정 대상이 아니다.
현재 공개 가능한 결론은 theory와 protocol뿐이다. Main 40 points와 extension 72 points의 완전한 artifact gate가 끝나기 전에는 estimator 우열, drift boundary, model/domain generalization 수치를 제시하지 않는다.
stale score error가 현재 top-k margin의 절반보다 작으면 exact set가 보존된다
작은 KL과 unit ESS에서도 true ranking이 식별되지 않을 수 있다
reference가 재현될 때만 stale-vs-fresh comparison을 해석한다
evidence에 따라 stale reuse, fresh sampling, 판단 보류 중 하나를 선택한다