RLVR 데이터 선택 · Policy Drift · Top-k Reliability

When Do Stale Rollouts Preserve On-Policy Prompt Rankings?

과거 정책의 rollout으로 현재 정책에 유용한 prompt를 고를 때, 핵심은 평균 score 오차가 아니라 top-k 경계 margin에 비해 오차가 충분히 작은가이다. 이 논문은 exact recovery 조건, partial correction의 식별 불가능성, 그리고 실제 비교가 가능한지를 먼저 판정하는 split-half protocol을 분리한다.

Protocol draft · ICLR 2027 target 🗓️ 2026-09 (protocol draft) ✍️ Under double-blind review 🏷️ cs.LG
2 × 2
history distribution과 action value의 두 correction 축
Δk > 2e
stale score가 exact top-k를 보존하는 충분조건
10 families · 40 points
OLMo-3 7B × MATH-500/MBPP × 5 seeds × 4 checkpoints
결과 미동결
40개 전 지점의 계약과 reliability gate 통과 전 수치 주장 없음
1SUMMARY

한 문장 요약 — stale reuse는 margin 문제다

논문 그대로

현재 정책 $\pi$의 prompt utility를 과거 정책 $\beta$의 고정 rollout으로 추정하면 score error가 생긴다. 현재 top-$k$ 경계의 margin을 $\Delta_k$, 모든 prompt의 최대 score error를 $e$라 하면 $\Delta_k>2e$일 때 stale ranking과 on-policy ranking의 top-$k$ 집합이 정확히 같다. 반대로 history distribution이나 action value 중 한쪽만 보정한 정보로는 일반적으로 현재 ranking을 식별할 수 없다. 작은 KL과 높은 ESS만으로 이 빈 정보를 복원할 수는 없다.

쉽게 풀면

오래된 답안으로 지금 학생에게 맞는 문제를 고르는 상황이다. 문제 사이의 실제 차이가 추정 오차보다 충분히 크면 선택은 유지된다. 경계의 두 문제가 거의 동점이면, 전체 점수 상관이 높아도 마지막 한 자리가 바뀔 수 있다.

2MOTIVATION

왜 이 질문인가 — selection은 training과 다른 결정이다

논문 그대로

RLVR pipeline은 rollout 비용을 줄이기 위해 behavior-policy 응답을 재사용한다. 하지만 stale group으로 정책을 안정적으로 업데이트할 수 있다는 사실과, 그 응답이 현재 정책 기준의 prompt 순서를 보존한다는 사실은 다르다. 이 연구의 estimand는 downstream benchmark 점수가 아니라 held-out validation gradient와의 projected cosine alignment다. 실제로 소비되는 출력은 scalar score가 아니라 그 score로 선택한 top-$k$ 집합이다.

Figure 1. Prompt selection의 비용 스펙트럼. 값싼 history signal, stale gradient score, 일부 fresh audit, 전수 on-policy score는 비용과 정보량이 다르다. 본 논문은 stale score가 언제 top-k 결정을 보존하는지 묻고, 불확실한 경우에는 fresh evidence를 추가하거나 판단을 보류한다.
구분 1

Selection vs optimization

stale rollout으로 학습하는 안정성과 prompt ranking 보존은 같은 주장이 아니다

구분 2

Proxy vs outcome

gradient alignment는 selection utility proxy이며 실제 reward 향상은 matched update가 필요하다

구분 3

Failure vs unresolved

reference reliability gate 실패는 estimator 실패가 아니라 미판정이다

3ESTIMATION

추정량 — 두 축의 2×2 분해

논문 그대로

응답 $Y=(A_1,\dots,A_T)$, terminal reward $R(Y)$, token score $z_t=\nabla_\theta\log\pi(A_t\mid H_t)$를 둔다. Token ratio $r_j=\pi(A_j\mid H_j)/\beta(A_j\mid H_j)$, prefix product $P_t=\prod_{j<t}r_j$, suffix product $S_t=\prod_{j>t}r_j$를 쓰면 네 population estimator가 생긴다.

$$ g_{00}=\sum_t\mathbb E_\beta[r_tRz_t],\qquad g_{10}=\sum_t\mathbb E_\beta[P_tr_tRz_t] $$
$$ g_{01}=\sum_t\mathbb E_\beta[r_tS_tRz_t],\qquad g_{11}=\sum_t\mathbb E_\beta[P_tr_tS_tRz_t]=g_\pi $$
추정량history distributionaction value해석
$g_{00}$$\beta$$\beta$current-token ratio만 사용
$g_{10}$$\pi$$\beta$history/prefix만 보정
$g_{01}$$\beta$$\pi$continuation/action value만 보정
$g_{11}$$\pi$$\pi$exact unclipped population quantity
논문 그대로

$g_{10}$과 $g_{01}$은 서로 다른 bias를 제거한다. 한쪽 보정이 다른 쪽의 대체재라는 일반 보장은 없다. 이 등식은 finite horizon, support, exact unclipped ratio를 전제로 하므로 실제 capped estimator를 자동으로 certify하지 않는다.

4GUARANTEE

이론 1 — exact top-k recovery 조건

논문 그대로

현재 scalar utilities를 $s_{(1)}\ge\cdots\ge s_{(n)}$, stale estimates를 $\widetilde s_i$라 하자. Selection margin과 uniform error는 다음과 같다.

$$ \Delta_k=s_{(k)}-s_{(k+1)},\qquad e=\max_i|\widetilde s_i-s_i|. $$
$$ \boxed{\Delta_k>2e\quad\Longrightarrow\quad \operatorname{TopK}(\widetilde s)=\operatorname{TopK}(s)} $$
논문 그대로

각 true top-$k$ item과 각 unselected item 사이의 간격은 perturbation 뒤에도 최소 $\Delta_k-2e>0$이다. 이 결과는 임의의 scalar score에 적용되지만, empirical cosine proxy의 오차 $e$를 관측 없이 알고 있다는 뜻은 아니다.

5IDENTIFIABILITY

이론 2 — partial correction의 식별 한계

논문 그대로

History-only $g_{10}$ 또는 action-value-only $g_{01}$ 중 하나만 관측하는 어떤 selection rule에도, 관측 분포는 같지만 true top-$k$ 집합이 다른 두 binary-reward prompt pool이 존재한다. 이 구성은 임의로 작은 $D_{KL}(\pi\Vert\beta)$, unit normalized ESS, 모든 $K\ge2$의 GRPO group normalization 아래에서도 성립한다. 따라서 minimax exact-selection error는 적어도 $1/2$이다.

$$ u^\top g_\pi=-\frac{\epsilon}{2},\qquad u^\top g_c=+\frac{\epsilon}{2},\qquad D_{KL}(\pi\Vert\beta)=8\epsilon^2+O(\epsilon^4). $$
쉽게 풀면

이 정리는 모든 실제 stale estimator가 항상 실패한다는 말이 아니다. 한쪽 정보만 보고도 항상 안전하다고 선언하는 보편적 certificate가 없다는 말이다. 실제 빈도는 아래의 gated experiment가 측정해야 한다.

6MEASUREMENT

측정 — reference 자체를 먼저 감사한다

논문 그대로

Candidate와 validation의 on-policy samples를 ranking split R과 서로 독립인 reference halves A/B로 분리한다. R은 fresh comparator를 만들고, A/B scalar score 평균은 held-out utility reference가 된다. A와 B가 만든 top-$k$ overlap이 split-half reliability다. Tie stream도 독립적으로 고정해 accidental agreement가 reliability로 들어가지 않게 한다.

1 · Estimate

stale selectors와 R 기반 fresh selector가 prompt를 각각 ranking한다

2 · Audit

A/B top-k agreement와 independent-reference interval로 reference 재현성을 검사한다

3 · Decide

reliability와 positive fresh-gain gate 통과 후에만 retention과 regime label을 해석한다

논문 그대로

Gaussian latent-score model에서는 population split-half reliability가 independent selector가 noisy reference에 대해 보일 수 있는 expected precision의 ceiling을 결정한다. 이 ceiling은 Gaussianity, equal noise variance, conditional independence 가정 아래의 결과이며 모든 데이터에 대한 무가정 상한이 아니다. Exact certification의 표본 비용은 boundary gap $\Delta$에 대해 $\Omega(\sigma^2\Delta^{-2}\log(1/\delta))$로 증가한다.

7PRIMARY

메인 실험 — raw OLMo-3에서 연속 GRPO drift

구성고정 값
Base modelraw non-SFT allenai/Olmo-3-1025-7B
DomainsMATH-500 symbolic/numeric math, MBPP executable Python
Seeds · checkpoints5 seeds · RLVR steps 0/25/100/400
Matrix size10 dataset-seed families · 40 checkpoint points
Candidate · validationMATH-500 400/100, MBPP 512/100
Rolloutsbehavior/current/validation 8/32/8; current R/A/B=16/8/8
Selectiontop 10%; fixed 4096-d CountSketch over final four layers + final norm
Training4 H100 DDP ranks; 8 responses/group; one-epoch verifier-reward GRPO
Parameterizationq/v LoRA rank 16, alpha 32, dropout 0
논문 그대로

Positive checkpoints form one continuous adapter-and-optimizer lineage. One fresh group is used for one optimizer epoch, so the PPO-form ratio is one at the only loss evaluation up to numerical error; clip_epsilon=0.2 is recorded but not claimed as an effective trust region. Homogeneous-reward groups are valid zero-advantage groups and are counted rather than retried.

8DECISION

판정 규칙 — 수치를 보기 전에 고정

논문 그대로

비교는 split-half reliability의 one-sided 95% lower bound가 $2k/n$ 이상이고 fresh selection gain의 lower bound가 0보다 클 때만 admissible하다. Effective는 등록 seed의 80% 이상이 gate를 통과하고 stale gain lower bound가 양수이며 fresh gain의 50% 이상을 보존할 때다. Ineffective는 80% 이상이 gate를 통과하지만 stale gain upper bound가 0 이하일 때다. 그 밖은 inconclusive다.

Effective

측정 가능하고 유용

5개 main seed 중 최소 4개가 모든 조건을 충족

Ineffective

측정 가능하지만 이득 없음

gate는 통과했으나 stale gain의 상한이 0 이하

Inconclusive

판단 보류

reference, fresh gain, seed support 중 하나라도 부족

9INTEGRITY

실행 무결성 — cluster 결과가 들어오기 위한 계약

논문 그대로

세 개의 4×H100 node는 동일 command와 물리적으로 공유된 volume을 사용한다. Shared queue는 dataset-seed family 전체를 한 worker에 할당하고 0→25→100→400 lineage를 유지한다. Rollout은 exact policy/adaptor hash와 RNG domain에 묶이며, 중단된 generation은 마지막 complete prompt group 뒤에서, training은 마지막 5-step checkpoint에서 재개한다. Compute mode는 Hugging Face credential discovery와 network를 끄고, 준비된 local snapshot의 revision·content hash·schema·verifier self-test를 GPU 작업 전에 검사한다.

``bash git pull --ff-only bash scripts/run_olmo3_rlzero.sh run ``

논문 그대로

모든 40 points가 objective, lineage, generation, analysis, manifest contract를 통과하고 10,000회 bootstrap이 끝나기 전에는 aggregate table을 동결하지 않는다. Active cluster의 partial output은 논문 수치가 아니다.

10EXTENSION

일반화 확장 — 모델과 도메인을 실제로 바꾼다

등록 값
Modelsraw sparse OLMoE-1B-7B-0125; raw dense Qwen2.5-14B
DomainsKnights-and-Knaves logic; ARC-Challenge science; non-math MMLU-Pro knowledge
Objective모든 cell에서 동일한 one-epoch verifier-reward GRPO
Scale2 models × 3 domains × 3 seeds × 4 checkpoints = 72 points
논문 그대로

메인 OLMo-3/math/code와 다른 architecture, capacity, domain, verifier에서 observed-stratum replication을 본다. 단, 두 확장 모델은 architecture와 capacity가 함께 달라지므로 각각의 인과 효과를 분리하지 못한다. Single-domain run이므로 cross-domain transfer, mixed-domain curriculum, alternative RLVR objective, full-parameter training, multilingual generalization은 주장하지 않는다.

11SCOPE

SFT는 어디에 남는가 — objective ablation만

논문 그대로

이전 positive-only response cross-entropy sweep는 RLVR이 아니라 SFT였다. 그 20개 run은 appendix의 supervised policy-shift ablation으로만 보존한다. GRPO checkpoint로 resume하거나, main regime table을 채우거나, RLVR 결과로 인용할 수 없다. LoRA는 trainable parameterization일 뿐 학습 objective가 아니다.

12RELATED

관련 연구에서의 위치

Data influence

CROPI · LESS · RFTInf

stored trajectory나 gradient alignment로 training data의 영향도를 추정한다

Off-policy correction

prefix · multi-step · trajectory ratios

history와 continuation을 서로 다른 방식으로 보정한다

RLVR selection

GradAlign · LearnAlign · DEPO

prompt utility, difficulty, uncertainty, curriculum을 사용한다

Replay and allocation

POPO · M2PO · Mu-GRPO · prompt replay

stale group 학습, prompt identity replay, rollout 배분은 ranking 보존과 다른 intervention이다

Noisy top-k

ranking under uncertainty · best-subset identification

global correlation보다 boundary gap과 reference noise가 결정을 지배한다

13LIMITATIONS

현재 상태와 한계

논문 그대로

이론은 exact unclipped population ratios와 unnormalized directional gradient를 다룬다. 실험은 capped ratios, CountSketch, cosine normalization을 쓰므로 정리를 그대로 구현한 것은 아니다. Main study는 한 raw 7B model, 두 domain, LoRA, one-epoch GRPO에 한정된다. Bootstrap interval은 fixed prompt pool과 fixed stale selector에 조건부이며, 실제 selected-subset training 뒤의 downstream reward는 아직 측정 대상이 아니다.

논문 그대로

현재 공개 가능한 결론은 theory와 protocol뿐이다. Main 40 points와 extension 72 points의 완전한 artifact gate가 끝나기 전에는 estimator 우열, drift boundary, model/domain generalization 수치를 제시하지 않는다.

TAKEAWAY

한 장으로 끝내는 정리

Recover

$\Delta_k>2e$

stale score error가 현재 top-k margin의 절반보다 작으면 exact set가 보존된다

Cannot certify from half

$g_{10}$ 또는 $g_{01}$ 단독

작은 KL과 unit ESS에서도 true ranking이 식별되지 않을 수 있다

Measure first

R/A/B split

reference가 재현될 때만 stale-vs-fresh comparison을 해석한다

Decide honestly

reuse · refresh · abstain

evidence에 따라 stale reuse, fresh sampling, 판단 보류 중 하나를 선택한다