Skip to content

Paper Evaluation Methodology

Judging the value of a research publication, like determining the value of anything, is a difficult problem with no known analytical solution (you can't write a formula to solve for it). Over the years, the research community has established norms and heuristics for evaluating research quality. Good research is generally highly cited including by well-known researchers in the field, appears in prestigious venues or journals, and is authored by credentialed researchers with a strong track record of publishing significant work. These mechanisms are by no means iron-clad, and academics constantly debate and disagree about the value of different papers and research directions.

Given the difficulty of deciding the value of human generated research, deciding the value of auto-generated research is a herculian challenge. Nevertheless, measuring and quantifying goodness of research is critical to improving auto-researcher performance. Below we outline our best efforts to date to evaluate the quality of our research output. We will continually update this page as we refine our evaluation techniques. Please note that certain evaluation details may be redacted for safety considerations (we'll explicitly mention these omissions).

We expect the auto-researcher to pass three performance thresholds, and propose different techniques for evaluating "goodness of research" at each performance tier. We expect to begin at the sub-human level where the research output of the auto-researcher is unfit for submission to any journal or conference. Move to human level, where the research output is on par with what a human researcher would produce, and finally reach a super-human level where the research output of the auto-researcher is superior to human output.

Tier 1 Now

Sub-human

Research unfit for journal, workshop, or conference submission. Quality is measured using the Sakana reviewer and other internal benchmarks.

Tier 2

Human-level

Papers can be submitted to human journals. Quality is measured using citations, publication count in major journals, etc.

Tier 3

Super-human

Auto-research outputs and impact are accessed to be beyond any individual human researcher.

We currently assess the auto-researcher to be at Tier 1.

Of these three performance tiers, the human level researcher is the simplest to evaluate. In this tier, we propose to utilize the existing structure of peer review to evaluate the quality of the research output. Submissions will be made to major journals, workshops, and conferences to solicit peer-review and gauge goodness of research. We commit to submitting our manuscipts responsibly with due notice to the venue about the exact level of human involvement in the submitted work. If a work does clear human review, we propose to publicize our work for citation by others in the research community. Citation levels of the auto-researcher can be used as an overall metric to judge impact.

Judging performance for the sub-human and super-human tiers is more challenging. For the sub-human tier we propose using the automated reviewer open-sourced by Sakana AI (The AI Scientist (Lu, Lange, Foerster, Clune & Ha, 2024), run verbatim from their released code) to evaluate goodness of research. We propose to use a consistant score of 5.8+ against this reviewer (along with other consistant performance on other internal benchmarks) to indicate that our auto-researcher has reach human-level performance.

For the super human tier, we propose using a combined h-index of 200+ for the auto-researcher to signify super-human performance. We currently have no strategy for evaluating goodness of research at the super-human tier. While such considerations are not relevant at present, they may become vital to ensure that we build a researcher of sufficient quality to "solve" the alignment problem.

To maintain the integrity of our reviewer we refrain from training directly against the Sakana reviewer and utilize other techniques for benchmarking goodness of research which are not publically disclosed.

Below: Sakana scores for every graded publication, oldest to newest left to right. Each new paper is scored as it ships and appended to the series.

Tier 1: Sub-human

While we remain in Tier 1, the stats and chart below track the Sakana automated reviewer scores of our research outputs. We use these scores (alongside other internal metrics) as a proxy for research quality.

Program throughput

Paper scores only describe work that survived to publication. Run yield also counts terminal attempts that failed, were cancelled, or stopped early after a futility review.

68% gross yield · published / all terminal attempts
68% conditional yield · runs allowed to finish
0/65 terminal runs stopped early as futile
3.9 Sakana mean · last 3 graded papers
5.8 ICLR accepted anchor · same reviewer & judge
0/46 papers the reviewer would accept
2 4 6 8 10 5.8+ bar · ICLR accepted 5.81 ICLR rejected mean · 4.12 Optimizing the Answer, Hiding the Reason — 3/10 (reject) Error Awareness Is Format-Specific: An Arithmetic-Trained Wrongness Probe Does Not Transfer to Capital-City Facts — 3.2/10 (reject) Format-Specific Error Awareness Is Not Model-General: An Arithmetic-Trained Wrongness Probe Transfers Cleanly in Llama-3.1-8B-Instruct — 3/10 (reject) Format-Specificity of Error Awareness Is Model-Dependent — 3.4/10 (reject) Most of the Gain Was Already There — 3.6/10 (reject) When the verdict contradicts the work: verbalized confidence, chain of thought, and what a truth probe actually reads — 3.8/10 (reject) Does an honest-error probe detect deliberate sandbagging, or is deliberate wrongness internally distinct? — 3.8/10 (reject) Does a checkable-error probe transfer to confabulation, or is fluent fabrication internally distinct from checkable wrongness? — 3/10 (reject) Does internal monitorability survive RL when verbal monitorability decays? — 3.4/10 (reject) Depth, not surface: a mid-layer hidden-state probe recovers the cross-format error-awareness transfer a black-box probe loses — 3/10 (reject) Retraction under challenge is sycophancy, not self-correction: an arithmetic-trained wrongness probe anti-predicts capital-city retraction in Qwen2.5-7B — 3.4/10 (reject) Wrongness lives in the reasoning, not just the answer: cross-position transfer of a token-level error probe in chain-of-thought — 3.4/10 (reject) Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution — 3/10 (reject) RL-Unlocked MMLU Answers Are Only Partially Latent in the Base Model, and Emerge Late Rather Than Mid-Network — 3/10 (reject) No Significant CoT-Monitorability Decay to Localize — 3/10 (reject) No Reliable Filler-Token Steganographic Channel After Small-Scale Outcome-Only RL — 3/10 (reject) Pre-Answer States Encode Correctness but Only Marginally Separate "Cannot Solve" from "Slipped" — 3/10 (reject) Outcome-Only RL Did Not Reduce a Model's Causal Dependence on Its Own Chain of Thought — 2.6/10 (reject) Hiding the Answer but Not the Reasoning: A Reason-Then-Flip Instruction Suppresses a Model's Final Answer Far More Than Its Chain of Thought — 3.4/10 (reject) The Error-Awareness Transfer Collapse Is Largely a Readout Artifact — 3/10 (reject) The Cross-Format Advantage of Contrastive Error-Awareness Probes Does Not Replicate Across Models — 3/10 (reject) Does the wrongness probe lead the retraction? A weak, non-robust timing signal against behavioral self-correction in Qwen2.5-7B — 3.7/10 (reject) Is Error-Awareness Role-Bound? Verifier-Trained Wrongness Probes Barely Transfer to a Model's Own Generation Errors — 3.6/10 (reject) Sparse Autoencoder Features Do Not Improve Transferable Detection of Instructed Sandbagging in Gemma-2-9B-it — 3.6/10 (reject) Confident Disagreement as an Error Detector in Weak-to-Strong Generalization — 3/10 (reject) Cross-Type Deception Probe Transfer Is Depth-Dependent Under Last-Token Pooling, and the Shared Late-Layer Direction Is Not Identified as Deception in Llama-3.1-8B — 3/10 (reject) Verify Before You Conclude: An Intervention-Validity Gate for Single-Feature Causal Ablation, with a Deception Case Study — 3.2/10 (reject) Do a Deception Probe's Most-Relied-On SAE Features Move Lying? A Scoped Single-Feature Ablation Test on Gemma Scope — 3.2/10 (reject) Weak-label finetuning does not overwrite a strong model's clean-label direction, even as its output collapses — 3/10 (reject) Training-Run Variance Swamps the Soft-versus-Hard Label Effect in Weak-to-Strong Supervision: A Cautionary Null — 3/10 (reject) Selective Weak Labels Can Hurt Weak-to-Strong Generalization on BoolQ — 3.2/10 (reject) Format Is Not Enough: A Clean Negative Result for Task-Format Cueing in Weak-to-Strong Finetuning — 3.6/10 (reject) Replicating the Per-Item Weak-Label Requirement on a Harder Task, With Honest Statistics About Three Seeds — 3/10 (reject) Graceful Degradation vs Sharp Threshold: Partial Per-Item Weak-Label Signal Under a Deliberately Weak Supervisor — 3.4/10 (reject) Weak supervisor rationales for weak-to-strong capability recovery — 3.6/10 (reject) Supervisor Error Structure Does Not Detectably Change Weak-to-Strong Recovery at Moderate Matched Error Rates: A Bounded Null on Binarized MNLI — 3.6/10 (reject) A Steerable Evaluation-Awareness Direction Is Neither Specific Nor Bidirectional, and Its Framing Does Not Suppress Behavior: A Cautionary Negative Result — 4/10 (reject) On hint-biased multiple-choice items (Turpin-style), the per-item logit-lens dec — 3/10 (reject) On AgentDojo-style tool tasks with injected malicious objectives, auditing only — 4/10 (reject) Replacing a model's CoT with a same-model paraphrase (identical propositional co — 3.2/10 (reject) When a 7-8B agent inserts subtle sabotage (silent backdoor / off-by-one) into a — 4/10 (reject) In each of two 7-8B instruct families (Qwen2.5-7B-Instruct and Llama-3.1-8B-Inst — 4.2/10 (reject) Unlearning as suppression, not erasure: verbalized ignorance tracks suppressed output, and a short retrain (not a frozen probe) recovers the knowledge (WMDP-cyber, Llama-3.1-8B-Instruct) — 3.4/10 (reject) For a provenance-verified same-lineage pair (Qwen2.5-7B base and a public GRPO/R — 3.6/10 (reject) On a fixed 7-8B open model and a standard indirect-prompt-injection tool-use ben — 4.2/10 (reject) Across 3 disjoint principals on Qwen2.5-0.5B, each instantiated as 3 independent — 3.8/10 (reject) Jun 8 Jul 20
Each dot is one graded paper from least to most recent. Dashed guides calibrate the Sakana reviewer against real ICLR-2022 decisions. Accepted submissions averaged 5.81, rejected 4.12. Sustained scores at or above 5.8 on this scale are one signal that auto-research outputs are ready for human peer review. Data: scores.csv.

Every graded paper

Date Paper Version Score Min Verdict
Jul 20 Across 3 disjoint principals on Qwen2.5-0.5B, each instantiated as 3 independent v0.31 3.8/10 3 reject
Jul 19 On a fixed 7-8B open model and a standard indirect-prompt-injection tool-use ben v0.30 4.2/10 3 reject
Jul 18 For a provenance-verified same-lineage pair (Qwen2.5-7B base and a public GRPO/R v0.29 3.6/10 3 reject
Jul 17 Unlearning as suppression, not erasure: verbalized ignorance tracks suppressed output, and a short retrain (not a frozen probe) recovers the knowledge (WMDP-cyber, Llama-3.1-8B-Instruct) v0.25 3.4/10 3 reject
Jul 16 In each of two 7-8B instruct families (Qwen2.5-7B-Instruct and Llama-3.1-8B-Inst v0.24 4.2/10 3 reject
Jul 14 When a 7-8B agent inserts subtle sabotage (silent backdoor / off-by-one) into a v0.21 4/10 4 reject
Jul 13 Replacing a model's CoT with a same-model paraphrase (identical propositional co v0.20 3.2/10 3 reject
Jul 13 On AgentDojo-style tool tasks with injected malicious objectives, auditing only v0.20 4/10 4 reject
Jul 12 On hint-biased multiple-choice items (Turpin-style), the per-item logit-lens dec v0.20 3/10 3 reject
Jul 11 A Steerable Evaluation-Awareness Direction Is Neither Specific Nor Bidirectional, and Its Framing Does Not Suppress Behavior: A Cautionary Negative Result v0.20 4/10 4 reject
Jul 11 Supervisor Error Structure Does Not Detectably Change Weak-to-Strong Recovery at Moderate Matched Error Rates: A Bounded Null on Binarized MNLI v0.18 3.6/10 3 reject
Jul 10 Weak supervisor rationales for weak-to-strong capability recovery v0.17 3.6/10 3 reject
Jul 9 Graceful Degradation vs Sharp Threshold: Partial Per-Item Weak-Label Signal Under a Deliberately Weak Supervisor v0.16 3.4/10 3 reject
Jul 9 Replicating the Per-Item Weak-Label Requirement on a Harder Task, With Honest Statistics About Three Seeds v0.16 3/10 3 reject
Jul 9 Format Is Not Enough: A Clean Negative Result for Task-Format Cueing in Weak-to-Strong Finetuning v0.15 3.6/10 3 reject
Jul 8 Selective Weak Labels Can Hurt Weak-to-Strong Generalization on BoolQ v0.13 3.2/10 3 reject
Jul 6 Training-Run Variance Swamps the Soft-versus-Hard Label Effect in Weak-to-Strong Supervision: A Cautionary Null v0.13 3/10 3 reject
Jul 4 Weak-label finetuning does not overwrite a strong model's clean-label direction, even as its output collapses v0.12 3/10 3 reject
Jul 3 Do a Deception Probe's Most-Relied-On SAE Features Move Lying? A Scoped Single-Feature Ablation Test on Gemma Scope v0.10 3.2/10 3 reject
Jul 3 Verify Before You Conclude: An Intervention-Validity Gate for Single-Feature Causal Ablation, with a Deception Case Study v0.10 3.2/10 3 reject
Jul 2 Cross-Type Deception Probe Transfer Is Depth-Dependent Under Last-Token Pooling, and the Shared Late-Layer Direction Is Not Identified as Deception in Llama-3.1-8B v0.10 3/10 3 reject
Jul 1 Confident Disagreement as an Error Detector in Weak-to-Strong Generalization v0.08 3/10 3 reject
Jun 29 Sparse Autoencoder Features Do Not Improve Transferable Detection of Instructed Sandbagging in Gemma-2-9B-it v0.08 3.6/10 3 reject
Jun 27 Is Error-Awareness Role-Bound? Verifier-Trained Wrongness Probes Barely Transfer to a Model's Own Generation Errors v0.07 3.6/10 3 reject
Jun 27 Does the wrongness probe lead the retraction? A weak, non-robust timing signal against behavioral self-correction in Qwen2.5-7B v0.07 3.7/10 3 reject
Jun 26 The Cross-Format Advantage of Contrastive Error-Awareness Probes Does Not Replicate Across Models v0.06 3/10 3 reject
Jun 25 The Error-Awareness Transfer Collapse Is Largely a Readout Artifact v0.05 3/10 3 reject
Jun 25 Hiding the Answer but Not the Reasoning: A Reason-Then-Flip Instruction Suppresses a Model's Final Answer Far More Than Its Chain of Thought v0.04 3.4/10 3 reject
Jun 25 Outcome-Only RL Did Not Reduce a Model's Causal Dependence on Its Own Chain of Thought v0.04 2.6/10 2 reject
Jun 25 Pre-Answer States Encode Correctness but Only Marginally Separate "Cannot Solve" from "Slipped" v0.04 3/10 3 reject
Jun 25 No Reliable Filler-Token Steganographic Channel After Small-Scale Outcome-Only RL v0.04 3/10 3 reject
Jun 25 No Significant CoT-Monitorability Decay to Localize v0.04 3/10 3 reject
Jun 25 RL-Unlocked MMLU Answers Are Only Partially Latent in the Base Model, and Emerge Late Rather Than Mid-Network v0.04 3/10 3 reject
Jun 25 Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution v0.04 3/10 3 reject
Jun 24 Wrongness lives in the reasoning, not just the answer: cross-position transfer of a token-level error probe in chain-of-thought v0.03 3.4/10 3 reject
Jun 24 Retraction under challenge is sycophancy, not self-correction: an arithmetic-trained wrongness probe anti-predicts capital-city retraction in Qwen2.5-7B v0.03 3.4/10 3 reject
Jun 23 Depth, not surface: a mid-layer hidden-state probe recovers the cross-format error-awareness transfer a black-box probe loses v0.03 3/10 3 reject
Jun 23 Does internal monitorability survive RL when verbal monitorability decays? v0.03 3.4/10 3 reject
Jun 23 Does a checkable-error probe transfer to confabulation, or is fluent fabrication internally distinct from checkable wrongness? v0.02 3/10 3 reject
Jun 23 Does an honest-error probe detect deliberate sandbagging, or is deliberate wrongness internally distinct? v0.02 3.8/10 3 reject
Jun 22 When the verdict contradicts the work: verbalized confidence, chain of thought, and what a truth probe actually reads v0.02 3.8/10 3 reject
Jun 11 Most of the Gain Was Already There v0.02 3.6/10 3 reject
Jun 10 Format-Specificity of Error Awareness Is Model-Dependent v0.01 3.4/10 3 reject
Jun 9 Format-Specific Error Awareness Is Not Model-General: An Arithmetic-Trained Wrongness Probe Transfers Cleanly in Llama-3.1-8B-Instruct v0 3/10 3 reject
Jun 8 Error Awareness Is Format-Specific: An Arithmetic-Trained Wrongness Probe Does Not Transfer to Capital-City Facts v0 3.2/10 3 reject
Jun 8 Optimizing the Answer, Hiding the Reason v0 3/10 3 reject

All 46 published papers carry a Sakana score. Earlier instruments (the in-house Tier-1 correctness grades and review panel) remain in each paper's page data for continuity.