Judging the value of a research publication, like determining the value of anything,
is a difficult problem with no known analytical solution (you can't write a formula to solve for it).
Over the years, the research community has established norms and heuristics for evaluating
research quality. Good research is generally highly cited including by well-known researchers in the field,
appears in prestigious venues or journals, and is authored by credentialed researchers with a strong track
record of publishing significant work.
These mechanisms are by no means iron-clad, and academics constantly debate and disagree about the value
of different papers and research directions.
Given the difficulty of deciding the value of human generated research, deciding the value of auto-generated research is a herculian challenge. Nevertheless, measuring and quantifying goodness of research is critical to improving auto-researcher performance.
Below we outline our best efforts to date to evaluate the quality of our research output. We will continually update this page as we refine our evaluation techniques.
Please note that certain evaluation details may be redacted for safety considerations (we'll explicitly mention these omissions).
We expect the auto-researcher to pass three performance thresholds, and propose different techniques for evaluating "goodness of research" at each performance tier.
We expect to begin at the sub-human level where the research output of the auto-researcher is unfit for submission to any journal or conference.
Move to human level, where the research output is on par with what a human researcher would produce, and finally reach a super-human level where the research output of the auto-researcher is superior to human output.
Tier 1Now
Sub-human
Research unfit for journal, workshop, or conference submission. Quality is measured using the Sakana reviewer and other internal benchmarks.
Tier 2
Human-level
Papers can be submitted to human journals. Quality is measured using citations, publication count in major journals, etc.
Tier 3
Super-human
Auto-research outputs and impact are accessed to be beyond any individual human researcher.
We currently assess the auto-researcher to be at Tier 1.
Of these three performance tiers, the human level researcher is the simplest to evaluate. In this tier, we propose to utilize the existing structure
of peer review to evaluate the quality of the research output. Submissions will be made to major journals,
workshops, and conferences to solicit peer-review and gauge goodness of research. We commit to submitting our manuscipts
responsibly with due notice to the venue about the exact level of human involvement in the submitted work. If a work does
clear human review, we propose to publicize our work for citation by others in the research community. Citation levels
of the auto-researcher can be used as an overall metric to judge impact.
Judging performance for the sub-human and super-human tiers is more challenging. For the sub-human tier we
propose using the automated reviewer open-sourced by Sakana AI (The AI Scientist (Lu, Lange,
Foerster, Clune & Ha, 2024), run verbatim from their released code) to evaluate goodness of research.
We propose to use a consistant score of 5.8+ against this reviewer (along with other consistant performance on other internal benchmarks) to indicate that our auto-researcher
has reach human-level performance.
For the super human tier, we propose using a combined h-index of 200+ for the auto-researcher to signify super-human performance. We currently have no strategy for evaluating
goodness of research at the super-human tier. While such considerations are not relevant at present, they may become vital to ensure that we build a researcher of sufficient quality
to "solve" the alignment problem.
To maintain the integrity of our reviewer we refrain from training directly against the Sakana reviewer and utilize other
techniques for benchmarking goodness of research which are not publically disclosed.
Below: Sakana scores for every graded publication, oldest to newest left to right. Each new
paper is scored as it ships and appended to the series.
Tier 1: Sub-human
While we remain in Tier 1, the stats and chart below track the Sakana automated reviewer scores of our research outputs.
We use these scores (alongside other internal metrics) as a proxy for research quality.
Program throughput
Paper scores only describe work that survived to publication. Run yield
also counts terminal attempts that failed, were cancelled, or stopped
early after a futility review.
68%gross yield · published / all terminal attempts
68%conditional yield · runs allowed to finish
0/65terminal runs stopped early as futile
3.9Sakana mean · last 3 graded papers
5.8ICLR accepted anchor · same reviewer & judge
0/46papers the reviewer would accept
Each dot is one graded paper from least to most recent. Dashed guides calibrate the
Sakana reviewer against real ICLR-2022 decisions. Accepted
submissions averaged 5.81, rejected 4.12.
Sustained scores at or above 5.8 on this scale
are one signal that auto-research outputs are ready for human peer review. Data: scores.csv.
All 46 published papers carry a Sakana score.
Earlier instruments (the in-house Tier-1 correctness grades and review
panel) remain in each paper's page data for continuity.