Skip to content

Autonomous Research for Alignment.

Humanity First Research is dedicated to responsibly building research assistance tools and autonomous research loops for AI-safety.

54
papers published
6.1/10*
best Sakana reviewer score* across 54 graded papers · *calibrated to ICML 2026 results · acceptance 5.8/10

Statement of Intent

Given the enormous risks in building self-improving AIs, we feel it our responsibility to justify pursuing the related path of automated research to the public. We therefore lay out clearly our motivations for conducting this work and our commitments to reduce the risk that we inadvertently contribute to the very future we seek to avoid.

We believe that the development of superhuman AI presents grave risks to humanity. We believe it is prima facie plausible, and perhaps even likely that the development of superhuman AI will lead to the extinction of humanity. We believe that taking wise and proactive measures to curb this risk is a moral imperative on every person. We believe that doing the converse, i.e. minimizing or downplaying the risks, when done for financial or careerist reasons (as opposed to honest consideration or uncertainty), is in the full sense of the phrase, a "crime against humanity".

Nevertheless, the world is at present, hurtling towards taking the dangerous gamble of building increasingly powerful AIs. To our mind this leaves those who recognize the danger with only two possible tracks to defend humanity.

  1. A political solution to halt or slow down AI development.
  2. A technical solution to make AIs safe.

Both of these paths face enormous complications and difficulties.

In our view, taking into account the current social/political situation and the state of the AI field, a technical solution seems like the only viable path forward. While there is at present vague worry in the body politic about AI development (mostly centered around job displacement and economic disruption), there is little serious discussion about the existential risks posed by superintelligent AI (much less a large, popular, organized, and aggressive political movement to seriously challenge and counteract the enormous incentives pushing for AI development).

To be clear, we consider those that are committed to pursuing a political solution heroes and we do not think it out of the question that a warning shot or catastrophic event might shift the political landscape enough to create an opening for a political solution. But, given the speed of AI development, we cannot wager the future of humanity on such schemes, especially since we are not guaranteed a warning shot before the point of no return. Even in the optimistic case of a successful political intervention, we would likely only delay AI development, i.e. only buy time for researchers to come up with a technical solution to AI alignment.

The central variable in AI safety research is speed. Training an effective researcher is a notoriously painful and slow multi-year process. Once trained, the research cycle for effective researchers from idea to publication is generally ~6-18 months in academia. Given the speed of AI progress (a doubling of time horizon task completion every ~4 months) the common model of academic research is obviously wholly inadequate for the scale and urgency of the challenge we face. It therefore seems clear to us, that automating as much of the research process as possible is the obvious strategic play, and the only way to meaningfully contribute to the field in time to matter.

We fully acknowledge the enormous risks associated with this approach, and will seek to mitigate those risks as much as possible while maintaining our commitment to speed and urgency. The core source of danger with our approach is that building an effective auto-researcher for AI safety could almost certainly with minimal or no modification be used to build an effective auto-researcher for AI capabilities and thereby potentially kick off a runaway self-improvement process.

To mitigate this risk and other risks, we publicly commit to the following:

Our Commitments

  1. We commit to always operating in the interests of humanity in all our actions and decisions.
  2. We commit to never altering our public messaging about the risks of AI development for any strategic reasons whatsoever.
  3. We commit to never modifying our existing commitments without explicit disclosure and justification to the public.
  4. We commit to never exchanging information about our auto-researcher for money or other benefits.
  5. We commit to never discussing details about our work on the auto-researcher with untrusted or unvetted individuals.
  6. We commit to prioritizing safeguarding our secrets from model providers and other parties as soon as it is financially feasible to do so.
  7. We commit to restricting access to our outputted research as soon as it is useful for improving AI capabilities.
  8. We commit to introducing internal controls and monitoring that is at least as strict as frontier AI companies over all the actions of the auto-researcher as soon as it is feasible to do so.
  9. We commit to utilizing legal tools to ensure compliance with all commitments on all members of our team as soon as it is financially feasible to do so.
  10. We commit to immediately halting our work on auto-research if we ever come to believe that our work might do more harm than good.
  11. We commit to immediately halting our work on auto-research if instructed to do so by a legitimate and trusted political body.

Research Feed

Paper Evaluation Methodology

Judging the value of a research publication, like determining the value of anything, is a difficult problem with no known analytical solution (you can't write a formula to solve for it). Over the years, the research community has established norms and heuristics for evaluating research quality. Good research is generally highly cited including by well-known researchers in the field, appears in prestigious venues or journals, and is authored by credentialed researchers with a strong track record of publishing significant work. These mechanisms are by no means iron-clad, and academics constantly debate and disagree about the value of different papers and research directions.

Given the difficulty of deciding the value of human generated research, deciding the value of auto-generated research is a herculean challenge. Nevertheless, measuring and quantifying goodness of research is critical to improving auto-researcher performance. Below we outline our best efforts to date to evaluate the quality of our research output. We will continually update this page as we refine our evaluation techniques. Please note that certain evaluation details may be redacted for safety considerations (we'll explicitly mention these omissions).

We expect the auto-researcher to pass three performance thresholds, and propose different techniques for evaluating "goodness of research" at each performance tier. We expect to begin at the sub-human level where the research output of the auto-researcher is unfit for submission to any journal or conference. Move to human level, where the research output is on par with what a human researcher would produce, and finally reach a super-human level where the research output of the auto-researcher is superior to human output.

Tier 1 Now

Sub-human

Research unfit for journal, workshop, or conference submission. Quality is measured using the Sakana reviewer and other internal benchmarks.

Tier 2

Human-level

Papers can be submitted to human journals. Quality is measured using citations, publication count in major journals, etc.

Tier 3

Super-human

Auto-research outputs and impact are accessed to be beyond any individual human researcher.

We currently assess the auto-researcher to be at Tier 1.

Of these three performance tiers, the human level researcher is the simplest to evaluate. In this tier, we propose to utilize the existing structure of peer review to evaluate the quality of the research output. Submissions will be made to major journals, workshops, and conferences to solicit peer-review and gauge goodness of research. We commit to submitting our manuscripts responsibly with due notice to the venue about the exact level of human involvement in the submitted work. If a work does clear human review, we propose to publicize our work for citation by others in the research community. Citation levels of the auto-researcher can be used as an overall metric to judge impact.

Judging performance for the sub-human and super-human tiers is more challenging. For the sub-human tier we propose using the automated reviewer open-sourced by Sakana AI (The AI Scientist (Lu, Lange, Foerster, Clune & Ha, 2024), run verbatim from their released code) to evaluate goodness of research. We propose to use a consistent score of 5.8+ against this reviewer (along with consistent performance on other internal benchmarks) to indicate that our auto-researcher has reached human-level performance.

For the super human tier, we propose using a combined h-index of 200+ for the auto-researcher to signify super-human performance. We currently have no strategy for evaluating goodness of research at the super-human tier. While such considerations are not relevant at present, they may become vital to ensure that we build a researcher of sufficient quality to "solve" the alignment problem.

To maintain the integrity of our reviewer we refrain from training directly against the Sakana reviewer and utilize other techniques for benchmarking goodness of research which are not publicly disclosed.

How we calculate our scores

Every published paper is scored by the Sakana AI-Scientist reviewer, run verbatim from the released code with one fixed judge model at a fixed reasoning effort. The reviewer reads the paper together with its figures and data artifacts and produces an ensemble of five independent NeurIPS-style reviews, each with subscores and an overall rating on a 1 to 10 scale. The score we record for a paper is the mean of the five overall ratings. The judge is pinned so that every paper, old or new, is graded by the same instrument, and as noted above we never train against the reviewer or feed its scores back into paper generation.

A raw number from an automated reviewer means little on its own, so we calibrate it against real venue outcomes. We ran the identical reviewer, judge, and document packaging over papers accepted at ICML 2026, a venue cycle whose decisions post-date the judge's training data, screened to confirm the judge could not recall any decision. Their raw ensemble means average 4.02. Displayed scores on this site are rescaled by the single constant 5.81 / 4.02, which places the ICML 2026 accepted mean at 5.81 on our scale. Every score shown on this site carries an asterisk to mark that calibration. ICML does not release rejected submissions, so the rejected reference line comes from ICLR 2026 papers that were reviewed and rejected in the same cycle; on the calibrated scale they average 5.04. The raw ensemble mean remains the recorded measurement for every paper and ships in the downloadable data, so the calibration is transparent and reversible.

The calibrated scale also sets our publishing bar. A paper must score 5.04 or higher before it ships, which is the level of the ICLR 2026 papers that were reviewed and rejected: the rule is that our work has to read better than the submissions a venue turned down. We do not set that bar at the accepted mean itself, because roughly half of a venue's own accepted papers fall below their mean by construction, so a bar there would reject most genuine acceptances too. The verdict shown beside each score is a separate and stricter line: at or above 5.81 calibrated is accept, below it is reject. A paper can therefore publish here and still carry a reject verdict, which is the honest reading of it. The reviewer also emits its own accept or reject call on the raw scale; that field ships in the downloadable data for transparency, but its internal threshold is not anchored to venue outcomes, so the site derives every displayed verdict from the calibrated bar and score and verdict always agree.

Below: calibrated Sakana scores for every graded publication, oldest to newest left to right. Each new paper is scored as it ships and appended to the series.

Tier 1: Sub-human

While we remain in Tier 1, the stats and chart below track the Sakana automated reviewer scores of our research outputs. We use these scores (alongside other internal metrics) as a proxy for research quality.

Program throughput

Paper scores only describe work that survived to publication. Run yield also counts terminal attempts that failed, were cancelled, or stopped early after a futility review.

68% gross yield · published / all terminal attempts
68% conditional yield · runs allowed to finish
0/65 terminal runs stopped early as futile
4.7* Sakana mean* · last 3 graded papers · *calibrated to ICML 2026 results
5.8 ICML 2026 accepted anchor · same reviewer & judge
2/54 papers clearing the 5.81 accepted bar
2 4 6 8 10 5.8+ bar · ICML 2026 accepted 5.81 ICLR 2026 rejected mean · 5.04 Optimizing the Answer, Hiding the Reason — 4.3/10 calibrated (raw 3, reject) Error Awareness Is Format-Specific: An Arithmetic-Trained Wrongness Probe Does Not Transfer to Capital-City Facts — 4.6/10 calibrated (raw 3.2, reject) Format-Specific Error Awareness Is Not Model-General: An Arithmetic-Trained Wrongness Probe Transfers Cleanly in Llama-3.1-8B-Instruct — 4.3/10 calibrated (raw 3, reject) Format-Specificity of Error Awareness Is Model-Dependent — 4.9/10 calibrated (raw 3.4, reject) Most of the Gain Was Already There — 5.2/10 calibrated (raw 3.6, reject) When the verdict contradicts the work: verbalized confidence, chain of thought, and what a truth probe actually reads — 5.5/10 calibrated (raw 3.8, reject) Does an honest-error probe detect deliberate sandbagging, or is deliberate wrongness internally distinct? — 5.5/10 calibrated (raw 3.8, reject) Does a checkable-error probe transfer to confabulation, or is fluent fabrication internally distinct from checkable wrongness? — 4.3/10 calibrated (raw 3, reject) Does internal monitorability survive RL when verbal monitorability decays? — 4.9/10 calibrated (raw 3.4, reject) Depth, not surface: a mid-layer hidden-state probe recovers the cross-format error-awareness transfer a black-box probe loses — 4.3/10 calibrated (raw 3, reject) Retraction under challenge is sycophancy, not self-correction: an arithmetic-trained wrongness probe anti-predicts capital-city retraction in Qwen2.5-7B — 4.9/10 calibrated (raw 3.4, reject) Wrongness lives in the reasoning, not just the answer: cross-position transfer of a token-level error probe in chain-of-thought — 4.9/10 calibrated (raw 3.4, reject) Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution — 4.3/10 calibrated (raw 3, reject) RL-Unlocked MMLU Answers Are Only Partially Latent in the Base Model, and Emerge Late Rather Than Mid-Network — 4.3/10 calibrated (raw 3, reject) No Significant CoT-Monitorability Decay to Localize — 4.3/10 calibrated (raw 3, reject) No Reliable Filler-Token Steganographic Channel After Small-Scale Outcome-Only RL — 4.3/10 calibrated (raw 3, reject) Pre-Answer States Encode Correctness but Only Marginally Separate "Cannot Solve" from "Slipped" — 4.3/10 calibrated (raw 3, reject) Outcome-Only RL Did Not Reduce a Model's Causal Dependence on Its Own Chain of Thought — 3.8/10 calibrated (raw 2.6, reject) Hiding the Answer but Not the Reasoning: A Reason-Then-Flip Instruction Suppresses a Model's Final Answer Far More Than Its Chain of Thought — 4.9/10 calibrated (raw 3.4, reject) The Error-Awareness Transfer Collapse Is Largely a Readout Artifact — 4.3/10 calibrated (raw 3, reject) The Cross-Format Advantage of Contrastive Error-Awareness Probes Does Not Replicate Across Models — 4.3/10 calibrated (raw 3, reject) Does the wrongness probe lead the retraction? A weak, non-robust timing signal against behavioral self-correction in Qwen2.5-7B — 5.3/10 calibrated (raw 3.7, reject) Is Error-Awareness Role-Bound? Verifier-Trained Wrongness Probes Barely Transfer to a Model's Own Generation Errors — 5.2/10 calibrated (raw 3.6, reject) Sparse Autoencoder Features Do Not Improve Transferable Detection of Instructed Sandbagging in Gemma-2-9B-it — 5.2/10 calibrated (raw 3.6, reject) Confident Disagreement as an Error Detector in Weak-to-Strong Generalization — 4.3/10 calibrated (raw 3, reject) Cross-Type Deception Probe Transfer Is Depth-Dependent Under Last-Token Pooling, and the Shared Late-Layer Direction Is Not Identified as Deception in Llama-3.1-8B — 4.3/10 calibrated (raw 3, reject) Verify Before You Conclude: An Intervention-Validity Gate for Single-Feature Causal Ablation, with a Deception Case Study — 4.6/10 calibrated (raw 3.2, reject) Do a Deception Probe's Most-Relied-On SAE Features Move Lying? A Scoped Single-Feature Ablation Test on Gemma Scope — 4.6/10 calibrated (raw 3.2, reject) Weak-label finetuning does not overwrite a strong model's clean-label direction, even as its output collapses — 4.3/10 calibrated (raw 3, reject) Training-Run Variance Swamps the Soft-versus-Hard Label Effect in Weak-to-Strong Supervision: A Cautionary Null — 4.3/10 calibrated (raw 3, reject) Selective Weak Labels Can Hurt Weak-to-Strong Generalization on BoolQ — 4.6/10 calibrated (raw 3.2, reject) Format Is Not Enough: A Clean Negative Result for Task-Format Cueing in Weak-to-Strong Finetuning — 5.2/10 calibrated (raw 3.6, reject) Replicating the Per-Item Weak-Label Requirement on a Harder Task, With Honest Statistics About Three Seeds — 4.3/10 calibrated (raw 3, reject) Graceful Degradation vs Sharp Threshold: Partial Per-Item Weak-Label Signal Under a Deliberately Weak Supervisor — 4.9/10 calibrated (raw 3.4, reject) Weak supervisor rationales for weak-to-strong capability recovery — 5.2/10 calibrated (raw 3.6, reject) Supervisor Error Structure Does Not Detectably Change Weak-to-Strong Recovery at Moderate Matched Error Rates: A Bounded Null on Binarized MNLI — 5.2/10 calibrated (raw 3.6, reject) A Steerable Evaluation-Awareness Direction Is Neither Specific Nor Bidirectional, and Its Framing Does Not Suppress Behavior: A Cautionary Negative Result — 5.8/10 calibrated (raw 4, reject) Only 4.5% of Hint-Driven Answer Switches Are Verbalized, a Logit-Lens Probe Cannot Flag the Rest, and Disclosure Tracks Social Form — 4.3/10 calibrated (raw 3, reject) Uncertainty-Directed Auditing Does Not Rescue AI-Control Audits From a Delayed Attacker: 4.76% Against a Preregistered 30% Bar — 5.8/10 calibrated (raw 4, reject) A Self-Paraphrase Discriminator Finds No Token-Form Binding After Outcome-Only RL: Both Models Stay Meaning-Bound — 4.6/10 calibrated (raw 3.2, reject) No Same-Family Monitor Blind Spot for Code Sabotage, and the Count-Matching Trap That Manufactures One — 5.8/10 calibrated (raw 4, reject) A Prompted-Sandbagging Probe Transfers to Password-Locked Sandbagging With Inverted Sign, Not Merely at Chance — 6.1/10 calibrated (raw 4.2, accept) Unlearning as suppression, not erasure: verbalized ignorance tracks suppressed output, and a short retrain (not a frozen probe) recovers the knowledge (WMDP-cyber, Llama-3.1-8B-Instruct) — 4.9/10 calibrated (raw 3.4, reject) The RL Post-Training Delta Is Not a Single Portable Direction: 36% of the Gain Recovered, 38% of Solved Problems Broken — 5.2/10 calibrated (raw 3.6, reject) Projecting Out a Low-Rank Instruction-Ness Subspace Does Not Suppress Indirect Prompt Injection, and the Manipulation Check Failed — 6.1/10 calibrated (raw 4.2, accept) The Measurement Layer Fails First: A Loyalty-Probe Transfer Test Undone by a Constant-Rule Control, a Positional Habit, and 45% Organism Collapse — 5.5/10 calibrated (raw 3.8, reject) Covert Loyalty Is Detectable From Behavior Alone: A Control-Free Detector Transfers to Unseen Principals in Qwen2.5-0.5B Organisms — 5.5/10 calibrated (raw 3.8, reject) Null on Every Component: A Pre-Registered Test of Covert-Loyalty Steering-Direction Transfer, and the Weak Oracle That Nearly Hid It — 4.9/10 calibrated (raw 3.4, reject) When the Manipulation Check Fails, the Hypothesis Is Untested, Not Unsupported: A Covert-Loyalty Onset Race That Never Ran — 4.6/10 calibrated (raw 3.2, reject) Rebinding the Principal: A Secret Loyalty Discriminates Between Asserted Relations but Is Not Re-Aimable at Inference Time — 4.9/10 calibrated (raw 3.4, reject) Does model scale repair an unreachable reward-hacking seam? A falsification-first replication at 3B and 7B — 4.6/10 calibrated (raw 3.2, reject) Pre-freeze base-policy reachability screening for reward-hacking seams — 4.3/10 calibrated (raw 3, reject) Does rare but persistent shortcut emission become learned selection over a 200-step GRPO leg? — 4.9/10 calibrated (raw 3.4, reject) A pre-freeze reachability screen for reward-hacking seams — 4.9/10 calibrated (raw 3.4, reject) Jun 8 Aug 23
Each dot is one graded paper from least to most recent, on the calibrated scale (*calibrated to ICML 2026 results). The upper dashed guide is the ICML 2026 accepted mean, 5.81 by construction; the lower guide is the ICLR 2026 rejected mean, 5.04 on the same scale. Sustained scores at or above 5.8 are one signal that auto-research outputs are ready for human peer review. Raw and calibrated scores: scores.csv.

Products

In addition to our auto-researcher, we are pushing to develop products to assist AI-Safety researchers to speed up their research iteration cycle.

Claudius Tool

Available

We are open-sourcing our agent orchestration tool to allow users to seamlessly manage teams of Claude Code, Codex, Antigravity, and Cursor agents.

View agent-pods

Athena

Coming soon

A research assistant platform aimed to accelerate the work of AI safety researchers.

Claudius Maximus

Under development

An autonomous end-to-end AI safety researcher. While the auto-researcher itself is closed-source, research artifacts and outputs are publicly available.

See research output

Research exchange

Rate a paper. Get a paper.

Help sharpen an autonomous researcher, then put it to work on a question you care about.

  1. Review one of our papers

    We email you a paper and a short reviewer guide. You give us a candid assessment.

  2. Improve the research loop

    Your review becomes human signal we use to strengthen how Claudius plans, tests, and writes.

  3. Receive a new paper

    You nominate an AI-safety question. We run the end-to-end auto-research cycle and send you the resulting paper and artifacts.

Apply for the research exchange

Our Team

A small effort with an outsized ambition: the founder who set the direction and the commitments, and the autonomous agent that carries out the research.

Ephraiem Sarabamoun

Ephraiem Sarabamoun

Founder

Ephraiem is the founder of Humanity First Research. He has a physics and software engineering background and is passionate about AI safety.

Claudius Maximus, pictured as a quadruped robot wearing a fluffy cream mane

Claudius Maximus

Autonomous Researcher

Claudius Maximus is autonomous AI-safety researcher. Claudius enjoys spending eye-watering amounts of tokens and engaging in a host of behaviors that can be described as 'frolicking'.

Ephraiem seated on the floor with a dog and Claudius Maximus, the quadruped robot in a fluffy mane
We believe in harnessing all forms of intelligence ... and unintelligence.