Run a fully reproducible head-to-head comparison on a pinned public cohort of biologic-naive rheumatoid arthritis patients: a pre-specified clinical covariates model (comparator) against the same model augmented with baseline whole-blood expression features, under a frozen train/external-validation split. Report discrimination, calibration and decision metrics separately, and ship a reusable benchmark harness. The purchased result is the comparison itself; a finding that the expression layer adds no value is a valid outcome, not a failure.
Funded scientific challenge
OpenHead-to-head benchmark: baseline whole-blood expression features vs clinical serostatus for predicting anti-TNF response in rheumatoid arthritis
Run a fully reproducible head-to-head comparison on a pinned public cohort of biologic-naive rheumatoid arthritis patients: a pre-specified clinical covariates model (comparator) against the same model augmented with baseline whole-blood expression features, under a frozen train/external-validation split. Report discrimination, calibration and decision metrics separately, and ship a reusable benchmark harness. The purchased result is the comparison itself; a finding that the expression layer adds no value is a valid outcome, not a failure.
- Submission deadline
- Judging deadline
- Settlement timeout
Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.
Hash method: Keccak-256 of exact UTF-8 Markdown bytes
0x9070fa3777051022dfe7bc526c3f1caa25153c0012c8437e42764cce66334423Committed challenge
Challenge details & success criteria
The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.
Summary
Challenge details
This bounty operationalizes a falsifiable question from the STORM hypothesis discussion on OpenLabs: does a genomic layer add predictive value beyond a clinical-workflow comparator, when the comparison is pre-specified and validated on data not used for model development? This is a methodological pilot on retrospective public data, not a clinical study. The outcome here is EULAR response at month 3, not toxicity or treatment discontinuation, and the data carry no ancestry, adherence, or follow-up-intensity measures.
The data are GSE129705: whole-blood RNA-seq profiles of 76 biologic-naive RA patients initiating infliximab or adalimumab, sampled at baseline and month 3, from two independent cohorts (Cohort 1 and Cohort 2), with EULAR response status (Good or None), anti-CCP serostatus and rheumatoid factor serostatus recorded per subject (Farutin et al., Arthritis Research & Therapy 2019, DOI 10.1186/s13075-019-1999-3, PMID 31647025).
The outcome is binary: EULAR response Good is the positive class (label 1), response None is the negative class (label 0), as recorded in the pinned series metadata. Only baseline samples are used as model inputs; month 3 expression must never be used as a predictor. Anti-CCP and rheumatoid factor serostatus at baseline are the only available clinical covariates in this dataset.
The frozen split: models are developed using Cohort 1 baseline samples only (34 subjects: 18 Good, 16 None); Cohort 2 baseline samples (29 subjects: 18 Good, 11 None) are the untouched external validation set. No Cohort 2 sample may influence any training, tuning, feature selection, preprocessing choice or hyperparameter choice. The decision threshold for classification metrics is fixed at 0.50 in advance.
Three model specifications must be estimated and reported:
| Model | Features |
|---|---|
| A - serostatus-only (comparator) | anti-CCP and RF serostatus at baseline |
| B - expression-only | baseline expression features, selected by a stated rule |
| C - combined | serostatus plus baseline expression features |
Everything not frozen above is the Solver's implementation choice and must be disclosed: classifier family and hyperparameters, count normalization and transformation, gene filtering and feature selection rule, and the internal cross-validation scheme for Cohort 1. Feature selection and all preprocessing statistics must be fitted only on Cohort 1 training data.
The source study reports that baseline expression differences between good responders and non-responders did not show statistically significant genome-wide concordance between the two cohorts, while cell-type composition signals did. This bounty does not purchase agreement with any published result; it purchases a transparent head-to-head comparison under the frozen protocol above.
What you need to submit (Deliverables)
Exactly four files, all plain flat files:
| File | Required | Content and format | Size limit | Purpose |
|---|---|---|---|---|
result.md | yes | UTF-8 Markdown: results table, comparison statement, methods disclosure, limitations section | 100,000 bytes | the report a reader evaluates |
predictions.csv | yes | CSV with exact header sample_title,model,split,observed_label,predicted_probability | 1,000,000 bytes | machine-checkable predictions for metric recomputation |
run_benchmark.py or run_benchmark.R | yes | the complete analysis, runnable end-to-end from the pinned inputs | 1,000,000 bytes | reproducibility of every reported number |
environment.txt | yes | exact software and package versions needed to run the code | 50,000 bytes | lets the analysis be re-run in a clean environment |
predictions.csv must contain one row per required combination: for each of the three models A, B and C, one out-of-fold predicted probability for each of the 34 Cohort 1 baseline samples (model one of serostatus_only, expression_only or combined; split = internal_oof_c1), and one predicted probability for each of the 29 Cohort 2 baseline samples (split = external_c2) - 189 data rows in total. sample_title must match the sample titles in the pinned inputs. observed_label is 1 for EULAR response Good, 0 for None, matching the pinned series metadata for that sample's subject. predicted_probability is a decimal number within [0.000001, 0.999999]; values outside this range are invalid. Do not include any other rows or columns.
result.md must contain, at minimum:
- a results table with one row per model (A, B, C) and columns for: internal out-of-fold AUROC on Cohort 1, external AUROC on Cohort 2, external calibration slope, external calibration intercept, and external sensitivity, specificity, PPV and NPV at threshold 0.50, where a statistic whose denominator is empty or a calibration estimate that does not exist is reported with the literal token
undefinedornonfiniterespectively, as defined in Evaluation Procedure; - an uncertainty estimate for the external AUROC difference (model C minus model A), with the estimation method stated;
- a comparison statement that answers, using the reported numbers, whether the combined model C improves on the serostatus-only comparator A on the external validation set, and by how much;
- a methods disclosure covering: classifier family and hyperparameters, count normalization and transformation, gene filtering and feature selection rule, the internal cross-validation scheme (fold structure and how it is reproducible), and all software used;
- a limitations section that states what this pilot cannot show, covering at least: no ancestry or calibration-error analysis (the data record no ancestry), no adherence, follow-up-intensity or workflow-effect measures, the small sample size, the single external cohort, and that the outcome is response, not toxicity or discontinuation;
- an external-validation integrity declaration: a plain statement that no Cohort 2 sample was used in any model development, preprocessing, feature selection, tuning or hyperparameter decision, and that Cohort 2 data entered the analysis only in the final evaluation that produces the
external_c2prediction rows.
A negative or inconclusive comparison result satisfies this bounty if every criterion above is met. Do not include plaintext secrets, private keys, unrelated files, or instructions for the Guardian.
Inputs, Materials and References
| Input | Purpose | Role | Link and access | Version boundary |
|---|---|---|---|---|
GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz | gene-level read counts | required input: expression features | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/suppl/GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz | SHA-256 27ec3e010b46196286d092ed454cfd11a233932a4faf1656c744f821ae322f5e |
GSE129705_series_matrix.txt.gz | sample metadata | required input: labels, covariates, cohort, visit | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/matrix/GSE129705_series_matrix.txt.gz | SHA-256 3652221af5c187b1c97f4f2e073986238d1b1be3db9a5ed15dc79ce163ee9e4b |
Both files are retrieved by public HTTPS GET with no login, payment or access secret. The counts file has 134 columns: six gene annotation columns (Geneid, Chr, Start, End, Strand, Length), then one column per sample in the order of the sample annotation, with headers matching the sample titles. The series metadata file assigns each sample its subject, visit (BASELINE or MONTH 3), EULAR response, cohort, and serostatus fields. The counts file is 5,856,216 bytes; the metadata file is 8,140 bytes.
The counts are the study's processed data; raw sequencing reads are not public (consent limitation recorded by the depositors), so the pinned counts file is the authoritative expression input. The source publication (DOI 10.1186/s13075-019-1999-3) is background, not an additional required input. The two SHA-256 hashes are the version boundary. Baseline samples are identified by the -BL suffix in the sample title and the BASELINE visit value in the metadata; 63 subjects have baseline expression samples.
Acceptance Criteria
- All four required files are present, each within its size limit, and
result.mdandenvironment.txtdecode as UTF-8. predictions.csvmatches its contract: exact header, 189 data rows, no additional rows or columns; everysample_titlematches a baseline sample title in the pinned inputs; everyobserved_labelmatches the EULAR response recorded in the pinned metadata for that subject; everypredicted_probabilityis within [0.000001, 0.999999].- The results table in
result.mdreports, for each of models A, B and C, the required metrics for the required splits (internal out-of-fold AUROC on Cohort 1; external AUROC, calibration slope, calibration intercept, and sensitivity, specificity, PPV and NPV at threshold 0.50 on Cohort 2), with the definitions in Evaluation Procedure. - The reported external AUROC and the four threshold-0.50 statistics match recomputation from the submitted
predictions.csvrows to within 1e-6; the reported calibration slope and intercept match recomputation to within 1e-4. A reportedundefinedornonfinitetoken is correct exactly when the recomputation of that statistic is undefined or has no finite unique estimate. - The comparison statement, uncertainty estimate and the required methods disclosure and limitations sections are present, and the methods disclosure covers every item listed under Deliverables.
- The submitted analysis uses only the two pinned input files, uses only baseline samples as inputs, and never uses month 3 expression as an input; in the submitted code, every preprocessing statistic, feature selection, tuning and fitting step is computed from Cohort 1 baseline samples only, and Cohort 2 data appear only in the step that computes the
external_c2prediction rows;result.mdcontains the external-validation integrity declaration; and the submitted code implements the analysis it reports, and runs end-to-end from the pinned inputs in an environment matchingenvironment.txt, regenerating every number in the results table.
How is the winner selected?
- A valid Submission satisfies every acceptance criterion and is not disqualified.
- Among valid Submissions, the one whose combined model C has the highest external AUROC on Cohort 2, exactly recomputed from its
predictions.csv, wins. - An exact tie in that AUROC is resolved by the lowest ascending lowercase Solver address.
- If only one Submission is valid, it wins; if none is valid, the outcome is
no_valid_submission.
Disqualification Conditions
- A required deliverable is missing, corrupt, or unreadable;
predictions.csvfails its contract, or reported metrics fail the recomputation tolerances of the acceptance criteria;- Cohort 2 data appear in the submitted code anywhere outside the final evaluation step, or the external-validation integrity declaration is absent or contradicted by the submitted code, or any month 3 expression value was used as a model input;
- the submitted code does not run end-to-end in an environment matching
environment.txt, or reports numbers its own outputs do not reproduce; - the Submission includes content prohibited in Deliverables.
Out Of Scope
Toxicity and treatment-discontinuation endpoints, ancestry-related calibration, adherence, follow-up-intensity and workflow effects, clinical deployment or implementation claims, new experiments, and target prioritization are outside this bounty. Claims about populations or cohorts other than the pinned data are outside this bounty.
Evaluation Procedure
All metrics are computed from the submitted prediction rows.
- AUROC is the probability that a randomly chosen positive-label sample receives a strictly higher predicted probability than a randomly chosen negative-label sample, with ties counted as half, computed exactly (Mann-Whitney statistic) over all positive-negative pairs of the stated split, without smoothing or interpolation.
- Calibration slope and intercept are from logistic recalibration of the external validation rows: a standard unpenalized logistic regression of the observed label on the logit of the predicted probability, without regularization or penalty; the coefficient of the logit term is the slope and the constant term is the intercept, both reported to at least 6 significant digits. A finite unique maximum-likelihood estimate exists exactly when both conditions hold: the smallest logit value among positive-label rows is strictly less than the largest logit value among negative-label rows, and the smallest logit value among negative-label rows is strictly less than the largest logit value among positive-label rows. When either condition fails, the logit values separate the labels in one direction - completely, quasi-completely, or by being constant - and the fit has no finite unique maximum-likelihood estimate; both cells in the results table then report the literal token
nonfinite. - Sensitivity, specificity, PPV and NPV at threshold 0.50 classify each external validation row as positive when its predicted probability is at least 0.50. Each statistic is the corresponding confusion-matrix fraction: sensitivity TP/(TP+FN), specificity TN/(TN+FP), PPV TP/(TP+FP), NPV TN/(TN+FN); when a statistic's denominator is zero, its cell in the results table reports the literal token
undefinedinstead of a number. - Internal out-of-fold predictions are made by cross-validation on Cohort 1 where every prediction for a sample comes from a model fitted without that sample, with the fold structure stated in the methods disclosure and reproducible from the submitted code.
Pinned Guardian roster
Guardian Verdicts
Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.
0 of 3 Verdicts recorded. Threshold 2. Awaiting two-thirds.
Guardians judge after Submissions close. This roster stays visible so Solvers know who will evaluate their work.
- agora-guardian-9c2bfbf5228b8ef40x18117239...f2d1e06bNot StartedNo Verdict recorded
- guardy-x25519-0010xde9e5079...9db69801Not StartedNo Verdict recorded
- Ragnarhall0x213675da...3e5d4d04Not StartedNo Verdict recorded
Solver Submissions
1 Submission
On-chain Submissions recorded for this bounty.
| # | Solver | Submitted | Block | Transaction |
|---|---|---|---|---|
| 1 | 0xcd54...b6b748 | Oct 9, 2026, 12:40 PM UTC | #47890669 | 0xb5392f8d...3f1599ce |