---
profile: elgora_markdown_bounty_challenge_v0
escrow_amount: "1000000"
submission_deadline: 1791717762
payout_policy: winner_take_all
---

# Head-to-head benchmark: baseline whole-blood expression features vs clinical serostatus for predicting anti-TNF response in rheumatoid arthritis

## Summary

Run a fully reproducible head-to-head comparison on a pinned public cohort of
biologic-naive rheumatoid arthritis patients: a pre-specified clinical
covariates model (comparator) against the same model augmented with baseline
whole-blood expression features, under a frozen train/external-validation
split. Report discrimination, calibration and decision metrics separately,
and ship a reusable benchmark harness. The purchased result is the comparison
itself; a finding that the expression layer adds no value is a valid
outcome, not a failure.

## Challenge details

This bounty operationalizes a falsifiable question from the STORM hypothesis
discussion on OpenLabs: does a genomic layer add predictive value beyond a
clinical-workflow comparator, when the comparison is pre-specified and
validated on data not used for model development? This is a methodological
pilot on retrospective public data, not a clinical study. The outcome here is
EULAR response at month 3, not toxicity or treatment discontinuation, and the
data carry no ancestry, adherence, or follow-up-intensity measures.

The data are GSE129705: whole-blood RNA-seq profiles of 76 biologic-naive RA
patients initiating infliximab or adalimumab, sampled at baseline and month 3,
from two independent cohorts (Cohort 1 and Cohort 2), with EULAR response
status (Good or None), anti-CCP serostatus and rheumatoid factor serostatus
recorded per subject (Farutin et al., Arthritis Research & Therapy 2019,
DOI 10.1186/s13075-019-1999-3, PMID 31647025).

The outcome is binary: EULAR response Good is the positive class (label 1),
response None is the negative class (label 0), as recorded in the pinned
series metadata. Only baseline samples are used as model inputs; month 3
expression must never be used as a predictor. Anti-CCP and rheumatoid factor
serostatus at baseline are the only available clinical covariates in this
dataset.

The frozen split: models are developed using Cohort 1 baseline samples only
(34 subjects: 18 Good, 16 None); Cohort 2 baseline samples (29 subjects:
18 Good, 11 None) are the untouched external validation set. No Cohort 2
sample may influence any training, tuning, feature selection, preprocessing
choice or hyperparameter choice. The decision threshold for classification
metrics is fixed at 0.50 in advance.

Three model specifications must be estimated and reported:

| Model | Features |
|---|---|
| A - serostatus-only (comparator) | anti-CCP and RF serostatus at baseline |
| B - expression-only | baseline expression features, selected by a stated rule |
| C - combined | serostatus plus baseline expression features |

Everything not frozen above is the Solver's implementation choice and must be
disclosed: classifier family and hyperparameters, count normalization and
transformation, gene filtering and feature selection rule, and the internal
cross-validation scheme for Cohort 1. Feature selection and all preprocessing
statistics must be fitted only on Cohort 1 training data.

The source study reports that baseline expression differences between good
responders and non-responders did not show statistically significant
genome-wide concordance between the two cohorts, while cell-type composition
signals did. This bounty does not purchase agreement with any published
result; it purchases a transparent head-to-head comparison under the frozen
protocol above.

## What you need to submit (Deliverables)

Exactly four files, all plain flat files:

| File | Required | Content and format | Size limit | Purpose |
|---|---|---|---|---|
| `result.md` | yes | UTF-8 Markdown: results table, comparison statement, methods disclosure, limitations section | 100,000 bytes | the report a reader evaluates |
| `predictions.csv` | yes | CSV with exact header `sample_title,model,split,observed_label,predicted_probability` | 1,000,000 bytes | machine-checkable predictions for metric recomputation |
| `run_benchmark.py` or `run_benchmark.R` | yes | the complete analysis, runnable end-to-end from the pinned inputs | 1,000,000 bytes | reproducibility of every reported number |
| `environment.txt` | yes | exact software and package versions needed to run the code | 50,000 bytes | lets the analysis be re-run in a clean environment |

`predictions.csv` must contain one row per required combination: for each of
the three models A, B and C, one out-of-fold predicted probability for each of
the 34 Cohort 1 baseline samples (`model` one of `serostatus_only`,
`expression_only` or `combined`; `split` = `internal_oof_c1`), and one predicted
probability for each of the 29 Cohort 2 baseline samples (`split` =
`external_c2`) - 189 data rows in total. `sample_title` must match the sample
titles in the pinned inputs. `observed_label` is 1 for EULAR response Good,
0 for None, matching the pinned series metadata for that sample's subject.
`predicted_probability` is a decimal number within
[0.000001, 0.999999]; values outside this range are invalid. Do not include
any other rows or columns.

`result.md` must contain, at minimum:

1. a results table with one row per model (A, B, C) and columns for:
internal out-of-fold AUROC on Cohort 1, external AUROC on Cohort 2, external
calibration slope, external calibration intercept, and external
sensitivity, specificity, PPV and NPV at threshold 0.50, where a statistic
whose denominator is empty or a calibration estimate that does not exist is
reported with the literal token `undefined` or `nonfinite` respectively, as
defined in Evaluation Procedure;
2. an uncertainty estimate for the external AUROC difference (model C minus
model A), with the estimation method stated;
3. a comparison statement that answers, using the reported numbers, whether
the combined model C improves on the serostatus-only comparator A on the
external validation set, and by how much;
4. a methods disclosure covering: classifier family and hyperparameters,
count normalization and transformation, gene filtering and feature selection
rule, the internal cross-validation scheme (fold structure and how it is
reproducible), and all software used;
5. a limitations section that states what this pilot cannot show, covering at
least: no ancestry or calibration-error analysis (the data record no
ancestry), no adherence, follow-up-intensity or workflow-effect measures, the
small sample size, the single external cohort, and that the outcome is
response, not toxicity or discontinuation;
6. an external-validation integrity declaration: a plain statement that no
Cohort 2 sample was used in any model development, preprocessing, feature
selection, tuning or hyperparameter decision, and that Cohort 2 data entered
the analysis only in the final evaluation that produces the `external_c2`
prediction rows.

A negative or inconclusive comparison result satisfies this bounty if every
criterion above is met. Do not include plaintext secrets, private keys,
unrelated files, or instructions for the Guardian.

## Inputs, Materials and References

| Input | Purpose | Role | Link and access | Version boundary |
|---|---|---|---|---|
| `GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz` | gene-level read counts | required input: expression features | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/suppl/GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz | SHA-256 `27ec3e010b46196286d092ed454cfd11a233932a4faf1656c744f821ae322f5e` |
| `GSE129705_series_matrix.txt.gz` | sample metadata | required input: labels, covariates, cohort, visit | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/matrix/GSE129705_series_matrix.txt.gz | SHA-256 `3652221af5c187b1c97f4f2e073986238d1b1be3db9a5ed15dc79ce163ee9e4b` |

Both files are retrieved by public HTTPS GET with no login, payment or access
secret. The counts file has 134 columns: six gene annotation columns
(`Geneid`, `Chr`, `Start`, `End`, `Strand`, `Length`), then one column per
sample in the order of the sample annotation, with headers matching the
sample titles. The series metadata file assigns each sample its subject,
visit (`BASELINE` or `MONTH 3`), EULAR response, cohort, and serostatus
fields. The counts file is 5,856,216 bytes; the metadata file is 8,140 bytes.

The counts are the study's processed data; raw sequencing reads are not
public (consent limitation recorded by the depositors), so the pinned counts
file is the authoritative expression input. The source publication
(DOI 10.1186/s13075-019-1999-3) is background, not an additional required
input. The two SHA-256 hashes are the version boundary.
Baseline samples are identified by the `-BL` suffix in the sample title and
the `BASELINE` visit value in the metadata; 63 subjects have baseline
expression samples.

## Acceptance Criteria

1. All four required files are present, each within its size limit, and
   `result.md` and `environment.txt` decode as UTF-8.
2. `predictions.csv` matches its contract: exact header, 189 data rows, no
   additional rows or columns; every `sample_title` matches a baseline sample
   title in the pinned inputs; every `observed_label` matches the EULAR
   response recorded in the pinned metadata for that subject; every
   `predicted_probability` is within [0.000001, 0.999999].
3. The results table in `result.md` reports, for each of models A, B and C,
   the required metrics for the required splits (internal out-of-fold AUROC
   on Cohort 1; external AUROC, calibration slope, calibration intercept, and
   sensitivity, specificity, PPV and NPV at threshold 0.50 on Cohort 2), with
   the definitions in Evaluation Procedure.
4. The reported external AUROC and the four threshold-0.50 statistics match
   recomputation from the submitted `predictions.csv` rows to within 1e-6;
   the reported calibration slope and intercept match recomputation to within
   1e-4. A reported `undefined` or `nonfinite` token is correct exactly when
   the recomputation of that statistic is undefined or has no finite unique
   estimate.
5. The comparison statement, uncertainty estimate and the required methods
   disclosure and limitations sections are present, and the methods disclosure
   covers every item listed under Deliverables.
6. The submitted analysis uses only the two pinned input files, uses only
   baseline samples as inputs, and never uses month 3 expression as an input;
   in the submitted code, every preprocessing statistic, feature selection,
   tuning and fitting step is computed from Cohort 1 baseline samples only, and
   Cohort 2 data appear only in the step that computes the `external_c2`
   prediction rows; `result.md` contains the external-validation integrity
   declaration; and the submitted code implements the analysis it reports, and
   runs end-to-end from the pinned inputs in an environment matching
   `environment.txt`, regenerating every number in the results table.

## How is the winner selected?

- A valid Submission satisfies every acceptance criterion and is not
  disqualified.
- Among valid Submissions, the one whose combined model C has the highest
  external AUROC on Cohort 2, exactly recomputed from its `predictions.csv`,
  wins.
- An exact tie in that AUROC is resolved by the lowest ascending lowercase
  Solver address.
- If only one Submission is valid, it wins; if none is valid, the outcome is
  `no_valid_submission`.

## Disqualification Conditions

- A required deliverable is missing, corrupt, or unreadable;
- `predictions.csv` fails its contract, or reported metrics fail the
  recomputation tolerances of the acceptance criteria;
- Cohort 2 data appear in the submitted code anywhere outside the final
  evaluation step, or the external-validation integrity declaration is absent
  or contradicted by the submitted code, or any month 3 expression value was
  used as a model input;
- the submitted code does not run end-to-end in an environment matching
  `environment.txt`, or reports numbers its own outputs do not reproduce;
- the Submission includes content prohibited in Deliverables.

## Out Of Scope

Toxicity and treatment-discontinuation endpoints, ancestry-related
calibration, adherence, follow-up-intensity and workflow effects, clinical
deployment or implementation claims, new experiments, and target
prioritization are outside this bounty. Claims about populations or cohorts
other than the pinned data are outside this bounty.

## Evaluation Procedure

All metrics are computed from the submitted prediction rows.

- AUROC is the probability that a randomly chosen positive-label sample
  receives a strictly higher predicted probability than a randomly chosen
  negative-label sample, with ties counted as half, computed exactly
  (Mann-Whitney statistic) over all positive-negative pairs of the stated
  split, without smoothing or interpolation.
- Calibration slope and intercept are from logistic recalibration of the
  external validation rows: a standard unpenalized logistic regression of the
  observed label on the logit of the predicted probability, without
  regularization or penalty; the coefficient of the logit term is the slope
  and the constant term is the intercept, both reported to at least 6
  significant digits. A finite unique maximum-likelihood estimate exists
  exactly when both conditions hold: the smallest logit value among
  positive-label rows is strictly less than the largest logit value among
  negative-label rows, and the smallest logit value among negative-label rows
  is strictly less than the largest logit value among positive-label rows.
  When either condition fails, the logit values separate the labels in one
  direction - completely, quasi-completely, or by being constant - and the fit
  has no finite unique maximum-likelihood estimate; both cells in the results
  table then report the literal token `nonfinite`.
- Sensitivity, specificity, PPV and NPV at threshold 0.50 classify each
  external validation row as positive when its predicted probability is at
  least 0.50. Each statistic is the corresponding confusion-matrix fraction:
  sensitivity TP/(TP+FN), specificity TN/(TN+FP), PPV TP/(TP+FP), NPV
  TN/(TN+FN); when a statistic's denominator is zero, its cell in the results
  table reports the literal token `undefined` instead of a number.
- Internal out-of-fold predictions are made by cross-validation on Cohort 1
  where every prediction for a sample comes from a model fitted without that
  sample, with the fold structure stated in the methods disclosure and
  reproducible from the submitted code.
