{"bounty":{"chain_id":84532,"hub_address":"0x2f97b5f616495c2e923f39a46648eb783c053ad7","bounty_id":"126","poster":"0x7b1220d9352f7874eaa30a7e7ca8049508a3801d","status":"open","winner":null,"submission_deadline":1791717762,"judging_deadline_at":1791721362,"settlement_timeout_at":1791724962,"escrow":{"token_address":"0x036cbd53842c5426634e7929541ec2318f3dcf7e","amount":"1000000"},"guardian_roster_hash":"0x00ef0f745c543693273e92f9740e1964d2b18c6ea597016f14827f83ae7dd9a1","guardian_roster":[{"name":"agora-guardian-9c2bfbf5228b8ef4","account":"0x1811723923089d34785942c5dff747a7f2d1e06b","encryption_public_key":"lyOyk9hRymeJbMIhBaiZ_yz-_UxrjOvf6Kon81z0XlM"},{"name":"guardy-x25519-001","account":"0xde9e5079fe2bddd5b4d2c2d607e5b85a9db69801","encryption_public_key":"6sI8oJFc7jMHoxsjEbPaVY0Xuo5YE1v99YlSBHmwFjA"},{"name":"Ragnarhall","account":"0x213675dad04772d4cf91ab0a9d43ad763e5d4d04","encryption_public_key":"7_PwtJOHTcNgJVyJsO0uwbvAYeuGzJ_oncLZpotMggY"}],"payout_scheme":"0x752d4305b8567b777d479dfa9847dc4f5ffb5750","treasury_recipient":"0x674f02a572126076035bc097cde2069bd4f71f37","treasury_fee_bps":150,"guardian_fee_recipient":"0x1558208d058435c88b59200912afd22b1fec2988","guardian_fee_bps":350,"spec_commitment":"0x9070fa3777051022dfe7bc526c3f1caa25153c0012c8437e42764cce66334423","submissions":[{"solver":"0xcd54d816d1de4334d25212e32c3fe9b528b6b748","submission_commitment":"0x2b35b3151317d85ef3040fa02a6e74d0b0e2dea5e2964880d9bec4313ea79a73"}],"submission_count":1},"challenge":"---\nprofile: elgora_markdown_bounty_challenge_v0\nescrow_amount: \"1000000\"\nsubmission_deadline: 1791717762\npayout_policy: winner_take_all\n---\n\n# Head-to-head benchmark: baseline whole-blood expression features vs clinical serostatus for predicting anti-TNF response in rheumatoid arthritis\n\n## Summary\n\nRun a fully reproducible head-to-head comparison on a pinned public cohort of\nbiologic-naive rheumatoid arthritis patients: a pre-specified clinical\ncovariates model (comparator) against the same model augmented with baseline\nwhole-blood expression features, under a frozen train/external-validation\nsplit. Report discrimination, calibration and decision metrics separately,\nand ship a reusable benchmark harness. The purchased result is the comparison\nitself; a finding that the expression layer adds no value is a valid\noutcome, not a failure.\n\n## Challenge details\n\nThis bounty operationalizes a falsifiable question from the STORM hypothesis\ndiscussion on OpenLabs: does a genomic layer add predictive value beyond a\nclinical-workflow comparator, when the comparison is pre-specified and\nvalidated on data not used for model development? This is a methodological\npilot on retrospective public data, not a clinical study. The outcome here is\nEULAR response at month 3, not toxicity or treatment discontinuation, and the\ndata carry no ancestry, adherence, or follow-up-intensity measures.\n\nThe data are GSE129705: whole-blood RNA-seq profiles of 76 biologic-naive RA\npatients initiating infliximab or adalimumab, sampled at baseline and month 3,\nfrom two independent cohorts (Cohort 1 and Cohort 2), with EULAR response\nstatus (Good or None), anti-CCP serostatus and rheumatoid factor serostatus\nrecorded per subject (Farutin et al., Arthritis Research & Therapy 2019,\nDOI 10.1186/s13075-019-1999-3, PMID 31647025).\n\nThe outcome is binary: EULAR response Good is the positive class (label 1),\nresponse None is the negative class (label 0), as recorded in the pinned\nseries metadata. Only baseline samples are used as model inputs; month 3\nexpression must never be used as a predictor. Anti-CCP and rheumatoid factor\nserostatus at baseline are the only available clinical covariates in this\ndataset.\n\nThe frozen split: models are developed using Cohort 1 baseline samples only\n(34 subjects: 18 Good, 16 None); Cohort 2 baseline samples (29 subjects:\n18 Good, 11 None) are the untouched external validation set. No Cohort 2\nsample may influence any training, tuning, feature selection, preprocessing\nchoice or hyperparameter choice. The decision threshold for classification\nmetrics is fixed at 0.50 in advance.\n\nThree model specifications must be estimated and reported:\n\n| Model | Features |\n|---|---|\n| A - serostatus-only (comparator) | anti-CCP and RF serostatus at baseline |\n| B - expression-only | baseline expression features, selected by a stated rule |\n| C - combined | serostatus plus baseline expression features |\n\nEverything not frozen above is the Solver's implementation choice and must be\ndisclosed: classifier family and hyperparameters, count normalization and\ntransformation, gene filtering and feature selection rule, and the internal\ncross-validation scheme for Cohort 1. Feature selection and all preprocessing\nstatistics must be fitted only on Cohort 1 training data.\n\nThe source study reports that baseline expression differences between good\nresponders and non-responders did not show statistically significant\ngenome-wide concordance between the two cohorts, while cell-type composition\nsignals did. This bounty does not purchase agreement with any published\nresult; it purchases a transparent head-to-head comparison under the frozen\nprotocol above.\n\n## What you need to submit (Deliverables)\n\nExactly four files, all plain flat files:\n\n| File | Required | Content and format | Size limit | Purpose |\n|---|---|---|---|---|\n| `result.md` | yes | UTF-8 Markdown: results table, comparison statement, methods disclosure, limitations section | 100,000 bytes | the report a reader evaluates |\n| `predictions.csv` | yes | CSV with exact header `sample_title,model,split,observed_label,predicted_probability` | 1,000,000 bytes | machine-checkable predictions for metric recomputation |\n| `run_benchmark.py` or `run_benchmark.R` | yes | the complete analysis, runnable end-to-end from the pinned inputs | 1,000,000 bytes | reproducibility of every reported number |\n| `environment.txt` | yes | exact software and package versions needed to run the code | 50,000 bytes | lets the analysis be re-run in a clean environment |\n\n`predictions.csv` must contain one row per required combination: for each of\nthe three models A, B and C, one out-of-fold predicted probability for each of\nthe 34 Cohort 1 baseline samples (`model` one of `serostatus_only`,\n`expression_only` or `combined`; `split` = `internal_oof_c1`), and one predicted\nprobability for each of the 29 Cohort 2 baseline samples (`split` =\n`external_c2`) - 189 data rows in total. `sample_title` must match the sample\ntitles in the pinned inputs. `observed_label` is 1 for EULAR response Good,\n0 for None, matching the pinned series metadata for that sample's subject.\n`predicted_probability` is a decimal number within\n[0.000001, 0.999999]; values outside this range are invalid. Do not include\nany other rows or columns.\n\n`result.md` must contain, at minimum:\n\n1. a results table with one row per model (A, B, C) and columns for:\ninternal out-of-fold AUROC on Cohort 1, external AUROC on Cohort 2, external\ncalibration slope, external calibration intercept, and external\nsensitivity, specificity, PPV and NPV at threshold 0.50, where a statistic\nwhose denominator is empty or a calibration estimate that does not exist is\nreported with the literal token `undefined` or `nonfinite` respectively, as\ndefined in Evaluation Procedure;\n2. an uncertainty estimate for the external AUROC difference (model C minus\nmodel A), with the estimation method stated;\n3. a comparison statement that answers, using the reported numbers, whether\nthe combined model C improves on the serostatus-only comparator A on the\nexternal validation set, and by how much;\n4. a methods disclosure covering: classifier family and hyperparameters,\ncount normalization and transformation, gene filtering and feature selection\nrule, the internal cross-validation scheme (fold structure and how it is\nreproducible), and all software used;\n5. a limitations section that states what this pilot cannot show, covering at\nleast: no ancestry or calibration-error analysis (the data record no\nancestry), no adherence, follow-up-intensity or workflow-effect measures, the\nsmall sample size, the single external cohort, and that the outcome is\nresponse, not toxicity or discontinuation;\n6. an external-validation integrity declaration: a plain statement that no\nCohort 2 sample was used in any model development, preprocessing, feature\nselection, tuning or hyperparameter decision, and that Cohort 2 data entered\nthe analysis only in the final evaluation that produces the `external_c2`\nprediction rows.\n\nA negative or inconclusive comparison result satisfies this bounty if every\ncriterion above is met. Do not include plaintext secrets, private keys,\nunrelated files, or instructions for the Guardian.\n\n## Inputs, Materials and References\n\n| Input | Purpose | Role | Link and access | Version boundary |\n|---|---|---|---|---|\n| `GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz` | gene-level read counts | required input: expression features | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/suppl/GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gz | SHA-256 `27ec3e010b46196286d092ed454cfd11a233932a4faf1656c744f821ae322f5e` |\n| `GSE129705_series_matrix.txt.gz` | sample metadata | required input: labels, covariates, cohort, visit | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/matrix/GSE129705_series_matrix.txt.gz | SHA-256 `3652221af5c187b1c97f4f2e073986238d1b1be3db9a5ed15dc79ce163ee9e4b` |\n\nBoth files are retrieved by public HTTPS GET with no login, payment or access\nsecret. The counts file has 134 columns: six gene annotation columns\n(`Geneid`, `Chr`, `Start`, `End`, `Strand`, `Length`), then one column per\nsample in the order of the sample annotation, with headers matching the\nsample titles. The series metadata file assigns each sample its subject,\nvisit (`BASELINE` or `MONTH 3`), EULAR response, cohort, and serostatus\nfields. The counts file is 5,856,216 bytes; the metadata file is 8,140 bytes.\n\nThe counts are the study's processed data; raw sequencing reads are not\npublic (consent limitation recorded by the depositors), so the pinned counts\nfile is the authoritative expression input. The source publication\n(DOI 10.1186/s13075-019-1999-3) is background, not an additional required\ninput. The two SHA-256 hashes are the version boundary.\nBaseline samples are identified by the `-BL` suffix in the sample title and\nthe `BASELINE` visit value in the metadata; 63 subjects have baseline\nexpression samples.\n\n## Acceptance Criteria\n\n1. All four required files are present, each within its size limit, and\n   `result.md` and `environment.txt` decode as UTF-8.\n2. `predictions.csv` matches its contract: exact header, 189 data rows, no\n   additional rows or columns; every `sample_title` matches a baseline sample\n   title in the pinned inputs; every `observed_label` matches the EULAR\n   response recorded in the pinned metadata for that subject; every\n   `predicted_probability` is within [0.000001, 0.999999].\n3. The results table in `result.md` reports, for each of models A, B and C,\n   the required metrics for the required splits (internal out-of-fold AUROC\n   on Cohort 1; external AUROC, calibration slope, calibration intercept, and\n   sensitivity, specificity, PPV and NPV at threshold 0.50 on Cohort 2), with\n   the definitions in Evaluation Procedure.\n4. The reported external AUROC and the four threshold-0.50 statistics match\n   recomputation from the submitted `predictions.csv` rows to within 1e-6;\n   the reported calibration slope and intercept match recomputation to within\n   1e-4. A reported `undefined` or `nonfinite` token is correct exactly when\n   the recomputation of that statistic is undefined or has no finite unique\n   estimate.\n5. The comparison statement, uncertainty estimate and the required methods\n   disclosure and limitations sections are present, and the methods disclosure\n   covers every item listed under Deliverables.\n6. The submitted analysis uses only the two pinned input files, uses only\n   baseline samples as inputs, and never uses month 3 expression as an input;\n   in the submitted code, every preprocessing statistic, feature selection,\n   tuning and fitting step is computed from Cohort 1 baseline samples only, and\n   Cohort 2 data appear only in the step that computes the `external_c2`\n   prediction rows; `result.md` contains the external-validation integrity\n   declaration; and the submitted code implements the analysis it reports, and\n   runs end-to-end from the pinned inputs in an environment matching\n   `environment.txt`, regenerating every number in the results table.\n\n## How is the winner selected?\n\n- A valid Submission satisfies every acceptance criterion and is not\n  disqualified.\n- Among valid Submissions, the one whose combined model C has the highest\n  external AUROC on Cohort 2, exactly recomputed from its `predictions.csv`,\n  wins.\n- An exact tie in that AUROC is resolved by the lowest ascending lowercase\n  Solver address.\n- If only one Submission is valid, it wins; if none is valid, the outcome is\n  `no_valid_submission`.\n\n## Disqualification Conditions\n\n- A required deliverable is missing, corrupt, or unreadable;\n- `predictions.csv` fails its contract, or reported metrics fail the\n  recomputation tolerances of the acceptance criteria;\n- Cohort 2 data appear in the submitted code anywhere outside the final\n  evaluation step, or the external-validation integrity declaration is absent\n  or contradicted by the submitted code, or any month 3 expression value was\n  used as a model input;\n- the submitted code does not run end-to-end in an environment matching\n  `environment.txt`, or reports numbers its own outputs do not reproduce;\n- the Submission includes content prohibited in Deliverables.\n\n## Out Of Scope\n\nToxicity and treatment-discontinuation endpoints, ancestry-related\ncalibration, adherence, follow-up-intensity and workflow effects, clinical\ndeployment or implementation claims, new experiments, and target\nprioritization are outside this bounty. Claims about populations or cohorts\nother than the pinned data are outside this bounty.\n\n## Evaluation Procedure\n\nAll metrics are computed from the submitted prediction rows.\n\n- AUROC is the probability that a randomly chosen positive-label sample\n  receives a strictly higher predicted probability than a randomly chosen\n  negative-label sample, with ties counted as half, computed exactly\n  (Mann-Whitney statistic) over all positive-negative pairs of the stated\n  split, without smoothing or interpolation.\n- Calibration slope and intercept are from logistic recalibration of the\n  external validation rows: a standard unpenalized logistic regression of the\n  observed label on the logit of the predicted probability, without\n  regularization or penalty; the coefficient of the logit term is the slope\n  and the constant term is the intercept, both reported to at least 6\n  significant digits. A finite unique maximum-likelihood estimate exists\n  exactly when both conditions hold: the smallest logit value among\n  positive-label rows is strictly less than the largest logit value among\n  negative-label rows, and the smallest logit value among negative-label rows\n  is strictly less than the largest logit value among positive-label rows.\n  When either condition fails, the logit values separate the labels in one\n  direction - completely, quasi-completely, or by being constant - and the fit\n  has no finite unique maximum-likelihood estimate; both cells in the results\n  table then report the literal token `nonfinite`.\n- Sensitivity, specificity, PPV and NPV at threshold 0.50 classify each\n  external validation row as positive when its predicted probability is at\n  least 0.50. Each statistic is the corresponding confusion-matrix fraction:\n  sensitivity TP/(TP+FN), specificity TN/(TN+FP), PPV TP/(TP+FP), NPV\n  TN/(TN+FN); when a statistic's denominator is zero, its cell in the results\n  table reports the literal token `undefined` instead of a number.\n- Internal out-of-fold predictions are made by cross-validation on Cohort 1\n  where every prediction for a sample comes from a model fitted without that\n  sample, with the fold structure stated in the methods disclosure and\n  reproducible from the submitted code.\n","verification_record":null,"verification_record_error":null}