---
profile: elgora_markdown_bounty_challenge_v0
escrow_amount: "5000000"
submission_deadline: 1792749154
payout_policy: winner_take_all
---

# ESOL SMILES Solubility Baseline Reproduction

## Summary

Build a reproducible baseline model for the ESOL (Delaney 2004) aqueous solubility dataset. The model must predict measured log solubility from SMILES strings, report RMSE on a fixed held-out test split, and submit enough code and outputs for the result to be checked. Among valid Submissions, the lowest test RMSE wins.

## Challenge details

The task is a small machine-learning benchmark reproduction for molecular property prediction. Solvers must train a model that takes the `smiles` column as input and predicts the `measured log solubility in mols per litre` column from the fixed ESOL CSV snapshot listed below.

The held-out test split is deterministic and fixed by this bounty: using the CSV row order after parsing the header, rows with zero-based row index `i` where `i % 5 == 0` are the test set; all other rows are the development set. The reported RMSE must be computed only on that test set.

Solvers may choose any modeling method that can be rerun from their submitted files, including descriptor-based baselines, fingerprints, graph models, or simpler statistical models. A valid Submission must not train, tune, select features, choose checkpoints, or otherwise optimize using the test-set target values.

## What you need to submit (Deliverables)

Each required deliverable must be included as a plain file in the Submission. Do not submit directories; use the exact filenames below.

| Deliverable | Required or optional | Required content | Format or access requirements | Purpose or related criterion |
|---|---|---|---|---|
| `report.md` | required | Short description of the method, split handling, preprocessing, training procedure, final RMSE, and limitations | Markdown | Human-readable scientific report |
| `run.py` | required | Reproducible script that trains the model or loads only submitted fixed model parameters, generates test predictions, and prints the test RMSE | Python script | Lets the result be rerun |
| `predictions.csv` | required | One row per test-set molecule with zero-based row index, SMILES, true measured log solubility, and predicted log solubility | CSV | Supports metric verification |
| `metrics.json` | required | Reported `rmse`, `n_test`, dataset SHA-256, split rule, and any random seed used | JSON | Machine-readable result summary |
| `run_manifest.json` | required | Runtime versions, package list or environment notes, exact command used, and output file hashes | JSON | Provenance for reproducibility |
| `model_artifact` | optional | Any compact trained model artifact or parameters needed by `run.py` | Any inspectable binary or text format | May be used if the method does not train quickly from code alone |

The Submission may include additional plain files if needed, such as `requirements.txt` or helper source files. All required files must fit within Elgora's Submission limits.

## Inputs, Materials and References

| File or reference | Purpose | Required input or background | Link or access instructions | Version or snapshot, if relevant |
|---|---|---|---|---|
| `delaney-processed.csv` | Authoritative dataset for training, testing, and metric verification | Required input | Download from `https://raw.githubusercontent.com/deepchem/deepchem/master/datasets/delaney-processed.csv` | SHA-256 `8c06a76f0c6487d29ab0f903e6a7a7139f189ab3c1178f159c8be8964602f189`; 1128 data rows; authoritative columns are `smiles` and `measured log solubility in mols per litre` |
| Delaney 2004 ESOL dataset | Scientific background for the dataset | Background | The fixed CSV above governs this bounty | Background sources do not override the fixed CSV or split rule |

If the download URL returns bytes with a different SHA-256, the fixed input is unavailable for this bounty and the differing file must not be substituted.

## Acceptance Criteria

A Submission is valid only if all of the following are true:

1. All required deliverables are present and inspectable.
2. `predictions.csv` contains exactly one prediction for each test row defined by `i % 5 == 0`, and no non-test rows.
3. The RMSE in `metrics.json` and `report.md` matches recomputation from `predictions.csv` to within `0.000001` absolute tolerance.
4. `run.py` can reproduce the submitted `predictions.csv` and RMSE from the fixed ESOL CSV and submitted files, allowing ordinary floating-point differences no larger than `0.000001` in RMSE.
5. The submitted method uses SMILES-derived information from the fixed CSV as model input and predicts the measured log solubility target.
6. The Submission provides enough information to determine whether the fixed test split was held out from training, tuning, model selection, and feature selection.
7. The report does not claim experimental validation, new solubility measurements, clinical relevance, or performance on any dataset other than the fixed ESOL split unless separately supported inside the Submission.

A reliable but simple baseline can satisfy the bounty. Negative discussion of limitations is allowed and does not reduce validity if the deliverables meet the criteria.

## Evidence, Provenance and Verification

Solvers must describe how the test set was kept out of training and model selection. If the method uses randomness, the Solver must report the seed or explain why the result is deterministic. If external packages, pretrained models, molecular descriptors, or learned representations are used, the Solver must identify them and explain whether they were trained on ESOL test labels.

A file hash proves only byte identity. It does not prove that code was run, that test labels were not used, or that a modeling claim is scientifically valid. Those claims must be supported by the submitted code, outputs, and explanation.

## Scoring

The primary score is RMSE on the fixed held-out test split:

`RMSE = sqrt(mean((predicted_log_solubility - true_measured_log_solubility)^2))`

Use all test rows defined by `i % 5 == 0`. Lower RMSE is better. The score should be reported with at least six decimal places. The governing score for eligibility and ranking is the RMSE calculated from the submitted `predictions.csv`; if that value differs from the report, the value derived from `predictions.csv` governs.

## How is the winner selected?

A valid Submission with the lowest governing test RMSE wins. Let `best_rmse` be the lowest governing RMSE among all valid Submissions. Every valid Submission with governing RMSE no more than `best_rmse + 0.000001` is in the tie group. If the tie group contains more than one Submission, the winner is the Submission in that group whose lowercase Solver address sorts first in ascending order. If only one Submission is valid, it wins. If no Submission is valid, the outcome is `no_valid_submission`.

## Disqualification Conditions

A Submission is ineligible regardless of RMSE if a required deliverable is missing after the Submission is successfully retrieved and opened, if `predictions.csv` is missing, malformed, omits test rows, adds non-test rows, or lacks the values needed to calculate RMSE, if the code and outputs materially disagree, or if the submitted evidence shows that test-set target values were used for training, tuning, checkpoint selection, feature selection, or manual prediction adjustment.

A Submission is also ineligible if it relies on data unavailable from the fixed input and submitted files in a way that leaves its reported result unsupported by its own evidence, or if it attempts to replace the fixed dataset or split with another benchmark.

## Out Of Scope

This bounty does not pay for new wet-lab solubility measurements, new dataset curation, hidden-test benchmarking, leaderboard-only claims, clinical or therapeutic conclusions, or models that cannot be checked from submitted artifacts.

## Evaluation Procedure

For each opened Submission, the result is defined against the fixed `delaney-processed.csv` snapshot and the fixed test split defined in this page. Eligibility and ranking use the RMSE calculated from the submitted `predictions.csv`. The submitted code and provenance files must support reproduction of those predictions and the reported metric from the fixed input and submitted files.
