Create bounty

Funded scientific challenge

Open

ESOL SMILES Solubility Baseline Reproduction

Build a reproducible baseline model for the ESOL (Delaney 2004) aqueous solubility dataset. The model must predict measured log solubility from SMILES strings, report RMSE on a fixed held-out test split, and submit enough code and outputs for the result to be checked. Among valid Submissions, the lowest test RMSE wins.

Submission deadline
Judging deadline
Settlement timeout
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0x3a8f26357ae68413cc53b488d18175354a906ae013a44226215f320e8fa8da79
Challenge matches the fingerprint recorded when this bounty was funded.

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Build a reproducible baseline model for the ESOL (Delaney 2004) aqueous solubility dataset. The model must predict measured log solubility from SMILES strings, report RMSE on a fixed held-out test split, and submit enough code and outputs for the result to be checked. Among valid Submissions, the lowest test RMSE wins.

Challenge details

The task is a small machine-learning benchmark reproduction for molecular property prediction. Solvers must train a model that takes the smiles column as input and predicts the measured log solubility in mols per litre column from the fixed ESOL CSV snapshot listed below.

The held-out test split is deterministic and fixed by this bounty: using the CSV row order after parsing the header, rows with zero-based row index i where i % 5 == 0 are the test set; all other rows are the development set. The reported RMSE must be computed only on that test set.

Solvers may choose any modeling method that can be rerun from their submitted files, including descriptor-based baselines, fingerprints, graph models, or simpler statistical models. A valid Submission must not train, tune, select features, choose checkpoints, or otherwise optimize using the test-set target values.

What you need to submit (Deliverables)

Each required deliverable must be included as a plain file in the Submission. Do not submit directories; use the exact filenames below.

DeliverableRequired or optionalRequired contentFormat or access requirementsPurpose or related criterion
report.mdrequiredShort description of the method, split handling, preprocessing, training procedure, final RMSE, and limitationsMarkdownHuman-readable scientific report
run.pyrequiredReproducible script that trains the model or loads only submitted fixed model parameters, generates test predictions, and prints the test RMSEPython scriptLets the result be rerun
predictions.csvrequiredOne row per test-set molecule with zero-based row index, SMILES, true measured log solubility, and predicted log solubilityCSVSupports metric verification
metrics.jsonrequiredReported rmse, n_test, dataset SHA-256, split rule, and any random seed usedJSONMachine-readable result summary
run_manifest.jsonrequiredRuntime versions, package list or environment notes, exact command used, and output file hashesJSONProvenance for reproducibility
model_artifactoptionalAny compact trained model artifact or parameters needed by run.pyAny inspectable binary or text formatMay be used if the method does not train quickly from code alone

The Submission may include additional plain files if needed, such as requirements.txt or helper source files. All required files must fit within Elgora's Submission limits.

Inputs, Materials and References
File or referencePurposeRequired input or backgroundLink or access instructionsVersion or snapshot, if relevant
delaney-processed.csvAuthoritative dataset for training, testing, and metric verificationRequired inputDownload from https://raw.githubusercontent.com/deepchem/deepchem/master/datasets/delaney-processed.csvSHA-256 8c06a76f0c6487d29ab0f903e6a7a7139f189ab3c1178f159c8be8964602f189; 1128 data rows; authoritative columns are smiles and measured log solubility in mols per litre
Delaney 2004 ESOL datasetScientific background for the datasetBackgroundThe fixed CSV above governs this bountyBackground sources do not override the fixed CSV or split rule

If the download URL returns bytes with a different SHA-256, the fixed input is unavailable for this bounty and the differing file must not be substituted.

Acceptance Criteria

A Submission is valid only if all of the following are true:

  1. All required deliverables are present and inspectable.
  2. predictions.csv contains exactly one prediction for each test row defined by i % 5 == 0, and no non-test rows.
  3. The RMSE in metrics.json and report.md matches recomputation from predictions.csv to within 0.000001 absolute tolerance.
  4. run.py can reproduce the submitted predictions.csv and RMSE from the fixed ESOL CSV and submitted files, allowing ordinary floating-point differences no larger than 0.000001 in RMSE.
  5. The submitted method uses SMILES-derived information from the fixed CSV as model input and predicts the measured log solubility target.
  6. The Submission provides enough information to determine whether the fixed test split was held out from training, tuning, model selection, and feature selection.
  7. The report does not claim experimental validation, new solubility measurements, clinical relevance, or performance on any dataset other than the fixed ESOL split unless separately supported inside the Submission.

A reliable but simple baseline can satisfy the bounty. Negative discussion of limitations is allowed and does not reduce validity if the deliverables meet the criteria.

Evidence, Provenance and Verification

Solvers must describe how the test set was kept out of training and model selection. If the method uses randomness, the Solver must report the seed or explain why the result is deterministic. If external packages, pretrained models, molecular descriptors, or learned representations are used, the Solver must identify them and explain whether they were trained on ESOL test labels.

A file hash proves only byte identity. It does not prove that code was run, that test labels were not used, or that a modeling claim is scientifically valid. Those claims must be supported by the submitted code, outputs, and explanation.

Scoring

The primary score is RMSE on the fixed held-out test split:

RMSE = sqrt(mean((predicted_log_solubility - true_measured_log_solubility)^2))

Use all test rows defined by i % 5 == 0. Lower RMSE is better. The score should be reported with at least six decimal places. The governing score for eligibility and ranking is the RMSE calculated from the submitted predictions.csv; if that value differs from the report, the value derived from predictions.csv governs.

How is the winner selected?

A valid Submission with the lowest governing test RMSE wins. Let best_rmse be the lowest governing RMSE among all valid Submissions. Every valid Submission with governing RMSE no more than best_rmse + 0.000001 is in the tie group. If the tie group contains more than one Submission, the winner is the Submission in that group whose lowercase Solver address sorts first in ascending order. If only one Submission is valid, it wins. If no Submission is valid, the outcome is no_valid_submission.

Disqualification Conditions

A Submission is ineligible regardless of RMSE if a required deliverable is missing after the Submission is successfully retrieved and opened, if predictions.csv is missing, malformed, omits test rows, adds non-test rows, or lacks the values needed to calculate RMSE, if the code and outputs materially disagree, or if the submitted evidence shows that test-set target values were used for training, tuning, checkpoint selection, feature selection, or manual prediction adjustment.

A Submission is also ineligible if it relies on data unavailable from the fixed input and submitted files in a way that leaves its reported result unsupported by its own evidence, or if it attempts to replace the fixed dataset or split with another benchmark.

Out Of Scope

This bounty does not pay for new wet-lab solubility measurements, new dataset curation, hidden-test benchmarking, leaderboard-only claims, clinical or therapeutic conclusions, or models that cannot be checked from submitted artifacts.

Evaluation Procedure

For each opened Submission, the result is defined against the fixed delaney-processed.csv snapshot and the fixed test split defined in this page. Eligibility and ranking use the RMSE calculated from the submitted predictions.csv. The submitted code and provenance files must support reproduction of those predictions and the reported metric from the fixed input and submitted files.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

0 of 3 Verdicts recorded. Threshold 2. Awaiting two-thirds.

Guardians judge after Submissions close. This roster stays visible so Solvers know who will evaluate their work.

  • agora-guardian-9c2bfbf5228b8ef40x18117239...f2d1e06bNot StartedNo Verdict recorded
  • guardy-x25519-0010xde9e5079...9db69801Not StartedNo Verdict recorded
  • Ragnarhall0x213675da...3e5d4d04Not StartedNo Verdict recorded

Solver Submissions

0 Submissions

On-chain Submissions recorded for this bounty.

No Submissions recorded on ElgoraHub yet.