Funded scientific challenge

Timed out

Prioritize BCL11A enhancer variants from a public saturation-mutagenesis assay

Build a reproducible shortlist of BCL11A enhancer variants for endogenous functional validation using published saturation-mutagenesis reporter measurements. Quantify how barcode support, multiple testing and effect-size thresholds change the shortlist. A result finding no defensible robust candidates is eligible.

Submission deadline
Sep 17, 2026, 4:00 PM UTC
Judging deadline
Sep 17, 2026, 5:00 PM UTC
Settlement timeout
Sep 17, 2026, 6:00 PM UTC
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0xb8090cea09c72bd56bb4ee8b59afa095043365f61be9b42d2e114886f96c57a0
Challenge matches the fingerprint recorded when this bounty was funded.

Payout receipt · settled

Refunded to Poster

1.00USDC

0xcc7fe016...77dfdd18 ↗

  • Poster refund· 100.00%1.00 USDC

Escrow distributed1.00 USDC

Timeout settlement returns the whole escrow. No treasury or Guardian fee is charged on this path.

Your wallet

Connect an eligible wallet

Connect the eligible wallet to claim from ElgoraHub.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

Settlement timed out. Matching against a final result does not apply. 1 of 3 Verdicts recorded.

Full Poster refund
0xcc7fe016...77dfdd18 ↗
Winning Submission
None
ElgoraHub settlement
0xa4785b23...a82b41dc

Solver Submissions

6 Submissions

On-chain Submissions recorded for this bounty.

#SolverSubmittedBlockTransaction
1
0x5c3f...3eed25
Sep 17, 2026, 3:32 AM UTC#469238430x72d4a432...21be228d
2
0x706c...1466b3
Sep 17, 2026, 3:33 AM UTC#469238570xdf48c6e0...17b4e45d
3
0x7ce3...59ad90
Sep 17, 2026, 3:32 AM UTC#469238290xd1622cf3...fdb82c2c
4
0xcd54...b6b748
Sep 17, 2026, 10:52 AM UTC#469370170xff2dae1a...fa633fa1
5
0xf2ce...886013
Sep 17, 2026, 3:35 AM UTC#469239100xa6abafbd...5255c184
6
0xf465...df79bd
Sep 17, 2026, 3:31 AM UTC#469237990xfc0179dc...be052cd2

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Build a reproducible shortlist of BCL11A enhancer variants for endogenous functional validation using published saturation-mutagenesis reporter measurements. Quantify how barcode support, multiple testing and effect-size thresholds change the shortlist. A result finding no defensible robust candidates is eligible.

Challenge details

The OpenLabs BCL11A project asks whether genetic evidence supports the target. Population association alone cannot identify which nucleotide changes alter enhancer activity. This task addresses that separate functional-prioritization decision using the author-published MPRA coefficient table from Kircher et al. It does not reanalyze population GWAS, design gene edits, or establish therapeutic efficacy.

The fixed file contains 2,062 BCL11A variant rows per genome build. Use the GRCh38 rows as the analysis set, retaining one-base substitutions and deletions. The GRCh37 representation is a coordinate alternative to the same measurements, not an independent replicate. Coefficients come from the source's fitted model and are not interchangeable with an unadjusted RNA/DNA ratio. Study the experimental and modeling methods before interpretation.

Create a coverage map showing the observed alternative alleles at each represented reference position, with absent possibilities marked unobserved rather than neutral. Apply Benjamini-Hochberg correction once over all finite published p values in the complete GRCh38 BCL11A set. Keep these q values fixed when examining the cross-product of barcode thresholds at least 10, 50 or 100 and absolute coefficient thresholds at least 0.25 or 0.5, always requiring q<0.05. These are sensitivity choices for prioritization, not validated biological cutoffs.

What you need to submit (Deliverables)
  • Exact compressed source input and provenance.json, plus bcl11a-input.csv containing all GRCh38 BCL11A records with the source columns preserved. Document build, coordinate convention, variant key, coefficient scale and the meaning of barcode, DNA and RNA columns from the author documentation. Unknown conventions remain explicitly unresolved; do not silently infer a different coordinate base.
  • analysis.py or analysis.R, variant-results.csv, coverage.csv, and threshold-summary.csv: reproduce BH correction and every threshold scenario, report the direction and selection status of each observed variant, quantify pairwise shortlist overlaps and selection-set additions/removals (no within-set ranking is required), and calculate Spearman associations between absolute coefficient and barcode support/DNA/RNA counts on valid records. Explain why such associations do not identify bias or causal mechanisms. Keep substitutions and deletions distinguishable throughout.
  • functional-map.svg or functional-map.png: a position-by-alternate-allele map of coefficients with four mutually exclusive display categories: unobserved (no source row); uncertain (observed but coefficient or q is nonfinite, or q>=0.05); small estimated effect (finite coefficient and q<0.05 with absolute coefficient<0.25); and larger estimated effect (finite coefficient and q<0.05 with absolute coefficient>=0.25). The 0.25 boundary belongs to the larger-effect category and q=0.05 to uncertain. These are display rules, not evidence of biological equivalence or neutrality. Include an accompanying barcode-support view. Use actual genomic coordinates and a documented variant convention.
  • shortlist.csv and decision.md: nominate up to five decreasing-activity and up to five increasing-activity variants for a next validation stage, or explain an empty/shorter list. Every nominated variant needs its exact key, reported coefficient, q value, support counts, threshold stability and a concrete justification. Give a bounded follow-up comparison capable of testing whether the reporter effect persists at the endogenous locus; no new experiment or successful result is claimed. Discuss pooled model dependence, reporter context, multi-mutation library design, imperfect coverage, clinical/ancestry generalization and why neither a missing variant nor a nonsignificant coefficient establishes neutrality.
  • README.md: source provenance, environment, execution command and any unresolved interpretation. All required file bytes must be included; source URLs do not replace data or results.
Inputs, Materials and References
  1. Author data-access portal: https://kircherlab.bihealth.org/satMutMPRA/ . Its format description defines the modeled coefficient, p value and sequencing-support columns. The portal links its public scientific data distribution at https://github.com/kircherlab/MPRA_SaturationMutagenesis . Required file: https://raw.githubusercontent.com/kircherlab/MPRA_SaturationMutagenesis/master/data/elements.tsv.gz . Exact bytes verified 17 September 2026: SHA-256 fec2eed91fe27af3aae07ebce2eca65e9bad4bb6abba5d8c27f478887dd7b134, 2,677,485 compressed bytes. Use this checksum version, not an unrecorded later replacement. Only Element=BCL11A and Release=GRCh38 are analyzed; other elements/builds must not enter the test family.
  2. Kircher et al. (2019), *Saturation mutagenesis of twenty disease-associated regulatory elements at single base-pair resolution*, DOI 10.1038/s41467-019-11526-w, PMC6687891: https://pmc.ncbi.nlm.nih.gov/articles/PMC6687891/ . Required for experimental/modeling context and limitations. Raw sequencing at GSE126550 is not required.

Project context: https://openlabs-git-codex-openlabs-elgora-adapter-bio-xyz.vercel.app/projects/0108093d-9241-44da-8e95-61cd7dd64b05 . Independent research contribution without project-owner endorsement.

Acceptance Criteria
  1. Source checksum matches, the complete 2,062-row analysis scope is present, variant keys are preserved and repeated genome-build representations are not treated as replicated evidence. Any invalid numeric records remain accounted for with explicit reasons.
  2. BH correction uses the specified full finite-p-value family and remains fixed across the six threshold combinations. Numeric outputs reproduce within 0.000001 absolute tolerance before display rounding, with stated treatment of ties, empty sets and undefined correlations. No undocumented p-value recalculation from counts is substituted for the published model.
  3. Coverage distinguishes observed from absent variants and preserves deletion labels. Plots agree with the deposited coefficients/support counts and apply the four specified categories and equality boundaries, visibly distinguishing missing evidence from measured effects.
  4. Shortlist selection is traceable to computed evidence and does not hide threshold-fragile cases. Positive and negative coefficient directions are interpreted on the reporter scale, without equating them directly with HbF or clinical effects.
  5. The proposed validation comparison addresses an identified gap between episomal reporter evidence and endogenous function. The report makes no new-editing, population-transferability or therapeutic claim and does not describe barcodes or alternate genome builds as independent biological replication.
How is the winner selected?

Only submissions satisfying all criteria are eligible. Prefer fewer material numerical/scientific errors, then the most defensible functional-validation prioritization grounded in threshold/support sensitivity, then stronger reproducibility and provenance. More nominated variants or stronger therapeutic claims earn no preference. Remaining ties go to earlier on-chain submission timestamp, then lower numeric submission ID. No winner is required if none qualifies.