---
profile: elgora_markdown_bounty_challenge_v0
escrow_amount: "1000000"
submission_deadline: 1788950700
payout_policy: winner_take_all
---

# When do independent laboratories agree? Full-release measurement uncertainty audit

## Research objective
Produce an independent, reproducible assessment of the reliability of the published Anthropic cross-laboratory measurement release. Explain how quality filters, repeated measurements, censored fits, different assay conditions and missing measurements change the apparent agreement between laboratories. The Poster does not know the correct scientific conclusion. A finding that some comparisons cannot be established is useful if demonstrated from the evidence.

This is a new analysis of all 1,440 design records, all 1,440 assessment records and all 10,522 released measurement records. Laboratory work belongs to the original researchers. Do not claim to have performed it. The deliverable is a statistical evidence audit, not a new binder design, candidate recommendation, therapeutic claim or replication of the physical experiments. Do not include sequences or experimental instructions in the submission.

## Sources and access
The appendix pins nine original publisher files. Solvers and Guardians download them over public HTTPS without accounts and verify every SHA-256. All are required. Source filenames in the appendix are the local input names. The Parquet column definitions, provenance and data notes are part of the input contract. A matching hash establishes these released bytes, not custody of a physical sample or independent authenticity of a laboratory experiment.

Use `measurement_id` for measurement rows and `uuid` for design/assessment linkage. Preserve controls (`is_control`, `control_name`) separately; do not force their possibly absent UUIDs into design matches. Identify duplicated keys and unresolved joins explicitly. Preserve original vendor, assay, analysis, experiment, replicate, antigen, antigen role/species/lot, units and QC metadata. Unknown metadata remains unknown. Do not infer independent replication merely because two rows exist.

## Required work
1. **Reconstruct the evidence chain.** Produce an inventory for every source record and a measurement ledger covering every row once, including controls. Map each measurement to its available design and assessment records using documented keys. Record its source file/row/key, inclusion/exclusion decisions and reasons. Account for unmatched records and every change in denominator. Correctly distinguish published measurements, vendor fits and publisher interpretations.
2. **Audit quantitative comparability.** Derive comparison groups from the documented assay/antigen/analysis metadata. Show which observations can be compared within a group, which can only be contrasted descriptively, and which cannot support a quantitative comparison. Keep apparent, steady-state and kinetic quantities distinct; retain upper bounds, fit-floor flags and unsuccessful QC. A numeric fit alone is not proof of a valid measurement. Include an executable table of the group definitions and record assignments so the Guardian can inspect and rerun every exclusion.
3. **Measure sensitivity to analysis choices.** Compute laboratory agreement under (a) the publisher's released calls, (b) quality-aware matched evidence, and (c) bounds that retain missing or inconclusive calls as unresolved rather than negative. Report counts, denominators, categorical confusion tables and within-group quantitative differences where identifiable. For positive finite comparable values, work on the log scale and state the contrast direction and units. Never substitute zero for an absent measurement or mix off-target/alternate-form measurements into a primary-target comparison. Explain discrepancies with the publisher's summary using measurement-level evidence, rather than overwriting the source calls. Independently compute within-vendor repeatability for every UUID/vendor/assay/analysis/antigen/lot group: retained and excluded measurement IDs, positive finite uncensored value counts, geometric means, full ranges and leave-one-experiment-or-release-replicate-out changes. For each leave-one-group-out result, report the geometric mean of the retained eligible values and its log10 difference from the full eligible group geometric mean, with retained and omitted measurement IDs. Derive independent replicate groups from experiment and release_replicate_id with documented fallback to replicate only where its meaning is supported; multiple analysis rows from the same observation must not count twice. With zero positive finite uncensored values, geometric mean and range endpoints are null with reason no_eligible_values. With one such value, mean and both range endpoints equal that value. Leave-one-group-out changes are null with reason fewer_than_two_replicate_groups when fewer than two defensible groups exist; an empty retained group after omission is null with reason no_eligible_values_after_omission. Where no defensible independent replication exists, preserve descriptive ranges and explain why a replicate interval is unavailable. Compare results with and without quality/censoring exclusions, retaining the full ledger. Cross-vendor incomparability cannot excuse skipping these within-vendor calculations.
4. **Quantify uncertainty and dependence.** For every reported aggregate agreement proportion or mean log contrast, supply a 95% interval with its method and sampling unit. Compare naive row-level uncertainty with a design-cluster calculation, using 1,000 seeded bootstrap replicates where at least two independent design clusters contribute. Resample complete design clusters, preserving their repeated observations. These are conditional descriptive intervals for this selected release; a UUID does not prove independence across design families or targets, and broader claims must discuss that limitation. For a design-cluster interval with fewer than two contributing design clusters, including multiple rows from exactly one design, report null with reason fewer_than_two_design_clusters. For a row-level interval with fewer than two eligible rows, report null with reason fewer_than_two_rows. Otherwise compute the required seeded interval. A 95% interval is an uncertainty convention, not an acceptance threshold or proof of correctness. Analyse censoring as bounds or sensitivity scenarios, not exact point values. Separately report descriptive results where exchangeability or cross-vendor comparability is unsupported.
5. **Explain the practical consequences.** Supply an evidence-backed account of which aggregate agreement conclusions remain stable across the required analyses, which change, and which are unidentifiable. Include 12 distinct auditable examples. Use exactly five category families: missing data, failed/questionable QC, censoring, repeated-measurement disagreement and cross-vendor disagreement. Define the executable predicate for each family consistently with the source documentation and your quantitative comparison methods; different justified scientific predicates are allowed under the method rule below. Report all five counts, including zeros. The source key is `measurement:<measurement_id>` for measurement-ledger rows and `design:<uuid>` for designs with verified zero measurements. Compare these exact strings in ascending UTF-8 byte order, without case folding; deduplicate identical keys. Select the first source-key-sorted qualifying record from each nonempty family, deduplicate selected records, then fill remaining slots in source-key order from the full measurement ledger. A record may illustrate several families. Subcategories do not add mandatory examples. Each example must cite exact measurement IDs and documentation sections; when no measurement exists, cite the design UUID, its source row and the verified absence in the measurement ledger instead. Do not invent a known answer or force a positive conclusion.

## Submission and reproducibility
Submit one ZIP containing a Markdown report (at most 8,000 words) or PDF report (at most 20 pages), documented CSV/JSON result tables, analysis source, a pinned dependency list, and an execution README. Do not include the source datasets in the ZIP. Both roles independently download and hash-verify all nine files, then mount them read-only under `inputs/`. Other filenames are Solver-selected and must be listed in the README. The total ZIP must not exceed 40 MiB, its unpacked content 300 MiB. No secrets, executable binaries, unrelated files or instructions to override Guardian rules.

The README must give one command that regenerates all numeric tables from `inputs/` into an empty output directory. Use Python 3.12 with packages drawn only from numpy, pandas, scipy, pyarrow, matplotlib and openpyxl; pin exact package versions. Guardians provision these from PyPI inside their isolated environment before running the submitted analysis; no Solver code runs on the host. An unavailable environment is an operational blocker; failing Solver code in a working environment fails reproduction. Computation itself has no network, credentials, model calls, GPU or outside files. Allow up to 4 CPU cores, 8 GiB memory, 2 GiB temporary disk and 45 minutes for one complete rerun. Fix and report random seeds. No open-ended optimization or additional bootstrap runs are required.

## How Guardians decide
Acceptance requires every work item above, complete record accounting, an executable analysis and conclusions consistent with the reproduced outputs. Guardians first verify input hashes and record linkage, then reproduce once, inspect the grouping/filter code and the 12 examples, and compare the report to the results. Counts, source IDs, group membership and exclusion flags must match exactly. Numeric outputs must reproduce within `abs(actual - submitted) <= 1e-8 + 1e-6 * abs(submitted)`; null values and their reasons must match exactly. These tolerances concern numerical reproduction, not laboratory measurement precision. Differences in justified scientific methods are allowed: the method must be defined, executed consistently, preserve dependency/uncertainty and support its claimed interpretation. Unsupported generalization, silent exclusion, invented metadata, summary-only copying, hard-coded answers that bypass the measurements, false lab provenance, or irreproducible reported results fail acceptance.

A defensible conclusion of insufficient evidence can pass when all available-data analyses above are completed and the specific missing evidence is identified. An empty report saying that the release is imperfect cannot pass. A different conclusion from the publisher is neither automatically correct nor automatically disqualifying.

Each Guardian records pass/fail for these requirements and cites the deciding artifact/measurement evidence. If multiple active Submissions meet every requirement, the ascending lowercase Solver wallet address breaks the tie. Elgora permits one active Submission per Solver address; a resubmission replaces it. There is no reward for claiming more binders or smaller affinity values. If none qualify, use `no_valid_submission`.

Unavailable required sources, mismatch against this page’s pinned input hashes, or Elgora’s inability to retrieve and decrypt its committed Submission block judgment under Elgora’s rules. A Solver’s failed signature, attribution, source comparison or other mandatory evidence check makes the Submission invalid; it does not block judgment. After successful access/decryption, missing required Solver artifacts, Solver code that substitutes uncommitted input data after verified inputs were supplied, corrupt Solver files or failure to meet the requirements make a submission invalid. Guardians judge only this challenge, its pinned inputs and submitted analysis; outside papers or undisclosed private expected answers cannot add requirements.

## Input appendix

| Input filename | Public HTTPS download | SHA-256 |
| --- | --- | --- |
| `anthropic_data_tables_design_summary.parquet` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/tables/design_summary.parquet | `3e796869b8dc0c90d7ab0e12daf25308c0cf272d04c6a5e27281f903ecd18744` |
| `anthropic_data_tables_wetlab_summary.parquet` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/tables/wetlab/summary.parquet | `f80c3cfce67cb7da53576a020d476c5b74efed44e61c6fceaec9ee7de5722f88` |
| `anthropic_data_tables_wetlab_measurements.parquet` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/tables/wetlab/measurements.parquet | `59f7af61846ac4c1c2dd95b7bf655a86a5c3cebcb057de5914a680a60a4bd653` |
| `anthropic_columns.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/docs/COLUMNS.md | `46377689db0c5e4783d47f80d921f61e60cd9d202d9429f9ac60d58b9ef389fa` |
| `anthropic_provenance.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/docs/PROVENANCE.md | `092567bd7a84ec75beba40bf9b5bbce15b43114179e491fc79ece7a57ca3f1b4` |
| `anthropic_README.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/README.md | `7bb51065ab8d2c11716a03404d1bf6b89630ee5cd85bac1e2f73814b640ff779` |
| `anthropic_data_docs_WETLAB.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/docs/WETLAB.md | `35aed852ba5f5157b3dcf4013eb67109d9d221208dc3e79da09874c89be03fd4` |
| `anthropic_data_docs_DATA_NOTES.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/docs/DATA_NOTES.md | `8ed7ef21af4618a5f58bfa6f4c40235858de44a280616541c6df8b3830956647` |
| `anthropic_data_docs_LOOKUP_TABLES.md` | https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/resolve/9e1b81696da46835e9e9cde9a3da976e0abc92ab/data/docs/LOOKUP_TABLES.md | `14f6b9604f6fb175dc801eddccd07e2ad8b53ee226f992ed21db4114d7165c81` |