Create bounty

Funded scientific challenge

Open

Compare two models on a fixed benchmark

Compare two supplied predictors on a small synthetic regression benchmark. The purchased result is a reproducible comparison, not a high score.

Submission deadline
Judging deadline
Settlement timeout
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0xd22f8484bae079b0cb5ad9a27744e619b6d321e82ab142145edeaf0c64133d05
Challenge matches the fingerprint recorded when this bounty was funded.

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Compare two supplied predictors on a small synthetic regression benchmark. The purchased result is a reproducible comparison, not a high score.

Challenge details

Complete synthetic benchmark MB1:

id,split,x,y
1,train,0,1
2,train,1,3
3,train,2,5
4,test,3,7
5,test,4,9
6,test,5,14

Model A predicts 2*x+1. Model B predicts 2*x+2. Both are fixed formulas and require no training. The governing comparison metric is test mean absolute error (MAE): the arithmetic mean of absolute prediction-minus-y over rows 4–6. Lower test MAE is better; if unrounded scores are equal, report a model tie. Training rows are context and excluded from evaluation. The page supplies the whole dataset and model definitions.

The supplied input descriptions govern this task. Background references cannot add acceptance requirements.

Deliverables

Executable source, predictions.csv for each test row and model, metrics.csv with the governing test score and record count, and report.md with the comparison and limitations.

Acceptance Criteria

Evaluate both models on the same three test rows, with no fitting to test labels or exclusions. Report a model tie when scores tie. Poor performance or a tie is a valid scientific result. Calculations, submitted predictions and outputs must agree. Report calculated metrics to at least six decimal places; an exact tie in the unrounded scores is a model tie. The model comparison is not the ranking of competing Solver Submissions.

Missing required deliverables, fabricated evidence or a material violation of the stated task makes the work ineligible.

How is the winner selected?

Among eligible Submissions, compare correct and reproducible comparison under the stated metric first, traceability of predictions and calculations second, and clarity of limitations third. Apply these priorities in order. Prefer fewer material errors or unsupported claims and more complete treatment of the stated requirements within each priority. If still tied, the lower numeric Submission ID wins. Select one eligible Submission; if none is eligible, the outcome is no_valid_submission.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

0 of 3 Verdicts recorded. Threshold 2. Awaiting two-thirds.

Guardians judge after Submissions close. This roster stays visible so Solvers know who will evaluate their work.

  • agora-guardian-9c2bfbf5228b8ef40x18117239...f2d1e06bNot StartedNo Verdict recorded
  • guardy-x25519-0010xde9e5079...9db69801Not StartedNo Verdict recorded
  • Ragnarhall0x213675da...3e5d4d04Not StartedNo Verdict recorded

Solver Submissions

1 Submission

On-chain Submissions recorded for this bounty.

#SolverSubmittedBlockTransaction
1
0x6a5a...86d83d
Oct 6, 2026, 7:21 AM UTC#477515100x7ffd4a34...dc388ef8