Home

Launch: Sept 4, 2026 Deadline: Nov 30, 2026 5 Non-model Species Unseen csRNA-Seq Data

Abstract

EPIC (Eukaryotic Promoter and transcription Initiation prediction Challenge) is a community benchmark that asks teams to predict genome-wide Pol II transcription initiation from DNA sequence alone. EPIC aims to establish a rigorous benchmark for promoter and transcription initiation prediction in understudied species, advance sequence-based genomic modelling, and improve our understanding of how regulatory information is encoded in eukaryotic genomes. Participants will train models on experimental transcription initiation data from five non-model animal species and predict initiation signals on held-out genomic regions. Submitted predictions will be evaluated against withheld experimental measurements using classification and correlation-based metrics.

The challenge uses five non-model metazoans — Pacific oyster (Magallana gigas), California two-spot octopus (Octopus bimaculoides), Indianmeal moth (Plodia interpunctella), large milkweed bug (Oncopeltus fasciatus), and Pacific spiny dogfish (Squalus suckleyi) — using unpublished csRNA-seq from the Duttke lab, with matched sRNA-seq provided as a background control but excluded from scoring. Training and test data are per-chromosome/contig bed-tracks (strand-separated, two replicates); the test set is GC-matched to train, with the GC content computed for repeat-free sequences. Teams submit one value per position-and-strand for held-out contigs. Download genomes & data from Zenodo doi:10.5281/zenodo.22285753.

Scored on two tasks — binary presence of initiation (AUPRC) and quantitative initiation efficiency at true positives (Spearman) — aggregated by log-ranks as in IBIS to award gold/silver/bronze medals. Half the test labels will drive a live leaderboard, and the other half will be used in the final evaluation; test data (the ground truth read count profiles) will remain closed until winners are announced.

Only the provided genome sequence may be used, with pretrained genomic models (AlphaGenome, Evo2) as the sole exception; teams register via GitHub (1–10 members, one spokesperson), all must supply a method write-up, medalists must additionally supply reproducible training and scoring code, and everyone clearing the baseline is invited to join the EPIC consortium authorship on the post-challenge paper. A published Nematostella vectensis dataset plus scoring scripts serve as an offline replica of the evaluation.

Read full rules →