Abstract

EPIC (Eukaryotic Promoter and transcription Initiation prediction Challenge) is a community benchmark that asks teams to predict genome-wide Pol II transcription initiation from DNA sequence alone. EPIC aims to establish a rigorous benchmark for promoter and transcription initiation prediction in understudied species, advance sequence-based genomic modelling, and improve our understanding of how regulatory information is encoded in eukaryotic genomes. Participants will train models on experimental transcription initiation data from five non-model animal species and predict initiation signals on held-out genomic regions. Submitted predictions will be evaluated against withheld experimental measurements using classification and correlation-based metrics.

The challenge uses five non-model metazoans — Pacific oyster (Magallana gigas), California two-spot octopus (Octopus bimaculoides), Indianmeal moth (Plodia interpunctella), large milkweed bug (Oncopeltus fasciatus), and Pacific spiny dogfish (Squalus suckleyi) — using unpublished csRNA-seq from the Duttke lab, with matched sRNA-seq provided as a background control but excluded from scoring. Training and test data are per-chromosome/contig bed-tracks (strand-separated, two replicates); the test set is GC-matched to train, with the GC content computed for repeat-free sequences. Teams submit one value per position-and-strand for held-out contigs, and are scored on two tasks — binary presence of initiation (AUPRC) and quantitative initiation efficiency at true positives (Spearman) — aggregated by log-ranks as in IBIS to award gold/silver/bronze medals. Half the test labels will drive a live leaderboard, and the other half will be used in the final evaluation; test data (the ground truth read count profiles) will remain closed until winners are announced. Only the provided genome sequence may be used, with pretrained genomic models (AlphaGenome, Evo2) as the sole exception; teams register via GitHub (1–10 members, one spokesperson), all must supply a method write-up, medalists must additionally supply reproducible training and scoring code, and everyone clearing the baseline is invited to join the EPIC consortium authorship on the post-challenge paper. A published Nematostella vectensis dataset plus scoring scripts serve as an offline replica of the evaluation.

EPIC aim

We aim at assembling a community-driven suite of computational tools for evaluation of genome-wide transcription initiation site prediction software in non-model eukaryotic species. An open infrastructure of this kind will advance our understanding of promoter architectures and transcriptional initiation mechanisms, as well as their respective stability and plasticity across the evolutionary tree of life.