General rules
The participants submit predicted transcription initiation levels at particular genomic positions, that is, the relative counts of 5' ends of the csRNA-Seq reads. The participants must rely on the training data provided by the challenge organizers (the genome sequences, the RepeatMasker tracks, and the preprocessed csRNA-Seq and sRNA-Seq tracks for the training subset of chromosomes/contigs, see below). The usage of external data (including existing gene annotations) is not allowed, except for the pre-trained genomic models such as AlphaGenome or Evo2. Any pretrained model is permitted (including fine-tuning and probing), but teams must be ready to disclose the model and its version/checkpoint date in their write-up.
For model training, to allow for accounting for the experimental variability, we will provide the data from two independent experimental replicates. For evaluation (testing), we will merge the data from the same replicates. The participants are welcome to build a species-specific model or any type of multi-species model instead, as long as the submissions technically pass the formatting requirements and do not rely on external data.
The test data will not be published until the announcement of the winners. The organizers reserve the right to assess their own solutions at the post-challenge stage, but these solutions will not be included in the model ranking nor affect the selection of the winners of the challenge.
All teams must provide a method write-up accompanying the Final submission. The winning teams (gold-silver-bronze medalists) must provide a reproducible protocol/pipeline/code to derive the submitted files from the training data, including a scoring tool to scan DNA sequences with the developed model. All participating teams are not obliged but very welcome to open the model training and prediction code. The winners will be personally invited to co-author the post-challenge manuscript aiming at a high-impact journal. All participants of the challenge who submitted at least one valid solution scoring above the basic baseline will be invited to the EPIC consortium, serving as joint authorship of the post-challenge paper.
Anyone except the organizers (and their lab members) can assemble a team. Each participating team should register online with a single spokesperson submitting solutions on behalf of the team, which can include from 1 to 10 members. Registration is possible only by authorizing with an active GitHub account. We kindly ask you to avoid using multiple accounts for a single team and registering with provocative, impolite, or offensive team names. During the Leaderboard stage, each team will be limited to 10 submissions per day (each submitted file per species counts as one submission). This limit might be raised depending on the feedback from the community. At the final stage, we will evaluate only the last submission per species from each team made before the leaderboard closure.