Running ASPEN

To effectively run the ASPEN (ATAC-Seq PipEliNe) on the Biowulf High-Performance Computing (HPC) system, please follow the detailed user guide below:

🛠️ Prerequisites

  • Biowulf Account: Ensure you have an active Biowulf account.
  • Data Preparation: Store your raw ATAC-Seq paired-end FASTQ files in a directory accessible from Biowulf.

🌐 Setting Up the Environment

🚀 Load the ASPEN Module on Biowulf

To access ASPEN, load the ccbrpipeliner module:

module load ccbrpipeliner/8

This command adds aspen to your system's PATH, allowing you to execute pipeline commands directly.

Note: If you're operating outside of Biowulf, ensure that dependencies such as snakemake, python, and singularity are installed and accessible in your system's PATH.

📝 Create a Sample Manifest

ASPEN requires a sample manifest file (samples.tsv) to identify and organize your input data. This tab-separated file should include the following columns:

  • replicateName: A unique name for each biological replicate — i.e. each independently processed biological sample (separate cell culture, animal, or patient). Every row must have a distinct value.
  • sampleName: The condition or group label shared by all biological replicates from the same experimental group (e.g. CONTROL or TREATMENT). Multiple rows will share the same sampleName.
  • path_to_R1_fastq: Absolute path to the Read 1 FASTQ file.
  • path_to_R2_fastq: Absolute path to the Read 2 FASTQ file (required for paired-end data).

Note

Symlinks for R1 and R2 files will be created in the results directory, named as <replicateName>.R1.fastq.gz and <replicateName>.R2.fastq.gz, respectively. Therefore, original filenames do not need to be altered.

Note

The replicateName is used as a prefix for individual peak calls, while the sampleName serves as a prefix for consensus peak calls.

Biological vs. technical replicates

ASPEN expects one row per biological replicate. If you sequenced the same sample across multiple lanes or sequencing runs (technical replicates), you must concatenate those FASTQ files into a single file before creating your manifest — ASPEN does not merge lanes internally.

Biological replicates are independent samples:

  • Biological: independent biological samples (separate cultures, animals, patients, etc.). Use one row per sample in samples.tsv.
  • Technical: the same sample re-sequenced across multiple lanes or runs. cat the FASTQs together first, then use one row.

Example of concatenating technical replicates before running ASPEN:

cat sample1_L001_R1.fastq.gz sample1_L002_R1.fastq.gz > sample1_R1.fastq.gz
cat sample1_L001_R2.fastq.gz sample1_L002_R2.fastq.gz > sample1_R2.fastq.gz

DESeq2 (used in diffatac) requires at least 2 biological replicates per group. Technical replicates do not count as biological replicates and will not satisfy this requirement.

Note

For differential ATAC analysis, create a contrasts.tsv file with two columns (Group1 and Group2 ... aka Sample1 and Sample2, without headers) and place it in the output directory after initialization. Ensure each group/sample in the contrast has at least two biological replicates, as DESeq2 requires this for accurate contrast calculations.

🏃 Running the ASPEN Pipeline

ASPEN operates through a series of modes to facilitate various stages of the analysis.

🗂️ Initialize the Working Directory

Begin by initializing your working directory, which will house configuration files and results. Replace with your desired output directory path:

aspen -m=init -w=<path_to_output_folder>

This command generates a config.yaml and a placeholder samples.tsv in the specified directory. Edit these files to reflect your experimental setup, replacing the placeholder samples.tsv with your prepared manifest. If performing differential analysis, include the contrasts.tsv file at this stage.

Note

To explore all possible options of the aspen command you can either run it without any arguments or run aspen --help

Here is what help looks like:

Note

This is illustrative example output captured at doc-writing time — exact values (e.g. pipeline_home, git commit/tag, aspen_version) will differ depending on which ASPEN version/branch is installed at your site. Run aspen --help yourself to see the current values for your installation.

##########################################################################################

Welcome to

╔══════════════════════════════════╗
║  ASPEN PIPELINE                  ║
║  v1.3.0                          ║
╚══════════════════════════════════╝

ATAC-Seq Analysis Pipeline

##########################################################################################

This pipeline was built by CCBR (https://bioinformatics.ccr.cancer.gov/ccbr)
Please contact Vishal Koparde for comments/questions (vishal.koparde@nih.gov)

##########################################################################################

Here is a list of genome supported by aspen:

  * hg19          [Human]
  * hg38          [Human]
  * mm10          [Mouse]
  * mmul10        [Macaca mulatta(Rhesus monkey) or rheMac10]
  * bosTau9       [Bos taurus(cattle)]
  * hs1           [Human T2T-CHM13]
  * hs1_chrR      [Human T2T-CHM13 + chrR rDNA unit]

aspen calls peaks using the following tools:

 * MACS2
 * Genrich        [RECOMMENDED FOR USE]

USAGE:
  bash ./aspen -w/--workdir=<WORKDIR> -m/--runmode=<RUNMODE>

Required Arguments:
1.  WORKDIR     : [Type: String]: Absolute or relative path to the output folder with write permissions.

2.  RUNMODE     : [Type: String] Valid options:
    * init      : initialize workdir
    * dryrun    : dry run snakemake to generate DAG
    * run       : run with slurm
    * runlocal  : run without submitting to sbatch
    ADVANCED RUNMODES (use with caution!!)
    * unlock    : unlock WORKDIR if locked by snakemake NEVER UNLOCK WORKDIR WHERE PIPELINE IS CURRENTLY RUNNING!
    * reconfig  : recreate config file in WORKDIR (debugging option) EDITS TO config.yaml WILL BE LOST!
    * reset     : DELETE workdir dir and re-init it (debugging option) EDITS TO ALL FILES IN WORKDIR WILL BE LOST!
    * printbinds: print singularity binds (paths)
    * local     : same as runlocal

Optional Arguments:

--genome|-g     : genome eg. hg38
--manifest|-s   : absolute path to samples.tsv. This will be copied to output folder                    (--runmode=init only)
--help|-h       : print this help

Example commands:
  bash ./aspen -w=/my/output/folder -m=init
  bash ./aspen -w=/my/output/folder -m=dryrun
  bash ./aspen -w=/my/output/folder -m=run

##########################################################################################

VersionInfo:
  python          : python/3.10
  snakemake       : snakemake
  pipeline_home   : /data/CCBR_Pipeliner/Pipelines/ASPEN/feature_spikeins
  git commit/tag  : 8d197d39927be3f60558911bd8b2756f36835deb    v1.3.0
  aspen_version   : v1.3.0

##########################################################################################

⚙️ Configurational Changes

MACS2

MACS2 parameters can be changed by editing this block in the config.yaml:

macs2:
    extsize: 200
    shiftsize: 100
    p: 0.01
    qfilter: 0.05
    annotatePeaks: True

Genrich

Genrich parameters can be changed by editing this block in the config.yaml:

genrich:
    s: 5
    m: 6
    q: 1
    l: 100
    g: 100
    d: 100
    qfilter: 0.05
    annotatePeaks: True

Contrasts

If contrasts are to be calculated then fixed-width peaks are used with the following changeable options:

# peak fixed width
fixed_width: 500

# contrasts info
contrasts: "WORKDIR/contrasts.tsv"
contrasts_fc_cutoff: 2
contrasts_fdr_cutoff: 0.05

🧬 Enabling Spike-In Normalization (Optional)

ASPEN supports spike-in normalization, which is useful for controlling technical variability or comparing global shifts in chromatin accessibility across samples. Spike-in reads (e.g., from Drosophila melanogaster or E. coli) are aligned separately and used to compute normalization factors that are applied to host genome accessibility counts.

Not sure whether you need this for your experiment? See "Should I turn on spike-in normalization?" in the Overview docs for a decision guide.

To enable spike-in normalization, edit the config.yaml file that was generated during init. You can find it in your output directory (<path_to_output_folder>/config.yaml).

Open the file and locate the following lines:

spikein: False
# spikein: True

spikein_genome: "dmelr6.32" # Drosophila mel.
# spikein_genome: "ecoli_k12" # E. coli

To activate spike-in normalization:

  1. Set spikein to True
  2. Uncomment or change the spikein_genome to match your experiment

📝 Example Configuration

For Drosophila spike-in:

spikein: True
spikein_genome: "dmelr6.32"

For E. coli spike-in:

spikein: True
spikein_genome: "ecoli_k12"

Note: The spike-in genome must be pre-indexed and available in the reference directory used by ASPEN. Please contact the pipeline maintainers if you need to add a new spike-in genome.

Once enabled, ASPEN will:

  • Align reads to both the host and spike-in genomes.
  • Quantify spike-in counts per sample.
  • Normalize accessibility counts using spike-in-derived scaling factors.
  • Report both normalized and raw counts in the output tables and reports.

🛠️ Dry Run the Pipeline

Before executing the full analysis and after editing the config.yaml as needed, perform a dry run to visualize the workflow and identify potential issues:

aspen -m=dryrun -w=<path_to_output_folder>

What successful dry-run output looks like

Example ASPEN dry-run console output

The dry-run summary confirms that ASPEN finished the preflight checks successfully, points you to dryrun.log for the full transcript, and ends with the exact run command to use next.

In practice, the wrapper prefixes mean:

  • STEP starts a major wrapper phase.
  • OK confirms that phase completed successfully.
  • INFO reports status details such as where output was written.
  • NEXT gives the follow-up command or monitoring action ASPEN expects you to take.

For dryrun, the short terminal summary is only the high-level result. The full dry-run transcript is written to dryrun.log; see the outputs page for the complete workdir artifact reference.

This step outlines the sequence of tasks (Directed Acyclic Graph - DAG) without actual execution, allowing you to verify the planned operations.

🚀 Execute the Pipeline

If the dry run output is satisfactory, proceed to execute the pipeline:

aspen -m=run -w=<path_to_output_folder>

What successful run submission output looks like

Example ASPEN run submission console output

The run summary shows that ASPEN created the submission script, wrote the pipeline.running marker updates during submission, and printed the NEXT monitoring hints for squeue, snakemake.log, and pipeline.status.json.

Those lines map directly to files in WORKDIR:

  • submit_script.sbatch is created before the master Slurm job is submitted.
  • pipeline.running is the human-readable status marker updated during execution.
  • pipeline.status.json is the machine-readable sidecar for scripting and automation.
  • snakemake.log is the detailed workflow execution log once the run starts.

Treat the NEXT lines as the wrapper's built-in "what should I do now?" guidance. They point you to the same files and commands documented below and in the outputs page, without requiring you to remember the monitoring commands yourself.

This command submits a master job to the Slurm workload manager, which orchestrates the entire analysis workflow, managing job submissions and monitoring progress.

ASPEN submits the generated Snakemake job with rule-level resources exposed to Slurm. The heavy rules that may need extra scratch space now define an attempt-aware resources.gres value, so the first submission uses the baseline cluster.json request and later retries can scale the requested lscratch allocation automatically if a rule is retried by Snakemake. This keeps the default configuration in cluster.json while still allowing individual rules to become more conservative on repeated failures.

  • 🛠️ Optional Argument:

--singcache or -c: Override the Singularity cache directory. On Biowulf, when you load the ccbrpipeliner module, ASPEN already uses SIFCACHE=/data/CCBR_Pipeliner/SIFS for you, so you usually do not need -c. Use -c only on another HPC system if you want to point ASPEN at your own cache directory or pull containers yourself.

💡 Example Command:

aspen -m=run -w=<path_to_output_folder> -c /data/${USER}/.singularity

This example is for a non-Biowulf HPC system where you want to manage your own Singularity cache location.

📊 Monitor ASPEN Runs

For day-to-day status checks, use the sidecar and pipeline.* files below first. They are the fastest and most reliable way to see whether ASPEN is running, completed, or failed. Reach for squeue and scontrol only when you want an advanced scheduler-level view or need to inspect an individual SLURM job.

If you do need to inspect the cluster directly, squeue shows the queue state and scontrol exposes detailed job metadata.

📝 Pipeline State Markers: the primary status check

ASPEN writes a set of state-tracking files directly into WORKDIR while a run is executing, so you can check status from the sidecar and pipeline.* files first, even without Slurm access (for example, from a laptop over ssh):

  • pipeline.running, pipeline.completed, pipeline.failed, pipeline.canceled — exactly one of these marker files exists at a time, reflecting the current state. While the pipeline is running, pipeline.running is periodically refreshed by a background progress monitor with a human-readable summary, including the percentage of Snakemake steps completed so far:

    cat <path_to_output_folder>/pipeline.running
    
  • pipeline.status.json — a machine-readable sidecar with the same information (state, reason, slurm_job_id, start/end timestamps, duration_seconds, tasks_done/tasks_total, exit_code), useful for scripting/automation:

    cat <path_to_output_folder>/pipeline.status.json
    
  • snakemake.log.jobby / snakemake.log.jobby.short — a jobby TSV summary of per-rule/job resource usage, generated as a best-effort step after the run finishes (even if the Slurm submission itself failed before Snakemake started).

To view all your active and pending jobs, execute:

squeue -u $USER --format="%.18i %.30j %.11P %.15T %.10r %.10M %.10l %.5D %.5C %.10m %.25b %.8N" --sort=-S

This command lists all jobs submitted by your user account, displaying details such as job IDs, partitions, job names, user names, job states, and the nodes allocated.

For more granular information about a specific job, including its child jobs spawned by ASPEN, use the scontrol command:

scontrol show job <jobid>

Replace with the specific Job ID of interest. This will provide comprehensive details about the job's configuration and status, aiding in effective monitoring and management of your ASPEN pipeline processes.

To quickly gauge the process of the entire pipeline run:

grep "done$" <path_to_output_folder>/snakemake.log