Handle high-throughput long-read DNA and RNA data with flashalign.
git clone https://github.com/97YearsOldProgrammer/flashalign.git
cd flashalign
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --target flashalign -j
# long reads against a reference genome
./build/flashalign align -a ref.fa reads.fq > aln.sam
# create an index first and then map
./build/flashalign index -x lr:hq ref.fa # writes ref.fa.faix
./build/flashalign align -a ref.fa.faix hifi.fq.gz > aln.sam
# use presets
./build/flashalign align -ax lr ref.fa ont.fq.gz > aln.sam # Oxford Nanopore genomic reads
./build/flashalign align -ax lr:hq ref.fa hifi.fq.gz > aln.sam # PacBio HiFi genomic reads
./build/flashalign align -ax splice ref.fa cdna.fq.gz > aln.sam # spliced long reads (strand unknown)
./build/flashalign align -ax splice:hq -u f ref.fa isoseq.fq.gz > aln.sam # PacBio Iso-Seq (transcript strand)
./build/flashalign align -ax splice --junc-bed anno.bed ref.fa cdna.fq.gz > aln.sam # use annotated junctions
./build/flashalign index -x splice -k14 ref.fa ref.k14.faix # noisy Nanopore direct RNA-seq: a k14 index,
./build/flashalign align -ax splice -uf ref.k14.faix drna.fq.gz > aln.sam # then map on it
./build/flashalign align -cx asm5 ref.fa asm.fa > aln.paf # assembly to assembly/ref alignment (experimental)
./build/flashalign align -x ava-ont reads.fq reads.fq > ovlp.paf # Nanopore read overlap (experimental)
./build/flashalign align -x ava-hifi reads.fq reads.fq > ovlp.paf # PacBio HiFi read overlap (experimental)
# man page for detailed command line options
man ./flashalign.1For Linux x86-64, a precompiled binary is on the release page. It runs on any x86-64 Linux with glibc 2.17 or newer and needs no installation:
curl -L https://github.com/97YearsOldProgrammer/flashalign/releases/download/v0.1.0/flashalign-0.1.0_x64-linux.tar.bz2 | tar -jxvf -
./flashalign-0.1.0_x64-linux/flashalign versionFlashAlign is on Bioconda, for Linux and macOS on x86-64 and ARM:
conda install -c conda-forge -c bioconda flashalignThe Python package (see the Developers' guide) installs with pip. Prebuilt wheels cover Linux on x86-64 and ARM and macOS on Apple silicon, for Python 3.12 or newer:
pip install flashalignOtherwise, FlashAlign is built from source with CMake 3.18 or newer, a C and C++17 compiler and zlib 1.2.9 or newer:
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --target flashalign -j
./build/flashalign versionThe result is one binary, build/flashalign. The same commands build on x86-64, where the
alignment kernels use SSE4.1, and on ARM, where they go through the bundled sse2neon shim.
Without options, flashalign align takes a reference and a read file, maps under the lr
preset and writes approximate mappings in the PAF format. No base-level alignment is done, so
there is no CIGAR and the coordinates are approximate:
flashalign align ref.fa reads.fq > approx-mapping.pafYou can ask for the CIGAR in the cg tag of PAF:
flashalign align -c ref.fa reads.fq > aln.pafor for base-level alignments in the SAM format:
flashalign align -a ref.fa reads.fq > aln.samReads may be FASTA or FASTQ, plain or gzip'd, and several read files may follow the reference.
A FASTA reference is indexed in memory on every run. To index it once, write the index with
flashalign index and give the index in place of the reference:
flashalign index -x lr:hq ref.fa # indexing; writes ref.fa.faix
flashalign align -a ref.fa.faix reads.fq > aln.sam # alignmentThe index records the preset it was built with, and align maps under that preset unless -x
is given. The index owns the seeding: -k, -s and -I are indexing options and can't be
changed during mapping, and an explicit -x at mapping time keeps the index's own -k and
-s.
FlashAlign is tuned to a read type with a preset, -x, which sets several options at the
same time. The default is lr.
flashalign align -ax lr ref.fa ont.fq.gz > aln.sam # Oxford Nanopore reads
flashalign align -ax lr:hq ref.fa hifi.fq.gz > aln.sam # PacBio HiFi readslr is for noisy long reads of ~10% error rate. lr:hq is for accurate long reads with an
error rate below 1%, such as PacBio HiFi reads. The two differ in seeding (-s9 and -s5)
and in scoring.
flashalign align -ax splice ref.fa cdna.fq.gz > aln.sam # Nanopore cDNA
flashalign index -x splice -k 14 ref.fa ref.k14.faix # Nanopore direct RNA: a k14 index
flashalign align -ax splice -u f ref.k14.faix drna.fq.gz > aln.sam
flashalign align -ax splice:hq -u f ref.fa isoseq.fq.gz > aln.sam # PacBio Iso-Seq-u b, the default, looks for splice sites on both strands; -u f looks on the transcript
strand only, for reads already on that strand, such as direct RNA and Iso-Seq. For noisy
Nanopore direct RNA, an index built with -k 14 finds more junctions, short first and last
exons above all. splice:hq differs from splice in scoring, in the vote's diagonal bin (48
against 64), in the fine chain's query gap (10,000 against 20,000; -g sets it) and in the
MAPQ coverage threshold (0.70 against 0.55).
FlashAlign can take annotated junctions and prefer them during base alignment:
paftools.js gff2bed anno.gtf > anno.bed
flashalign align -ax splice --junc-bed anno.bed ref.fa cdna.fq.gz > aln.sam--junc-bed takes gene annotations in the 12-column BED format, which paftools.js gff2bed
converts from GTF or GFF3, or intron positions in 6-column BED with the strand column. A splice
donor or acceptor found in the annotation gets a score bonus, --junc-bonus (9 by default).
flashalign align -cx asm5 ref.fa asm.fa > aln.paf # assembly to assembly/ref alignmentasm5, asm10 and asm20 are for an average divergence of about 0.1%, 1% and several percent,
with minimap2's scoring for its presets of the same names.
flashalign align -x ava-ont reads.fq reads.fq > ovlp.paf # Oxford Nanopore read overlap
flashalign align -x ava-hifi reads.fq reads.fq > ovlp.paf # PacBio HiFi read overlapThe output is PAF without CIGAR, each overlap once, from the shorter read (--dual=yes prints it
from both reads); -a and -c are refused. Read names must be unique.
ava-ont seeds with -k17 -s9 and ava-hifi with -k21 -s5; a read set mapped more than once is
indexed once with flashalign index -x ava-ont reads.fq reads.faix, in parts with -I for a large
set (Multi-part index). The Python module has no overlap presets.
PAF is the default and -a writes SAM. -o FILE outputs alignments to FILE [stdout],
whatever its name. For sorted BAM, pipe SAM to samtools:
flashalign align -a -x lr:hq ref.fa hifi.fq.gz | samtools sort -o aln.bamUnmapped reads are written to SAM unless --sam-hit-only is given, and to PAF only with
--paf-no-hit. Secondary alignments are written only with --secondary=yes. --cs, --MD and
--eqx add the cs tag, the MD tag and =/X CIGAR operators. The PAF columns, the tags and
the cs operations are listed under OUTPUT FORMAT in the manual.
--show-config prints every resolved option with the source of its value, then exits without
loading the reference or the reads; given an index, it reads only the index header:
flashalign align -x lr:hq --show-config
flashalign align --show-config ref.fa.faix-I caps the reference bases loaded into memory for indexing. A reference longer than that is
indexed in parts of whole contigs, and align maps the reads against one part at a time,
reading them once per part:
flashalign index -I 4G ref.fa
flashalign align -a ref.fa.faix reads.fq > aln.samMapping quality is incorrect given a multi-part index.
The manual page, flashalign.1, describes every option, the output format and
the exit status (man ./flashalign.1). flashalign align --help and flashalign index --help
list the common options. Bugs and questions go to the
issue page.
The manuscript is not yet public. Until then, please cite the software: doi:10.5281/zenodo.22985251.
FlashAlign is also a C++ library, with its public headers in
include/flashalign/, and a Python package, flashalign, that drives
the same library:
python -m pip install .
python demo.py ref.fa reads.fqThe package needs CPython 3.12 or newer and installs no flashalign command.
demo.py shows typical use; help(flashalign.Aligner) and
python/flashalign/_flashalign.pyi document the API.
- Built for long reads: a read has to carry enough anchors to be collapsed into one or more diagonals. Short reads with enough anchors also work, but without a speed advantage over minimap2.
- FlashAlign requires SSE4.1 instructions on x86 CPUs or NEON on ARM CPUs. A build without them is not provided.
- The assembly and overlap presets are experimental.
FlashAlign is released under the MIT License; see LICENSE.
THIRD_PARTY_NOTICES.md covers the third-party code.