Skip to content

Repository files navigation

Falcon Audio Codec

Falcon v2.0.1: an open-source MDCT audio codec, MIT licensed. Opus-class quality at equal size; decodes about 1.5× faster than libopus.

Read this in 日本語 (Japanese).

Falcon is an MDCT audio codec written in Rust and released under the MIT license. It follows one design rule: keep the decoder as light as possible, and buy quality back with encoder-side work.

On an 18-file test corpus, with every Falcon file no larger than the Opus file it is compared against:

  • Quality is on par with Opus, or better, depending on the metric. Falcon has the lower MR-STFT distance on 18 of 18 files and the lower log-mel distance on 13 of 18. A multi-detector perceptual A/B check finds no significant difference overall.
  • Decoding is faster than libopus, libvorbis and Tremor, and on par with stb_vorbis on x86-64. In a library-level benchmark on one thread, Falcon needs 0.66× the decode time of libopus and 0.79× that of libvorbis on an x86-64 laptop (0.68× and 0.84× on Apple M5). stb_vorbis is about as fast on x86-64 and 1.28× faster on Apple M5. Falcon uses more memory per stream than any of them (see Decode speed).

Version 2 redesigns the codec to avoid specific third-party patents (see Patents) and contains no code taken from other codec implementations.


Documentation: API reference · Changelog · Patent position · Third-party licenses · C/C++ bindings · Contributing

Table of contents


Highlights

Property Falcon 2.0
Quality vs Opus, equal file size Lower MR-STFT distance on 18/18 files and lower log-mel L1 on 13/18. The perceptual A/B check gives "same" overall: no detector is significantly worse, and harmonic amplitude modulation is better on 7 of 9 pitched files and worse on 2 (p = 0.18).
Decode speed (one thread, library level) 1114× real time over the corpus on an x86-64 laptop core and 1551× on Apple M5. About 1.5× faster than libopus and 1.2–1.35× faster than libvorbis and Tremor; stb_vorbis is on par on x86-64 and 1.28× faster on Apple M5.
Decoder memory About 320–345 KB of heap per stereo stream, plus a copy of the compressed file when decoding through the C API: more than libopus (27 KB) or stb_vorbis (about 210 KB).
Transform MDCT, 960 coefficients per 20 ms frame at 48 kHz, Vorbis power-complementary window. Transient frames switch to four short blocks (Edler window switching, signalled at zero bit cost).
Band energies 27 bands, 1 dB grid. Time DPCM on inter frames and frequency DPCM after a random-access point. Static canonical Huffman codes (8 tables) behind a per-channel all-zero bit.
Coefficients Gain-shape coding. The pulse count K is derived from the transmitted energies, so it costs no side bits. A static cell table per band, the per-cell pulse counts as one stars-and-bars composition index, and each cell as a Fischer pyramid enumeration index or dense Golomb-Rice magnitudes.
Noise and holes Encoder-signalled per-band coding classes (delta-coded): pulses only, pulses plus zero-bin fill at one of two fixed levels, or noise substitution. The decoder takes no fill decisions of its own.
Stereo Per-frame mid/side coding in the MDCT domain. The side channel's step takes half of its reference from the transmitted mid energies.
Special frames Bit-exact LSB-floor payload for 16-bit noise-floor frames. An overflow-extension mode (Ext) keeps every frame within the 12-bit frame-size field for up to 64 channels.
Integration Rust crates, a CLI, and a C ABI with a header-only C++17 wrapper (fuzz-tested against garbage and truncated input; MSVC and MinGW).
License MIT. No third-party codec code; every dependency is permissively licensed.

Every quality number in this README is measured at equal or smaller file size than the Opus reference. See Evaluation methodology for exactly what was measured and how, and Limitations for where Falcon is behind.

What's new in 2.0

2.0 replaces the coefficient, energy, fill and entropy layers of 1.x. It has two goals:

  1. Design around specific patents that a review found could read on 1.x.
  2. Remove the parts of 1.x whose code followed libopus.
1.x 2.0 Why
Two-step "spreading" rotation of the decoded pulses (CELT/Opus style) Removed. Zero-bin fill is controlled by an encoder-signalled class Designed around US 8,838,442. The rotation code followed libopus.
Recursive band halving with a quantized angle (theta) per node Static cell tables + one composition index + per-cell enumeration or Golomb-Rice Designed around US 9,009,036. No split recursion, no angles.
Decoder-side noise injection into bands that received no pulses ("anti-collapse") Encoder-signalled noise substitution (class 3) Designed around US 9,015,042. The decoder does not inspect energies to decide on noise.
rANS entropy coding Static canonical Huffman + Golomb-Rice Designed around US 11,234,023.
Mixed time/frequency energy predictor Pure time DPCM + frequency DPCM on random-access frames Removes a frame-rate flutter of harmonic amplitudes, and simplifies the energy path.
Pulse search using the libopus metric New search: rounding + remainder repair + cosine-improving pulse moves Clean-room rewrite (encoder only).
Transient detector: an attack in frame f also marked frame f+1 for short blocks Each frame decides its block type from its own window span only; nothing carries over between frames Designed around US 9,495,971 (transient hangover). Encoder only.
Raw scalar overflow fallback Ext frames: the normal layout at a coarser rate level The old fallback did not fit the frame-size field for many channels.

Compatibility: the bitstream is now version 0x10. 2.x decoders reject 1.x (0x04) files and 1.x decoders reject 2.x files. See Upgrading from 1.x.

Honesty note on 1.x: 1.x advertised that it beat Opus on all 18 files, based on MR-STFT alone. That claim did not hold up under the perceptual A/B check that 2.0 adds to the methodology. 2.0 reports all three views: MR-STFT, log-mel, and the perceptual check.

Design philosophy

"Make the decoder extremely lightweight; achieve quality through encoder optimization."

  • The decoder executes; it does not decide. Per frame it reads static prefix codes, decodes integer pulse vectors, applies one scaling pass per band (plus the fill that the encoder chose), and runs one inverse MDCT (or four short ones). It has no psychoacoustic model, no search, no adaptive tables and no heuristics that inspect the signal. This is what makes it cheap, and it also keeps it simple to audit.
  • The encoder is allowed to work hard. Silence and transient analysis, per-frame stereo decisions, the per-band coding-class decision by analysis-by-synthesis, the pulse search and a global rate-control search all run at encode time.
  • Zero side information where possible. The pulse count of every band is a deterministic function of data the decoder already has: the transmitted band energies and the global quality index. Only what cannot be derived is sent: the energies, the coding classes, and the pulse vectors.

Repository layout

Path What it is
falcon_core/ Shared kernel and the single source of truth for everything both sides must agree on: bitstream I/O, band tables, energy prediction, prefix codes (prefix), the step and pulse-count law (alloc, pvq), fixed-cell shape coding (cells), coding classes and fills (fill), window switching (transient), the LSB-floor payload (lsbfloor), and the rate law (rate)
falcon_encoder/ Encoder: MDCT analysis, silence detection (mode_detect), class decision (classify), transient look-ahead, M/S decision, rate control (encode)
falcon_decoder/ Decoder: frame parsing, reconstruction, fast inverse MDCT
falcon_cli/ The falcon command-line tool (encode, decode, info, bench-decode) and the integration tests
falcon_capi/ C ABI + header-only C++17 wrapper, CMake package, examples (README)
docs/ API reference and README images
tools/quality/ Evaluation harness: metrics, honesty gates, synthetic torture-corpus generator, iso-size comparison driver (Python)
tools/eval/ Older Rust A/B harness (falcon-eval)

How Falcon encodes: the signal chain

The encoder processes 20 ms frames (960 samples per channel at 48 kHz). encode_file first scans the whole input: it detects transients (with look-ahead), makes the initial M/S decision, and profiles LSB-floor frames. It then runs the frame loop, inside a rate-control search when a size or bitrate target is set.

1. Silence and the LSB floor

A frame is coded Silent only when both it and the previous frame fall below −100 dBFS of AC (DC-removed) energy. The MDCT window spans two frames, so the previous frame's tail must also be silent. Faint content (−95…−110 dBFS decay tails, room tone) is still coded, and a pure DC pedestal does not count as activity.

16-bit sources often end in a quantization-noise floor: a DC pedestal of about 1 LSB with single-LSB flicker. An MDCT cannot reproduce this at any sensible rate. Falcon codes such frames exactly in the time domain instead. A frame qualifies when every sample lies exactly on the 16-bit grid and within ±5 LSB of a DC level. Its content (per channel, a DC level plus a residual whose first difference is coded as Golomb-Rice zero runs) is carried in the next Silent frame's payload and added to the output after the zero-coefficient IMDCT. Reconstruction is bit-exact. Material that is not on the 16-bit grid never takes this path.

2. Transients and window switching

Each frame decides its own block type from an analysis of its own MDCT window span (the 1920 input samples its window covers, plus 40 ms of history that is re-read from the signal for every frame). The analysis compares 2.5 ms sub-block energies with a decaying envelope: +12 dB marks an attack, and an absolute gate keeps LSB flicker from triggering. No detector state or decision carries over from one frame's analysis to the next. A frame whose span holds an attack is coded as short blocks: four 480-sample MDCTs (240 bins each, interleaved into the usual 960-bin layout), which confines coding noise to about 5–10 ms around the attack. Start and Stop transition windows follow Edler's window switching and are TDAC-exact. As in AAC-style window switching, the frame before a short-block frame takes the Start window. The flag travels in the frame-header mode field (Mode::Flagged), so it costs no bits. The decoder derives each frame's window from the flags of the previous and current frame. Short-block runs are kept away from random-access frames.

3. Stereo: per-frame mid/side

For stereo input the encoder decides M/S or L/R per frame. It uses M/S when the side energy is small against the mid (with hysteresis: enter below 0.20, stay below 0.35). It rejects M/S when the per-band side-to-mid ratio is much higher than the broadband ratio, because then a loud centred bass would hide a wide mix. M/S is applied to the MDCT coefficients. The MDCT is linear and the overlap-add always runs in L/R, so switching per frame is exactly reconstructible.

4. MDCT and band energies

Each channel is transformed with an MDCT under the Vorbis power-complementary window (960 coefficients; 1920-sample span). The 960 coefficients are grouped into 27 bands, Bark-like up to 12 kHz, with the top octave split into 12–15 / 15–18.5 / 18.5–20 / 20–24 kHz so the envelope can follow the real high-frequency tilt. A band's energy is 20·log10(3.5 · rms) in dB.

Energies are quantized on a 1 dB grid with DPCM:

  • Inter frames: pure time prediction, pred_b = previous_b. A steady band repeats its level with a zero residual. (1.x mixed in the lower band's energy. On steady sloped spectra that made the 1 dB rounding alternate, which modulated every partial in the band at the frame rate.)
  • Random-access frames (the first coded frame after a reset): frequency prediction, pred_b = recon_{b−1}, pred_0 = −20 dB. These frames depend on no earlier frame.
  • The residual range is ±127 dB, so a band can jump from the −80 dB floor to full scale in one frame (clicks out of silence) and fall back in the next.
  • Silent frames decay the predictor towards silence identically on both sides.

The residuals are coded with static canonical Huffman codes. There are 8 tables built from discrete-Laplacian models, from very sharp to wide. The code length limit is 12 bits (JPEG Annex K length limiting), and residuals of 16 dB or more are sent as an ESCAPE code plus 8 raw bits. The encoder picks the cheapest table per frame (3 bits in the frame flags). Each channel first sends an all-zero bit, so a fully steady channel costs 1 bit. The tables are built with integer arithmetic at start-up, identically on every platform.

5. From energies to pulse counts (zero side information)

Each band's quantizer resolution is a pure function of the transmitted energies and the global rate scalar K. It is computed identically by encoder and decoder, so no allocation bits are sent. The normalized step of band b is

d_b = K · shape(b) / 10^((ρ·e_max,ref + γ·(e_ref − e_max,ref) + (e_b − e_ref)) / 20)
  • e_b is the band energy and e_max the frame's peak band energy (dB).
  • The exponents ρ = 0.5 (across frames) and γ = 0.32 (within a frame) are sub-linear. They buy relative, log-spectral precision instead of absolute precision.
  • shape(b) is a fixed perceptual weight: A-weighting-like, held high from 3.7 to 20 kHz, and 6× finer for the 100–920 Hz fundamentals and low harmonics, where coarse pulses make harmonic amplitudes flutter.
  • Mid-referenced side step: for the side channel of an M/S frame, the loudness reference e_ref is taken half-way to the mid channel's band energy. The side's coding noise is heard against the mid, so it should not be coded far finer than the mid. For all other channels e_ref = e_b.
  • A band more than 96 dB below the frame peak, or whose step reaches 0.7 (4.0 for loud bands above 15 kHz), gets no pulses. What happens to it then is decided by its coding class (step 6).
  • Transient step: a frame whose peak energy jumps by 9 dB or more over the previous frame gets a finer step (× 0.25), and the following frame × 0.8. Both sides derive this from the transmitted energies.

The step is mapped to a pulse count K_b by a fixed law (the expected L1 norm of a dead-zone quantizer on a unit Laplacian source), with at least one pulse per coded band and at most 64 pulses per bin.

6. Encoder-signalled coding classes

Every band of every channel has a coding class:

Class Decoder behaviour
0 — pulses Pulses only; a band with no pulses is silent
1 — low fill Pulses, plus every zero bin set to a random-sign value at 0.2× the band's decoded per-bin RMS; the band is then rescaled to its transmitted energy
2 — high fill Same with 0.45×
3 — noise No shape bits. The band is noise at its transmitted energy, tilted along the neighbouring bands' energies

The classes are a decision of the encoder, and the decoder only executes them. The encoder chooses them per band and frame:

  • A band on the energy floor always stays class 0.
  • A band with no pulses gets class 3, except in the 20–24 kHz band and around onsets. There a 20 ms noise burst would sound like pre- or post-echo, so the band stays silent.
  • A steady, not-too-dense band on a long-window frame is decided by analysis by synthesis. The encoder reconstructs the band exactly as the decoder will for classes 0, 1 and 2, and picks the class whose smoothed log-power spectrum is closest to the input. The input spectrum comes from a shift-invariant odd-frequency DFT, not the MDCT.
  • Classes carry a mild bias towards the current class.

Classes persist per channel. A frame sends a class section (announced by bit 0 of the frame flags) only when something changes. The section is coded as a delta: per channel a changed bit, the number of changes, the runs of unchanged bands (Golomb-Rice), and each new class as one of the three other values. All classes reset to 0 at random-access frames. Short-block frames treat classes 1 and 2 as 0.

7. Shapes: static cells, one composition index, per-cell codes

A band with K_b pulses is coded as an integer vector y of its N bins with Σ|y_i| = K_b. The decoder scales the whole vector to the transmitted band energy, so band energy is exact by construction.

  1. Static cells. Each band is split into cells of at most 16 bins by a fixed table: ceil(N/16) near-equal cells, with a separate table for short-block frames where each cell stays inside one time block. The table depends only on the band, never on K or on bit counts.
  2. Composition. The pulse counts of the cells, (k_0 … k_{C−1}) with Σ k_c = K_b, are sent as one stars-and-bars index: the rank of the bar positions among K_b + C − 1 slots, in truncated binary over C(K_b + C − 1, C − 1) values. Above 2 pulses per bin they are sent cell by cell instead.
  3. Cells. A cell with up to K_ENUM_MAX[n] pulses (a static table, ≤ 64) is sent as its rank among all V(n, k) signed integer vectors of dimension n and L1 norm k (Fischer's pyramid enumeration, truncated binary, u64 arithmetic). A denser cell is sent as Golomb-Rice magnitudes plus sign bits, with the last bin implied by k_c.

On short-block frames the band's bins are first regrouped by time block, so cells follow time. Silent blocks then decode to exact zeros and the attack block takes the budget. The encoder also deletes weak pre-onset residue that could only smear backwards as pre-echo.

8. Pulse search (encoder)

The shape that best matches the band's direction maximizes the cosine (Σ|x_i|·y_i)² / Σy_i². The encoder searches it in three steps:

  1. Round K|x_i|/Σ|x|.
  2. Repair the pulse count using the rounding remainders.
  3. Move single pulses from the smallest to the largest remainder at the current optimal gain, as long as each move strictly raises the cosine. When that stalls, try the exact best remove-then-add move. The loop is capped at N moves.

9. Rate control

A single quality index (0 = finest … 255 = coarsest, plus a 4-bit fraction in the header) sets K on a log scale, and file size is monotone in it. With --target-bytes or --target-kbps, the encoder binary-searches the index, then the fraction, for the finest setting whose whole file fits. The fraction lets it use about 99–100% of a byte budget.

10. Overflow extension (Ext frames)

The frame-size field is 12 bits (at most 4094 payload bytes per frame). Dense many-channel material can exceed that. Such a frame is re-packed from the same analysis at overflow-extension levels 1–6, which raise K geometrically towards a fixed ceiling, until it fits. It is then marked Mode::Ext, with the level in frame-flags bits 4–6. Level 7 carries only the energies, and every band decodes as noise. For example, 64 channels of white noise at the finest quality fit at level 5 or 6. Normal material never reaches this path.

How Falcon decodes

Per frame:

  1. Header, CRC. The decoder reads the 2-byte frame header and verifies the CRC-8 (enabled by default).
  2. Energies. For every channel: the all-zero bit, then 27 Huffman codes (table from the frame flags). It adds the prediction (time DPCM, or frequency DPCM after a random-access point).
  3. Classes. If bit 0 of the flags is set, it applies the class delta.
  4. Pulse counts. It derives every band's K_b from the energies (steps 5 and 10 above).
  5. Shapes. For every band with pulses and class ≠ 3: it reads the composition, then each cell, writes the integer vector, and scales it to the transmitted energy in one pass. Zero bins get the class 1/2 fill, and class 3 bands get tilted noise from a deterministic per-band PRNG.
  6. Stereo. For M/S frames it applies the inverse M/S on the coefficients.
  7. Synthesis. One fast inverse MDCT (N/4-point complex FFT) with the frame's window, or four short ones, then overlap-add.
  8. Silent frames. A zero-coefficient IMDCT keeps the overlap consistent, plus the LSB-floor payload when present.

The decoder keeps only a little state per channel: the IMDCT overlap, the energy predictor, the last peak energy and frames since an onset, the current classes, and the previous frame's transient flag.

Bitstream format

All multi-byte fields are little-endian. Bit fields inside frame payloads are packed LSB-first.

File header (25 bytes)

Field Size Description
Magic 4 B "FALC"
Version 1 B 0x10 (Falcon 2.x). Decoders reject any other value.
Flags 2 B Bits 0–7: quality index (0–255); bits 8–11: fractional quality index (sixteenths of one step); bits 12–15 reserved (0)
Sample rate 3 B Hz (24-bit)
Channels 1 B 1–64
Total frames 4 B Number of frames in the stream
Bitrate 2 B Nominal kbps (informational)
LFE mask 8 B One bit per channel (informational)

Frame

[frame header: 2 B][payload][CRC-8 over the payload: 1 B]

Frame-header field Bits Values
Frame size 12 Payload bytes + 1 (CRC); at most 4095
Mode 2 0 = Mdct, 1 = Flagged (transient: short blocks), 2 = Ext (overflow extension), 3 = Silent
Flags 2 Bit 0: random-access frame (predictors, classes and overlap reset); bit 1: M/S coupling (stereo)

The window of a frame follows from the transient flags of the previous and current frame: Normal, Start, Shorts, or Stop.

MDCT payload (modes 0, 1, 2)

One LSB-first bit stream:

Section Content
Flags (8 bits) Bit 0: class section present; bits 1–3: energy Huffman table (0–7); bits 4–6: Ext rate level (1–7; must be 0 in modes 0/1); bit 7: reserved (0)
Energies Per channel: all_zero (1 bit); if 0, 27 canonical Huffman codes of the zigzagged residuals, where ESCAPE + 8 raw bits covers residuals of 16 dB or more
Classes Only if flags bit 0: per channel a changed bit; if set, number of changes − 1 (Rice k=1), then per change the run of unchanged bands (Rice k=2) and the new class (truncated binary over the 3 other classes)
Shapes For every channel, then every band with K_b > 0 and class ≠ 3 (none at Ext level 7): the composition (one truncated-binary index, or per cell above 2 pulses/bin), then each cell (enumeration index, or dense Golomb-Rice magnitudes + signs)

Nothing else is sent. The pulse counts, cell layout, field widths and fill levels are all derived from data both sides already have.

Silent payload (mode 3)

  • [0]: plain silence.
  • [1], then per channel [dc: i8][k: u8], then one bit stream of Golomb-Rice zero runs of the residual's first difference: this is the LSB-floor payload. It is added to the output of the frame being emitted, which is the previous frame. k = 0xFF means DC only.

Building

Requirements: Rust stable with Cargo. The declared MSRV is 1.70, and development uses current stable. Windows, macOS and Linux are supported.

git clone https://github.com/pbtechlab/Falcon.git
cd Falcon
cargo build --release          # CLI at target/release/falcon(.exe)
cargo test --workspace --release

CPU target: .cargo/config.toml builds for x86-64-v3 (AVX2, FMA, BMI; Intel Haswell / AMD Excavator and later). That is what the published speed numbers use. For older x86-64 CPUs or other targets, override it, e.g. RUSTFLAGS="-C target-cpu=x86-64" cargo build --release, or remove the file. The bitstream format is unaffected.

C / C++ library (static + shared):

cargo build --release -p falcon_capi
cmake -S falcon_capi -B build && cmake --build build --config Release   # examples

See falcon_capi/README.md for the link lines of each toolchain (MSVC, MinGW, gcc/clang).

Command-line interface

falcon encode <input.wav> <output.falcon> [options]
falcon decode <input.falcon> <output.wav>
falcon info <input.falcon>
falcon bench-decode <input.falcon> [-i N]

encode

Reads integer-PCM WAV (16- or 24-bit, any channel count 1–64; the codec is tuned for 48 kHz) and writes a .falcon file.

Option Description
--target-bytes <N> Finest quality whose whole file is ≤ N bytes (binary search). Overrides the kbps options.
--target-kbps <R> Target average bitrate, converted to a byte budget from the duration
-b, --bitrate <KBPS> Default target when neither of the above is given (default 128). It is also written to the header.
--quality-index <0-255> Fixed quality (0 = finest). Disables rate control.
falcon encode in.wav out.falcon                       # ~128 kbps target
falcon encode in.wav out.falcon --target-kbps 96
falcon encode in.wav out.falcon --target-bytes 500000 # exact size budget
falcon encode in.wav out.falcon --quality-index 20    # fixed quality

It prints the resolved target, frames, bytes, average bitrate, encode time and compression ratio. Stereo input uses per-frame M/S automatically.

decode

Writes a 16-bit PCM WAV at the stream's sample rate and channel count, and reports the decode time and real-time factor. The first frame (TDAC warm-up) is dropped, so output sample n lines up with input sample n.

info

Prints the header: sample rate, channels, layout, total frames, nominal bitrate, duration, file size, and the LFE mask if set.

bench-decode

Loads the file into memory and decodes it N times (default 10). It prints the first-iteration time, the average decode time, the real-time factor and the throughput. It is meant for quick checks; the cross-codec numbers in Decode speed come from a C harness that calls the C API.

Rate control

Priority order:

  1. --quality-index sets a fixed quality with no search.
  2. --target-bytes sets an exact budget.
  3. --target-kbps sets a budget from the duration.
  4. --bitrate is the default budget.

Size is monotone in the quality index, so a binary search over the integer index and then over the 4-bit fraction finds the finest setting that fits. If even the finest setting (index 0) is smaller than the budget, the file simply comes out smaller. This happens on the synthetic sweep, which is 0.53× the Opus size at index 0.

Rust API

The workspace crates are falcon_encoder, falcon_decoder and falcon_core. Depend on them by path or git. PCM is f32, nominally in [-1.0, 1.0]; a frame is 960 samples per channel. Full reference: docs/API_REFERENCE.md.

Encode a whole buffer (de-interleaved: one Vec<f32> per channel):

use falcon_encoder::{encode::encode_file, EncoderConfig};
use std::{fs::File, io::BufWriter};

let channels: Vec<Vec<f32>> = load_my_audio();          // e.g. [left, right]
let config = EncoderConfig {
    target_bytes: Some(250_000),   // finest quality that fits 250 kB
    ..EncoderConfig::default()     // or set quality_index for a fixed quality
};
let out = BufWriter::new(File::create("out.falcon")?);
let stats = encode_file(out, &channels, 48_000, config)?;
println!("{} frames, {} bytes", stats.total_frames, stats.bytes_written);

EncoderConfig fields:

Field Default Meaning
target_bytes None Size target for the whole file; enables rate control
quality_index 43 Fixed quality (0 = finest … 255 = coarsest), used when target_bytes is None
quality_frac 0 Fraction of one index step (0–15)
stereo_coupling true Allow per-frame M/S on stereo input
ra_interval 50 Random-access frame every N frames (the CLI and C API use 10000, which in practice means only the first frame)
bitrate 128 Nominal kbps written to the header (informational)

Decode a whole stream:

use falcon_decoder::decode::decode_file;
let (pcm, sample_rate, channels) = decode_file(std::io::BufReader::new(File::open("out.falcon")?))?;
// pcm[c][n]: channel c, sample n (de-interleaved); first frame (warm-up) already dropped

Stream frame by frame:

use falcon_decoder::Decoder;
let mut dec = Decoder::new(reader)?;          // parses and validates the header
while dec.decode_frame()? != 0 {              // 960 samples per channel, 0 at end
    for ch in 0..dec.channels() {
        let frame: &[f32] = dec.get_samples(ch);
    }
}

falcon_encoder::Encoder offers the same frame-by-frame interface for encoding: interleaved frames of channels × 960 samples. It does not do the whole-file look-ahead (transient scheduling, rate control) that encode_file does. All fallible calls return falcon_core::FalconError. Corrupt or truncated streams are designed to produce errors rather than panics; this is fuzz-tested through the C API.

C and C++ API

falcon_capi builds a static and a shared library with a small, stable C99 ABI (falcon_capi/include/falcon.h) and a header-only, exception-free C++17 RAII wrapper (falcon.hpp).

#include <falcon.h>
/* streaming decode */
FalconDecoder* d = falcon_decoder_open(bytes, len);       /* copies the input */
float buf[1024 * 2]; int64_t n;
while ((n = falcon_decoder_read_f32(d, buf, 1024)) > 0) { /* n frames, interleaved */ }
falcon_decoder_seek(d, 48000);                            /* sample-exact */
falcon_decoder_close(d);
/* one-shot encode (target_bytes = 0 -> use target_kbps) */
uint8_t* enc; size_t enc_len;
falcon_encode_buffer(pcm, frames, 2, 48000, 0, 128, &enc, &enc_len);
falcon_free(enc);
#include <falcon.hpp>
falcon::Decoder dec = falcon::Decoder::open(bytes, len);
std::vector<float> block(1024 * dec.channels());
std::int64_t n = dec.read(block.data(), 1024);
  • All functions return error codes (FALCON_OK = 0, negative = error).
  • Buffers returned by the library are freed with falcon_free.
  • A decoder handle is not thread-safe, but separate handles are independent.
  • Garbage and truncated input is fuzz-tested to return errors.
  • Linking has been verified with MSVC and MinGW.

Details, CMake integration and per-toolchain link lines are in falcon_capi/README.md.

Quality at equal size

Iso-size quality vs Opus: Falcon's MR-STFT and log-mel error as a percentage of Opus's for each of the 18 corpus files. MR-STFT is below Opus on all 18; log-mel is below on 13 and above on 5 (sfx1, sfxtransient, harpsichord, bgm1, polychord).

Every Falcon file below was encoded with --target-bytes set to the size of the corresponding Opus file, so Falcon is never larger. Both metrics are distances to the original: lower is better, and bold marks the better codec. The last column is the per-file verdict of the perceptual A/B check (see methodology).

File Type Size vs Opus MR-STFT Falcon / Opus log-mel Falcon / Opus Perceptual A/B (per file)
bgm1 music (game BGM) 1.000 0.454 / 0.638 0.164 / 0.148 Falcon better
sfx1 SFX 0.999 0.444 / 0.626 0.188 / 0.127 same
sfx2 SFX 0.998 0.526 / 0.547 0.187 / 0.189 Falcon better
musicclassical music, classical 1.000 0.511 / 0.603 0.123 / 0.176 Falcon better
musicelectronic music, electronic 0.999 0.222 / 0.251 0.055 / 0.077 Falcon better
musicpop music, pop 0.999 0.677 / 0.883 0.171 / 0.198 Falcon better
sfxnoise SFX, noise 0.998 0.449 / 0.706 0.110 / 0.118 same
sfxtransient SFX, applause 1.000 0.586 / 0.876 0.142 / 0.099 Opus better
voicefemale speech, female 0.999 0.417 / 0.599 0.100 / 0.197 Falcon better
voicemale speech, male 1.000 0.406 / 0.553 0.109 / 0.181 Falcon better
voicesing singing 0.999 0.480 / 0.483 0.090 / 0.106 Falcon better
castanets synthetic 1.000 0.304 / 1.008 0.049 / 0.291 mixed
glockenspiel synthetic 0.999 0.540 / 0.919 0.518 / 1.444 Falcon better
harpsichord synthetic 0.998 0.521 / 0.678 0.175 / 0.128 Opus better
hftonal synthetic 0.999 0.434 / 1.421 0.118 / 0.190 same
pinkbursts synthetic 0.999 0.791 / 4.063 0.583 / 1.751 Opus better
polychord synthetic 1.000 0.435 / 0.883 0.121 / 0.117 Opus better
sweep synthetic 0.532 0.678 / 0.694 0.851 / 1.072 same

Summary: parity with Opus, and ahead on the spectral-distance metrics.

  • MR-STFT: Falcon is lower on 18/18 files.
  • Log-mel L1: Falcon is lower on 13/18. It is higher on bgm1, sfx1, sfxtransient, harpsichord and polychord.
  • Perceptual A/B check, aggregate (sign test over the 18 pairs per detector): the overall verdict is "same". No detector with validated accuracy is significantly worse for Falcon.
Detector (what it measures) Falcon better / worse p (sign test)
Overall-quality (MOS-style) estimate 8 / 10 0.81
Signal-to-noise ratio against the reference 11 / 7 0.48
Noisiness estimate 11 / 7 0.48
Harmonic amplitude modulation against the reference (9 pitched files) 7 / 2 0.18
Hum 4 / 3 1.00
Effective bit depth 0 / 0 1.00
  • Perceptual A/B check, per file: Falcon better on 9 files, Opus better on 4 (harpsichord, pinkbursts, polychord, sfxtransient), mixed on 1, and the same on 4.

Objective metrics are proxies. Listen to your own material.

Decode speed

These numbers come from one library-level benchmark. Every decoder is linked into the same small C program, decodes whole files from memory to PCM in 20 ms reads on one thread, and is timed without file I/O or process start-up. Falcon is called through its C API. The others are the libraries that games ship. The inputs are the files of the quality evaluation: Falcon and Opus at equal size, and Vorbis at its own quality setting.

Total over the 18-file corpus, in times real time (higher is better):

Decoder x86-64, P-core x86-64, E-core x86-64, LP E-core Apple M5, P-core Falcon's time ÷ this decoder's (x86-64 / M5)
Falcon 2.0 (C API) 1114× 661× 434× 1551× 1.00 / 1.00
libopus 1.6.1, float 731× 425× 279× 1056× 0.66 / 0.68
libopus 1.6.1, fixed point 629× 353× 234× 886× 0.56 / 0.57
opusfile 0.12 (Ogg + libopus float) 717× 414× 271× 1022× 0.64 / 0.66
stb_vorbis 1.22 1154× 612× 401× 1978× 1.04 / 1.28
libvorbis 1.3.7 (vorbisfile) 879× 484× 317× 1297× 0.79 / 0.84
Tremor (fixed-point Vorbis) 877× 526× 346× 1146× 0.79 / 0.74
IMA ADPCM (4:1, plain C) 5613× 2401× 1580× 9727× 5.0 / 6.3
  • Per file: Falcon is faster than libopus (float and fixed point) and opusfile on 18/18 files on both CPUs. It is faster than libvorbis on 15/18 (x86-64) and 12/18 (M5) files, than Tremor on 15/18 and 16/18, and than stb_vorbis on 10/18 and 0/18.
  • stb_vorbis is the one decoder here that keeps up with Falcon. On the laptop it is 4% faster on the performance core, and Falcon is 8% faster on both kinds of efficiency core. On Apple M5, stb_vorbis is 1.28× faster on the performance cores and about 1.2× faster on the efficiency cores.
  • IMA ADPCM decodes 5–6× faster than anything else, but its files are about 3× the size of the Opus and Falcon files here (4 bits per sample).

Many simultaneous voices: 256 stereo streams, each started at a different offset, decoded for 1 s of audio in 20 ms mixer ticks on one thread.

Decoder CPU for 256 voice-seconds, x86-64 / M5 Heap per stream after open Voice start (open + first 20 ms + close), x86-64 / M5
Falcon (C API) 291 ms / 182 ms 318 KB (x86-64), 344 KB (M5), plus a copy of the compressed file 156 µs / 71 µs
libopus, float 365 ms / 250 ms 27 KB 46 µs / 37 µs
libopus, fixed point 422 ms / 298 ms 27 KB 51 µs / 40 µs
opusfile, float 379 ms / 259 ms 104 KB 96 µs / 66 µs
stb_vorbis 338 ms / 154 ms 211 KB 369 µs / 92 µs
libvorbis 377 ms / 210 ms 98 KB (214 KB peak) 437 µs / 178 µs
Tremor 357 ms / 235 ms 115–151 KB (194 KB peak) 402 µs / 191 µs
IMA ADPCM 47 ms / 29 ms a few bytes per channel 4 µs / 13 µs
  • Falcon's per-stream memory is its main cost for many-voice use.
    • Each channel has its own transform plans and windows.
    • The C API also keeps its own copy of the whole compressed file, so 256 voices of the same 2 MB file hold 256 copies.
    • About 140 KB of tables are built once per process on the first decode.
    • The decoder makes about seven heap allocations (about 8 KB in total) per 20 ms frame.
  • libopus allocates nothing after opus_decoder_create. The caller supplies the PCM buffer.

Code size (machine code / read-only data, linked into a minimal program, arm64, unused code stripped):

Decoder Code Read-only data
Falcon (decoder only) 629 KB, of which rustfft 310 KB, Rust std 229 KB, rustdct 31 KB, Falcon itself 44 KB 67 KB
libopus, float (with libogg) 130 KB 43 KB
opusfile + libopus + libogg 148 KB 44 KB
stb_vorbis 43 KB 1 KB
libvorbis + vorbisfile + libogg 117 KB 67 KB
Tremor + libogg 57 KB 53 KB
IMA ADPCM 3 KB 0.3 KB

On x86-64 the Falcon DLL, which includes the encoder, has 1.4 MB of code. Most of it is rustfft's scalar, SSE and AVX kernels.

How this was measured:

  • Hardware: an Intel Core Ultra 9 185H laptop running Windows 11, with the thread pinned to one performance, efficiency or low-power efficiency core; and a MacBook Pro with Apple M5 running macOS 26, on its performance cores.
  • Falcon build: the workspace release profile (fat LTO, one codegen unit). On x86-64 it targets x86-64-v3; on the M5 it uses the default aarch64 target.
    • On Windows, Falcon is the MSVC-toolchain DLL.
    • A MinGW (x86_64-pc-windows-gnu) build of Falcon decodes 1.45× slower (771× in total). It links msvcrt's pow/exp/log, which the decoder calls for every band.
  • C builds: -O3. On x86-64, GCC 15.1 (MinGW-w64) with -march=x86-64-v3; on the M5, Apple clang 17 with the default target. libopus uses its default SIMD: SSE4.1/AVX2 runtime dispatch on x86-64 and NEON intrinsics on arm64. Its stack protector is off.
  • Timing: each run times open, whole-file decode and close. Every file gets at least 10 runs and at least 1 s of runs. The figure is the best run over two separate passes.
    • Raw libopus is fed packets that were extracted from the Ogg file beforehand, so its figures exclude Ogg parsing. opusfile, vorbisfile and stb_vorbis include their container parsing.
    • Each decoder writes its native output format: f32 for Falcon, float libopus, opusfile, stb_vorbis and libvorbis; int16 for fixed-point libopus, Tremor and ADPCM.
  • Output check: every decoder's output was checked against the original for time alignment and SNR.
  • Earlier numbers: earlier versions of this README compared falcon bench-decode with ffmpeg -threads 1 -stream_loop 19 -i <file> -f null -.
    • That method measured FFmpeg's built-in Opus and Vorbis decoders, not libopus and libvorbis as previously stated.
    • Re-running it on the same laptop gives 892× for Falcon, 514× for FFmpeg's Opus decoder and 1006× for its Vorbis decoder, in total. Falcon is faster than FFmpeg's Vorbis decoder on only 5 of 18 files.
    • On Apple M5 the totals are 1539×, 730× and 1737× (7 of 18 files).
    • The earlier Vorbis figures could not be reproduced and have been withdrawn.

Evaluation methodology

Corpus: 18 files. All are 48 kHz stereo 16-bit.

  • 11 real recordings: game BGM, SFX (impacts, noise beds, applause), classical, electronic and pop music, female and male speech, and singing. They are not redistributed.
  • 7 synthetic torture signals: click trains ("castanets"), inharmonic decaying partials ("glockenspiel"), a fast detuned arpeggio ("harpsichord"), dense detuned chords ("polychord"), a 20 Hz–20 kHz sweep, gated pink-noise bursts, and steady 8–15 kHz tones over a pink floor ("hftonal"). tools/quality/make_torture.py regenerates these bit-exactly, using fixed seeds and numpy only.

Size budget. The Opus reference files were made with libopus (VBR) and come out at 96–214 kbps depending on content. Each Falcon file is encoded with --target-bytes equal to its Opus file's size, so it is at most that large: 0.998–1.000× on 17 files and 0.53× on the sweep. Vorbis and IMA ADPCM appear only in the speed comparison.

Metrics. All metrics compare against the original. The two spectral distances are standard ones from the literature, not project-invented:

  • MR-STFT: multi-resolution STFT log-magnitude distance (FFT sizes 512/1024/2048, 75% overlap), as used in the neural-codec literature (SoundStream, EnCodec).

  • Log-mel L1: L1 distance between log-mel spectrograms (64 mel bands, up to 20 kHz).

  • Perceptual A/B check: an objective multi-detector audio-quality checker is run as a paired comparison. The reference is the original, A is the Opus decode, and B is the Falcon decode, over all 18 pairs.

    • Each detector's per-file differences go into a sign test (verdict threshold p < 0.1).
    • Only detectors whose accuracy has been validated decide the overall verdict.
    • The detectors estimate overall quality (MOS-style), SNR against the reference, noisiness, harmonic amplitude modulation against the reference, hum and effective bit depth. Bandwidth-cliff, clipping and dropout detectors are reported but do not decide the verdict.

    The checker is not part of this repository. The spectral metrics can be reproduced with tools/quality/.

Honesty checks. For these results, the cross-correlation lag between original and decode is exactly 0 samples on every file, and the RMS level offset is within 0.04 dB. tools/quality/gates.py implements these checks, together with NaN/Inf and clipping bounds, for your own runs. A codec cannot flatter its scores with a hidden time shift or gain change.

Reproducing the spectral metrics:

# needs ffmpeg on PATH and Python with numpy + soundfile (pesq / pystoi optional)
python tools/quality/make_torture.py corpus/wav_original   # synthetic part of the corpus
# add your own 48 kHz WAVs to corpus/wav_original and encode matching Opus files into corpus/opus/
python tools/quality/compare.py corpus/                    # iso-size encode, decode, metrics, report

compare.py encodes each reference at the byte size of its .opus counterpart, decodes all codecs with ffmpeg / falcon, and writes a markdown + CSV report to tools/quality/reports/. metrics.py also computes per-band energy error, HF retention, and PESQ/STOI for speech, when the optional packages are installed.

Limitations and known issues

  • Not "better than Opus everywhere".
    • On the perceptual check Falcon is at parity overall. Opus is rated better on harpsichord, pinkbursts, polychord and sfxtransient.
    • Harmonic amplitude modulation (a slow flutter of partial amplitudes against the original) is still higher than Opus's on harpsichord and polychord.
    • Log-mel distance is higher than Opus's on 5 of 18 files.
  • Rate floor. At the finest quality index Falcon cannot always spend the budget. The synthetic sweep reaches index 0 at 0.53× the Opus size, and the encoder never pads.
  • 48 kHz design. The band tables and constants assume 48 kHz. Other rates are stored and played back, but they are not tuned.
  • No packet-loss concealment, no low-delay mode. Frames are 20 ms with a 20 ms look-back (TDAC), plus encoder look-ahead for transient scheduling. Falcon targets files and streams, not interactive calls.
  • Seeking is linear. The CLI and C API encoders emit a random-access frame only at the start (see ra_interval), and there is no seek index.
  • Floating-point contract. The pulse-count law uses f32 powf/exp on both sides, so an encoder and a decoder must compute identical values. This is verified on the development platform (x86-64, Windows). Agreement between different platforms' math libraries is expected but has not been tested systematically. The prefix-code tables are built with integer arithmetic only.
  • Mono/stereo focus. Up to 64 channels are supported (M/S applies to stereo only). Dense many-channel content may use the coarser Ext frames.
  • Encoder speed. Rate control runs several full encodes in a binary search. Encoding is much slower than decoding, by design.

Upgrading from 1.x

  • Files: 1.x .falcon files (version 0x04) cannot be decoded by 2.x. Re-encode from the source audio.
  • Rust API: EncoderConfig::quality and the Quality enum are gone (they had no effect). PsychoacousticModel is removed. falcon_core::coeffcode is now falcon_core::rate. Mode::Lpc and Mode::Hybrid are now Mode::Flagged and Mode::Ext. See the migration table.
  • CLI: -q/--quality is removed. Use --quality-index, --target-bytes or --target-kbps.
  • C / C++: no API change (only the version string).

Patents and licensing

Falcon's code is MIT licensed, and the project charges no royalties. Version 2.0 was redesigned specifically to avoid four patents that could read on 1.x:

  • US 8,838,442 (spreading rotation)
  • US 9,009,036 (recursive gain-shape splitting)
  • US 9,015,042 (decoder-side collapse noise)
  • US 11,234,023 (rANS)

A preliminary search of the new 2.0 elements found no plausible reading of active claims on the shape coder, energy coding, class-controlled fill or pulse search. The encoder's transient decision was also reworked into a per-frame analysis of each frame's own window span, with no carry-over between frames, to avoid the cross-frame "hangover" indicator claimed by the Ericsson family US 9,495,971 et al. One open item is documented: the mid/side decoder structure, against a broad reading of US 11,527,252.

PATENTS.md has the full analysis, sources and limits. It is not legal advice and gives no warranty. Falcon contains no code from other codec implementations (NOTICE), and all dependencies are permissively licensed (THIRD_PARTY_LICENSES.md).

Opus and Vorbis are trademarks of their respective owners. They are named here only for factual comparison as open reference codecs. Falcon is an independent project and is not affiliated with or endorsed by them.

References

  • J. P. Princen, A. B. Bradley, "Analysis/synthesis filter bank design based on time domain aliasing cancellation," IEEE Trans. ASSP, 1986.
  • B. Edler, "Codierung von Audiosignalen mit überlappender Transformation und adaptiven Fensterfunktionen," Frequenz 43, 1989.
  • T. R. Fischer, "A pyramid vector quantizer," IEEE Trans. Information Theory, 1986.
  • D. A. Huffman, "A method for the construction of minimum-redundancy codes," Proc. IRE, 1952.
  • S. W. Golomb, "Run-length encodings," IEEE Trans. Information Theory, 1966; R. F. Rice, "Some practical universal noiseless coding techniques," JPL Publication 79-22, 1979.
  • J. Herre, D. Schulz, "Extending the MPEG-4 AAC codec by perceptual noise substitution," AES 104th Convention, 1998.
  • ITU-T Recommendation G.722.1, "Low-complexity coding at 24 and 32 kbit/s for hands-free operation in systems with low frame loss," 1999.
  • J. D. Johnston, A. J. Ferreira, "Sum-difference stereo transform coding," ICASSP 1992.
  • E. Zwicker, "Subdivision of the audible frequency range into critical bands," JASA, 1961.
  • Xiph.Org Foundation, Vorbis I specification (window function).
  • N. Zeghidour et al., "SoundStream: An end-to-end neural audio codec," 2021; A. Défossez et al., "High fidelity neural audio compression" (EnCodec), 2022 (MR-STFT methodology).

Contributing

Issues and pull requests are welcome. Please read CONTRIBUTING.md first. In particular, it asks for patent hygiene and clean-room code: no code from other codec implementations.

License

MIT, see LICENSE. Copyright (c) 2026 Falcon Audio Codec contributors.

About

A royalty-free, patent-free audio codec that beats Opus on quality at equal file size, with a decoder lighter than Vorbis. Rust libraries + C/C++ library + CLI.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages