Read this in 日本語 (Japanese).
Falcon is an MDCT audio codec written in Rust and released under the MIT license. It follows one design rule: keep the decoder as light as possible, and buy quality back with encoder-side work.
On an 18-file test corpus, with every Falcon file no larger than the Opus file it is compared against:
- Quality is on par with Opus, or better, depending on the metric. Falcon has the lower MR-STFT distance on 18 of 18 files and the lower log-mel distance on 13 of 18. A multi-detector perceptual A/B check finds no significant difference overall.
- Decoding is faster than libopus, libvorbis and Tremor, and on par with stb_vorbis on x86-64. In a library-level benchmark on one thread, Falcon needs 0.66× the decode time of libopus and 0.79× that of libvorbis on an x86-64 laptop (0.68× and 0.84× on Apple M5). stb_vorbis is about as fast on x86-64 and 1.28× faster on Apple M5. Falcon uses more memory per stream than any of them (see Decode speed).
Version 2 redesigns the codec to avoid specific third-party patents (see Patents) and contains no code taken from other codec implementations.
Documentation: API reference · Changelog · Patent position · Third-party licenses · C/C++ bindings · Contributing
- Highlights
- What's new in 2.0
- Design philosophy
- Repository layout
- How Falcon encodes: the signal chain
- How Falcon decodes
- Bitstream format
- Building
- Command-line interface
- Rate control
- Rust API
- C and C++ API
- Quality at equal size
- Decode speed
- Evaluation methodology
- Limitations and known issues
- Upgrading from 1.x
- Patents and licensing
- References
- Contributing
| Property | Falcon 2.0 |
|---|---|
| Quality vs Opus, equal file size | Lower MR-STFT distance on 18/18 files and lower log-mel L1 on 13/18. The perceptual A/B check gives "same" overall: no detector is significantly worse, and harmonic amplitude modulation is better on 7 of 9 pitched files and worse on 2 (p = 0.18). |
| Decode speed (one thread, library level) | 1114× real time over the corpus on an x86-64 laptop core and 1551× on Apple M5. About 1.5× faster than libopus and 1.2–1.35× faster than libvorbis and Tremor; stb_vorbis is on par on x86-64 and 1.28× faster on Apple M5. |
| Decoder memory | About 320–345 KB of heap per stereo stream, plus a copy of the compressed file when decoding through the C API: more than libopus (27 KB) or stb_vorbis (about 210 KB). |
| Transform | MDCT, 960 coefficients per 20 ms frame at 48 kHz, Vorbis power-complementary window. Transient frames switch to four short blocks (Edler window switching, signalled at zero bit cost). |
| Band energies | 27 bands, 1 dB grid. Time DPCM on inter frames and frequency DPCM after a random-access point. Static canonical Huffman codes (8 tables) behind a per-channel all-zero bit. |
| Coefficients | Gain-shape coding. The pulse count K is derived from the transmitted energies, so it costs no side bits. A static cell table per band, the per-cell pulse counts as one stars-and-bars composition index, and each cell as a Fischer pyramid enumeration index or dense Golomb-Rice magnitudes. |
| Noise and holes | Encoder-signalled per-band coding classes (delta-coded): pulses only, pulses plus zero-bin fill at one of two fixed levels, or noise substitution. The decoder takes no fill decisions of its own. |
| Stereo | Per-frame mid/side coding in the MDCT domain. The side channel's step takes half of its reference from the transmitted mid energies. |
| Special frames | Bit-exact LSB-floor payload for 16-bit noise-floor frames. An overflow-extension mode (Ext) keeps every frame within the 12-bit frame-size field for up to 64 channels. |
| Integration | Rust crates, a CLI, and a C ABI with a header-only C++17 wrapper (fuzz-tested against garbage and truncated input; MSVC and MinGW). |
| License | MIT. No third-party codec code; every dependency is permissively licensed. |
Every quality number in this README is measured at equal or smaller file size than the Opus reference. See Evaluation methodology for exactly what was measured and how, and Limitations for where Falcon is behind.
2.0 replaces the coefficient, energy, fill and entropy layers of 1.x. It has two goals:
- Design around specific patents that a review found could read on 1.x.
- Remove the parts of 1.x whose code followed libopus.
| 1.x | 2.0 | Why |
|---|---|---|
| Two-step "spreading" rotation of the decoded pulses (CELT/Opus style) | Removed. Zero-bin fill is controlled by an encoder-signalled class | Designed around US 8,838,442. The rotation code followed libopus. |
| Recursive band halving with a quantized angle (theta) per node | Static cell tables + one composition index + per-cell enumeration or Golomb-Rice | Designed around US 9,009,036. No split recursion, no angles. |
| Decoder-side noise injection into bands that received no pulses ("anti-collapse") | Encoder-signalled noise substitution (class 3) | Designed around US 9,015,042. The decoder does not inspect energies to decide on noise. |
| rANS entropy coding | Static canonical Huffman + Golomb-Rice | Designed around US 11,234,023. |
| Mixed time/frequency energy predictor | Pure time DPCM + frequency DPCM on random-access frames | Removes a frame-rate flutter of harmonic amplitudes, and simplifies the energy path. |
| Pulse search using the libopus metric | New search: rounding + remainder repair + cosine-improving pulse moves | Clean-room rewrite (encoder only). |
| Transient detector: an attack in frame f also marked frame f+1 for short blocks | Each frame decides its block type from its own window span only; nothing carries over between frames | Designed around US 9,495,971 (transient hangover). Encoder only. |
| Raw scalar overflow fallback | Ext frames: the normal layout at a coarser rate level |
The old fallback did not fit the frame-size field for many channels. |
Compatibility: the bitstream is now version 0x10. 2.x decoders reject
1.x (0x04) files and 1.x decoders reject 2.x files. See
Upgrading from 1.x.
Honesty note on 1.x: 1.x advertised that it beat Opus on all 18 files, based on MR-STFT alone. That claim did not hold up under the perceptual A/B check that 2.0 adds to the methodology. 2.0 reports all three views: MR-STFT, log-mel, and the perceptual check.
"Make the decoder extremely lightweight; achieve quality through encoder optimization."
- The decoder executes; it does not decide. Per frame it reads static prefix codes, decodes integer pulse vectors, applies one scaling pass per band (plus the fill that the encoder chose), and runs one inverse MDCT (or four short ones). It has no psychoacoustic model, no search, no adaptive tables and no heuristics that inspect the signal. This is what makes it cheap, and it also keeps it simple to audit.
- The encoder is allowed to work hard. Silence and transient analysis, per-frame stereo decisions, the per-band coding-class decision by analysis-by-synthesis, the pulse search and a global rate-control search all run at encode time.
- Zero side information where possible. The pulse count of every band is a deterministic function of data the decoder already has: the transmitted band energies and the global quality index. Only what cannot be derived is sent: the energies, the coding classes, and the pulse vectors.
| Path | What it is |
|---|---|
falcon_core/ |
Shared kernel and the single source of truth for everything both sides must agree on: bitstream I/O, band tables, energy prediction, prefix codes (prefix), the step and pulse-count law (alloc, pvq), fixed-cell shape coding (cells), coding classes and fills (fill), window switching (transient), the LSB-floor payload (lsbfloor), and the rate law (rate) |
falcon_encoder/ |
Encoder: MDCT analysis, silence detection (mode_detect), class decision (classify), transient look-ahead, M/S decision, rate control (encode) |
falcon_decoder/ |
Decoder: frame parsing, reconstruction, fast inverse MDCT |
falcon_cli/ |
The falcon command-line tool (encode, decode, info, bench-decode) and the integration tests |
falcon_capi/ |
C ABI + header-only C++17 wrapper, CMake package, examples (README) |
docs/ |
API reference and README images |
tools/quality/ |
Evaluation harness: metrics, honesty gates, synthetic torture-corpus generator, iso-size comparison driver (Python) |
tools/eval/ |
Older Rust A/B harness (falcon-eval) |
The encoder processes 20 ms frames (960 samples per channel at 48 kHz).
encode_file first scans the whole input: it detects transients (with
look-ahead), makes the initial M/S decision, and profiles LSB-floor frames.
It then runs the frame loop, inside a rate-control search when a size or
bitrate target is set.
A frame is coded Silent only when both it and the previous frame fall below −100 dBFS of AC (DC-removed) energy. The MDCT window spans two frames, so the previous frame's tail must also be silent. Faint content (−95…−110 dBFS decay tails, room tone) is still coded, and a pure DC pedestal does not count as activity.
16-bit sources often end in a quantization-noise floor: a DC pedestal of about 1 LSB with single-LSB flicker. An MDCT cannot reproduce this at any sensible rate. Falcon codes such frames exactly in the time domain instead. A frame qualifies when every sample lies exactly on the 16-bit grid and within ±5 LSB of a DC level. Its content (per channel, a DC level plus a residual whose first difference is coded as Golomb-Rice zero runs) is carried in the next Silent frame's payload and added to the output after the zero-coefficient IMDCT. Reconstruction is bit-exact. Material that is not on the 16-bit grid never takes this path.
Each frame decides its own block type from an analysis of its own MDCT
window span (the 1920 input samples its window covers, plus 40 ms of
history that is re-read from the signal for every frame). The analysis
compares 2.5 ms sub-block energies with a decaying envelope: +12 dB marks an
attack, and an absolute gate keeps LSB flicker from triggering. No detector
state or decision carries over from one frame's analysis to the next. A frame
whose span holds an attack is coded as short blocks: four 480-sample
MDCTs (240 bins each,
interleaved into the usual 960-bin layout), which confines coding noise to
about 5–10 ms around the attack. Start and Stop transition windows follow
Edler's window switching and are TDAC-exact. As in AAC-style window
switching, the frame before a short-block frame takes the Start window. The
flag travels in the frame-header mode field (Mode::Flagged), so it costs no
bits. The decoder
derives each frame's window from the flags of the previous and current
frame. Short-block runs are kept away from random-access frames.
For stereo input the encoder decides M/S or L/R per frame. It uses M/S when the side energy is small against the mid (with hysteresis: enter below 0.20, stay below 0.35). It rejects M/S when the per-band side-to-mid ratio is much higher than the broadband ratio, because then a loud centred bass would hide a wide mix. M/S is applied to the MDCT coefficients. The MDCT is linear and the overlap-add always runs in L/R, so switching per frame is exactly reconstructible.
Each channel is transformed with an MDCT under the Vorbis power-complementary
window (960 coefficients; 1920-sample span). The 960 coefficients are grouped
into 27 bands, Bark-like up to 12 kHz, with the top octave split into
12–15 / 15–18.5 / 18.5–20 / 20–24 kHz so the envelope can follow the real
high-frequency tilt. A band's energy is 20·log10(3.5 · rms) in dB.
Energies are quantized on a 1 dB grid with DPCM:
- Inter frames: pure time prediction,
pred_b = previous_b. A steady band repeats its level with a zero residual. (1.x mixed in the lower band's energy. On steady sloped spectra that made the 1 dB rounding alternate, which modulated every partial in the band at the frame rate.) - Random-access frames (the first coded frame after a reset): frequency
prediction,
pred_b = recon_{b−1},pred_0 = −20 dB. These frames depend on no earlier frame. - The residual range is ±127 dB, so a band can jump from the −80 dB floor to full scale in one frame (clicks out of silence) and fall back in the next.
- Silent frames decay the predictor towards silence identically on both sides.
The residuals are coded with static canonical Huffman codes. There are 8 tables built from discrete-Laplacian models, from very sharp to wide. The code length limit is 12 bits (JPEG Annex K length limiting), and residuals of 16 dB or more are sent as an ESCAPE code plus 8 raw bits. The encoder picks the cheapest table per frame (3 bits in the frame flags). Each channel first sends an all-zero bit, so a fully steady channel costs 1 bit. The tables are built with integer arithmetic at start-up, identically on every platform.
Each band's quantizer resolution is a pure function of the transmitted
energies and the global rate scalar K. It is computed identically by
encoder and decoder, so no allocation bits are sent. The normalized step of
band b is
d_b = K · shape(b) / 10^((ρ·e_max,ref + γ·(e_ref − e_max,ref) + (e_b − e_ref)) / 20)
e_bis the band energy ande_maxthe frame's peak band energy (dB).- The exponents
ρ = 0.5(across frames) andγ = 0.32(within a frame) are sub-linear. They buy relative, log-spectral precision instead of absolute precision. shape(b)is a fixed perceptual weight: A-weighting-like, held high from 3.7 to 20 kHz, and 6× finer for the 100–920 Hz fundamentals and low harmonics, where coarse pulses make harmonic amplitudes flutter.- Mid-referenced side step: for the side channel of an M/S frame, the
loudness reference
e_refis taken half-way to the mid channel's band energy. The side's coding noise is heard against the mid, so it should not be coded far finer than the mid. For all other channelse_ref = e_b. - A band more than 96 dB below the frame peak, or whose step reaches 0.7 (4.0 for loud bands above 15 kHz), gets no pulses. What happens to it then is decided by its coding class (step 6).
- Transient step: a frame whose peak energy jumps by 9 dB or more over the previous frame gets a finer step (× 0.25), and the following frame × 0.8. Both sides derive this from the transmitted energies.
The step is mapped to a pulse count K_b by a fixed law (the expected L1
norm of a dead-zone quantizer on a unit Laplacian source), with at least one
pulse per coded band and at most 64 pulses per bin.
Every band of every channel has a coding class:
| Class | Decoder behaviour |
|---|---|
| 0 — pulses | Pulses only; a band with no pulses is silent |
| 1 — low fill | Pulses, plus every zero bin set to a random-sign value at 0.2× the band's decoded per-bin RMS; the band is then rescaled to its transmitted energy |
| 2 — high fill | Same with 0.45× |
| 3 — noise | No shape bits. The band is noise at its transmitted energy, tilted along the neighbouring bands' energies |
The classes are a decision of the encoder, and the decoder only executes them. The encoder chooses them per band and frame:
- A band on the energy floor always stays class 0.
- A band with no pulses gets class 3, except in the 20–24 kHz band and around onsets. There a 20 ms noise burst would sound like pre- or post-echo, so the band stays silent.
- A steady, not-too-dense band on a long-window frame is decided by analysis by synthesis. The encoder reconstructs the band exactly as the decoder will for classes 0, 1 and 2, and picks the class whose smoothed log-power spectrum is closest to the input. The input spectrum comes from a shift-invariant odd-frequency DFT, not the MDCT.
- Classes carry a mild bias towards the current class.
Classes persist per channel. A frame sends a class section (announced by bit 0 of the frame flags) only when something changes. The section is coded as a delta: per channel a changed bit, the number of changes, the runs of unchanged bands (Golomb-Rice), and each new class as one of the three other values. All classes reset to 0 at random-access frames. Short-block frames treat classes 1 and 2 as 0.
A band with K_b pulses is coded as an integer vector y of its N bins
with Σ|y_i| = K_b. The decoder scales the whole vector to the transmitted
band energy, so band energy is exact by construction.
- Static cells. Each band is split into cells of at most 16 bins by a
fixed table:
ceil(N/16)near-equal cells, with a separate table for short-block frames where each cell stays inside one time block. The table depends only on the band, never onKor on bit counts. - Composition. The pulse counts of the cells,
(k_0 … k_{C−1})withΣ k_c = K_b, are sent as one stars-and-bars index: the rank of the bar positions amongK_b + C − 1slots, in truncated binary overC(K_b + C − 1, C − 1)values. Above 2 pulses per bin they are sent cell by cell instead. - Cells. A cell with up to
K_ENUM_MAX[n]pulses (a static table, ≤ 64) is sent as its rank among allV(n, k)signed integer vectors of dimensionnand L1 normk(Fischer's pyramid enumeration, truncated binary, u64 arithmetic). A denser cell is sent as Golomb-Rice magnitudes plus sign bits, with the last bin implied byk_c.
On short-block frames the band's bins are first regrouped by time block, so cells follow time. Silent blocks then decode to exact zeros and the attack block takes the budget. The encoder also deletes weak pre-onset residue that could only smear backwards as pre-echo.
The shape that best matches the band's direction maximizes the cosine
(Σ|x_i|·y_i)² / Σy_i². The encoder searches it in three steps:
- Round
K|x_i|/Σ|x|. - Repair the pulse count using the rounding remainders.
- Move single pulses from the smallest to the largest remainder at the
current optimal gain, as long as each move strictly raises the cosine.
When that stalls, try the exact best remove-then-add move. The loop is
capped at
Nmoves.
A single quality index (0 = finest … 255 = coarsest, plus a 4-bit fraction
in the header) sets K on a log scale, and file size is monotone in it.
With --target-bytes or --target-kbps, the encoder binary-searches the
index, then the fraction, for the finest setting whose whole file fits.
The fraction lets it use about 99–100% of a byte budget.
The frame-size field is 12 bits (at most 4094 payload bytes per frame).
Dense many-channel material can exceed that. Such a frame is re-packed from
the same analysis at overflow-extension levels 1–6, which raise K
geometrically towards a fixed ceiling, until it fits. It is then marked
Mode::Ext, with the level in frame-flags bits 4–6. Level 7 carries only the
energies, and every band decodes as noise. For example, 64 channels of white
noise at the finest quality fit at level 5 or 6. Normal material never
reaches this path.
Per frame:
- Header, CRC. The decoder reads the 2-byte frame header and verifies the CRC-8 (enabled by default).
- Energies. For every channel: the all-zero bit, then 27 Huffman codes (table from the frame flags). It adds the prediction (time DPCM, or frequency DPCM after a random-access point).
- Classes. If bit 0 of the flags is set, it applies the class delta.
- Pulse counts. It derives every band's
K_bfrom the energies (steps 5 and 10 above). - Shapes. For every band with pulses and class ≠ 3: it reads the composition, then each cell, writes the integer vector, and scales it to the transmitted energy in one pass. Zero bins get the class 1/2 fill, and class 3 bands get tilted noise from a deterministic per-band PRNG.
- Stereo. For M/S frames it applies the inverse M/S on the coefficients.
- Synthesis. One fast inverse MDCT (N/4-point complex FFT) with the frame's window, or four short ones, then overlap-add.
- Silent frames. A zero-coefficient IMDCT keeps the overlap consistent, plus the LSB-floor payload when present.
The decoder keeps only a little state per channel: the IMDCT overlap, the energy predictor, the last peak energy and frames since an onset, the current classes, and the previous frame's transient flag.
All multi-byte fields are little-endian. Bit fields inside frame payloads are packed LSB-first.
| Field | Size | Description |
|---|---|---|
| Magic | 4 B | "FALC" |
| Version | 1 B | 0x10 (Falcon 2.x). Decoders reject any other value. |
| Flags | 2 B | Bits 0–7: quality index (0–255); bits 8–11: fractional quality index (sixteenths of one step); bits 12–15 reserved (0) |
| Sample rate | 3 B | Hz (24-bit) |
| Channels | 1 B | 1–64 |
| Total frames | 4 B | Number of frames in the stream |
| Bitrate | 2 B | Nominal kbps (informational) |
| LFE mask | 8 B | One bit per channel (informational) |
[frame header: 2 B][payload][CRC-8 over the payload: 1 B]
| Frame-header field | Bits | Values |
|---|---|---|
| Frame size | 12 | Payload bytes + 1 (CRC); at most 4095 |
| Mode | 2 | 0 = Mdct, 1 = Flagged (transient: short blocks), 2 = Ext (overflow extension), 3 = Silent |
| Flags | 2 | Bit 0: random-access frame (predictors, classes and overlap reset); bit 1: M/S coupling (stereo) |
The window of a frame follows from the transient flags of the previous and current frame: Normal, Start, Shorts, or Stop.
One LSB-first bit stream:
| Section | Content |
|---|---|
| Flags (8 bits) | Bit 0: class section present; bits 1–3: energy Huffman table (0–7); bits 4–6: Ext rate level (1–7; must be 0 in modes 0/1); bit 7: reserved (0) |
| Energies | Per channel: all_zero (1 bit); if 0, 27 canonical Huffman codes of the zigzagged residuals, where ESCAPE + 8 raw bits covers residuals of 16 dB or more |
| Classes | Only if flags bit 0: per channel a changed bit; if set, number of changes − 1 (Rice k=1), then per change the run of unchanged bands (Rice k=2) and the new class (truncated binary over the 3 other classes) |
| Shapes | For every channel, then every band with K_b > 0 and class ≠ 3 (none at Ext level 7): the composition (one truncated-binary index, or per cell above 2 pulses/bin), then each cell (enumeration index, or dense Golomb-Rice magnitudes + signs) |
Nothing else is sent. The pulse counts, cell layout, field widths and fill levels are all derived from data both sides already have.
[0]: plain silence.[1], then per channel[dc: i8][k: u8], then one bit stream of Golomb-Rice zero runs of the residual's first difference: this is the LSB-floor payload. It is added to the output of the frame being emitted, which is the previous frame.k = 0xFFmeans DC only.
Requirements: Rust stable with Cargo. The declared MSRV is 1.70, and development uses current stable. Windows, macOS and Linux are supported.
git clone https://github.com/pbtechlab/Falcon.git
cd Falcon
cargo build --release # CLI at target/release/falcon(.exe)
cargo test --workspace --releaseCPU target:
.cargo/config.tomlbuilds for x86-64-v3 (AVX2, FMA, BMI; Intel Haswell / AMD Excavator and later). That is what the published speed numbers use. For older x86-64 CPUs or other targets, override it, e.g.RUSTFLAGS="-C target-cpu=x86-64" cargo build --release, or remove the file. The bitstream format is unaffected.
C / C++ library (static + shared):
cargo build --release -p falcon_capi
cmake -S falcon_capi -B build && cmake --build build --config Release # examplesSee falcon_capi/README.md for the link lines of each toolchain (MSVC, MinGW, gcc/clang).
falcon encode <input.wav> <output.falcon> [options]
falcon decode <input.falcon> <output.wav>
falcon info <input.falcon>
falcon bench-decode <input.falcon> [-i N]
Reads integer-PCM WAV (16- or 24-bit, any channel count 1–64; the codec is
tuned for 48 kHz) and writes a .falcon file.
| Option | Description |
|---|---|
--target-bytes <N> |
Finest quality whose whole file is ≤ N bytes (binary search). Overrides the kbps options. |
--target-kbps <R> |
Target average bitrate, converted to a byte budget from the duration |
-b, --bitrate <KBPS> |
Default target when neither of the above is given (default 128). It is also written to the header. |
--quality-index <0-255> |
Fixed quality (0 = finest). Disables rate control. |
falcon encode in.wav out.falcon # ~128 kbps target
falcon encode in.wav out.falcon --target-kbps 96
falcon encode in.wav out.falcon --target-bytes 500000 # exact size budget
falcon encode in.wav out.falcon --quality-index 20 # fixed qualityIt prints the resolved target, frames, bytes, average bitrate, encode time and compression ratio. Stereo input uses per-frame M/S automatically.
Writes a 16-bit PCM WAV at the stream's sample rate and channel count, and
reports the decode time and real-time factor. The first frame (TDAC warm-up)
is dropped, so output sample n lines up with input sample n.
Prints the header: sample rate, channels, layout, total frames, nominal bitrate, duration, file size, and the LFE mask if set.
Loads the file into memory and decodes it N times (default 10). It prints
the first-iteration time, the average decode time, the real-time factor
and the throughput. It is meant for quick checks; the cross-codec numbers in
Decode speed come from a C harness that calls the C API.
Priority order:
--quality-indexsets a fixed quality with no search.--target-bytessets an exact budget.--target-kbpssets a budget from the duration.--bitrateis the default budget.
Size is monotone in the quality index, so a binary search over the integer index and then over the 4-bit fraction finds the finest setting that fits. If even the finest setting (index 0) is smaller than the budget, the file simply comes out smaller. This happens on the synthetic sweep, which is 0.53× the Opus size at index 0.
The workspace crates are falcon_encoder, falcon_decoder and
falcon_core. Depend on them by path or git. PCM is f32, nominally in
[-1.0, 1.0]; a frame is 960 samples per channel. Full reference:
docs/API_REFERENCE.md.
Encode a whole buffer (de-interleaved: one Vec<f32> per channel):
use falcon_encoder::{encode::encode_file, EncoderConfig};
use std::{fs::File, io::BufWriter};
let channels: Vec<Vec<f32>> = load_my_audio(); // e.g. [left, right]
let config = EncoderConfig {
target_bytes: Some(250_000), // finest quality that fits 250 kB
..EncoderConfig::default() // or set quality_index for a fixed quality
};
let out = BufWriter::new(File::create("out.falcon")?);
let stats = encode_file(out, &channels, 48_000, config)?;
println!("{} frames, {} bytes", stats.total_frames, stats.bytes_written);EncoderConfig fields:
| Field | Default | Meaning |
|---|---|---|
target_bytes |
None |
Size target for the whole file; enables rate control |
quality_index |
43 |
Fixed quality (0 = finest … 255 = coarsest), used when target_bytes is None |
quality_frac |
0 |
Fraction of one index step (0–15) |
stereo_coupling |
true |
Allow per-frame M/S on stereo input |
ra_interval |
50 |
Random-access frame every N frames (the CLI and C API use 10000, which in practice means only the first frame) |
bitrate |
128 |
Nominal kbps written to the header (informational) |
Decode a whole stream:
use falcon_decoder::decode::decode_file;
let (pcm, sample_rate, channels) = decode_file(std::io::BufReader::new(File::open("out.falcon")?))?;
// pcm[c][n]: channel c, sample n (de-interleaved); first frame (warm-up) already droppedStream frame by frame:
use falcon_decoder::Decoder;
let mut dec = Decoder::new(reader)?; // parses and validates the header
while dec.decode_frame()? != 0 { // 960 samples per channel, 0 at end
for ch in 0..dec.channels() {
let frame: &[f32] = dec.get_samples(ch);
}
}falcon_encoder::Encoder offers the same frame-by-frame interface for
encoding: interleaved frames of channels × 960 samples. It does not do the
whole-file look-ahead (transient scheduling, rate control) that
encode_file does. All fallible calls return falcon_core::FalconError.
Corrupt or truncated streams are designed to produce errors rather than
panics; this is fuzz-tested through the C API.
falcon_capi builds a static and a shared library with a small, stable C99
ABI (falcon_capi/include/falcon.h) and a header-only, exception-free C++17
RAII wrapper (falcon.hpp).
#include <falcon.h>
/* streaming decode */
FalconDecoder* d = falcon_decoder_open(bytes, len); /* copies the input */
float buf[1024 * 2]; int64_t n;
while ((n = falcon_decoder_read_f32(d, buf, 1024)) > 0) { /* n frames, interleaved */ }
falcon_decoder_seek(d, 48000); /* sample-exact */
falcon_decoder_close(d);
/* one-shot encode (target_bytes = 0 -> use target_kbps) */
uint8_t* enc; size_t enc_len;
falcon_encode_buffer(pcm, frames, 2, 48000, 0, 128, &enc, &enc_len);
falcon_free(enc);#include <falcon.hpp>
falcon::Decoder dec = falcon::Decoder::open(bytes, len);
std::vector<float> block(1024 * dec.channels());
std::int64_t n = dec.read(block.data(), 1024);- All functions return error codes (
FALCON_OK= 0, negative = error). - Buffers returned by the library are freed with
falcon_free. - A decoder handle is not thread-safe, but separate handles are independent.
- Garbage and truncated input is fuzz-tested to return errors.
- Linking has been verified with MSVC and MinGW.
Details, CMake integration and per-toolchain link lines are in falcon_capi/README.md.
Every Falcon file below was encoded with --target-bytes set to the size of
the corresponding Opus file, so Falcon is never larger. Both metrics are
distances to the original: lower is better, and bold marks the better
codec. The last column is the per-file verdict of the perceptual A/B check
(see methodology).
| File | Type | Size vs Opus | MR-STFT Falcon / Opus | log-mel Falcon / Opus | Perceptual A/B (per file) |
|---|---|---|---|---|---|
| bgm1 | music (game BGM) | 1.000 | 0.454 / 0.638 | 0.164 / 0.148 | Falcon better |
| sfx1 | SFX | 0.999 | 0.444 / 0.626 | 0.188 / 0.127 | same |
| sfx2 | SFX | 0.998 | 0.526 / 0.547 | 0.187 / 0.189 | Falcon better |
| musicclassical | music, classical | 1.000 | 0.511 / 0.603 | 0.123 / 0.176 | Falcon better |
| musicelectronic | music, electronic | 0.999 | 0.222 / 0.251 | 0.055 / 0.077 | Falcon better |
| musicpop | music, pop | 0.999 | 0.677 / 0.883 | 0.171 / 0.198 | Falcon better |
| sfxnoise | SFX, noise | 0.998 | 0.449 / 0.706 | 0.110 / 0.118 | same |
| sfxtransient | SFX, applause | 1.000 | 0.586 / 0.876 | 0.142 / 0.099 | Opus better |
| voicefemale | speech, female | 0.999 | 0.417 / 0.599 | 0.100 / 0.197 | Falcon better |
| voicemale | speech, male | 1.000 | 0.406 / 0.553 | 0.109 / 0.181 | Falcon better |
| voicesing | singing | 0.999 | 0.480 / 0.483 | 0.090 / 0.106 | Falcon better |
| castanets | synthetic | 1.000 | 0.304 / 1.008 | 0.049 / 0.291 | mixed |
| glockenspiel | synthetic | 0.999 | 0.540 / 0.919 | 0.518 / 1.444 | Falcon better |
| harpsichord | synthetic | 0.998 | 0.521 / 0.678 | 0.175 / 0.128 | Opus better |
| hftonal | synthetic | 0.999 | 0.434 / 1.421 | 0.118 / 0.190 | same |
| pinkbursts | synthetic | 0.999 | 0.791 / 4.063 | 0.583 / 1.751 | Opus better |
| polychord | synthetic | 1.000 | 0.435 / 0.883 | 0.121 / 0.117 | Opus better |
| sweep | synthetic | 0.532 | 0.678 / 0.694 | 0.851 / 1.072 | same |
Summary: parity with Opus, and ahead on the spectral-distance metrics.
- MR-STFT: Falcon is lower on 18/18 files.
- Log-mel L1: Falcon is lower on 13/18. It is higher on bgm1, sfx1, sfxtransient, harpsichord and polychord.
- Perceptual A/B check, aggregate (sign test over the 18 pairs per detector): the overall verdict is "same". No detector with validated accuracy is significantly worse for Falcon.
| Detector (what it measures) | Falcon better / worse | p (sign test) |
|---|---|---|
| Overall-quality (MOS-style) estimate | 8 / 10 | 0.81 |
| Signal-to-noise ratio against the reference | 11 / 7 | 0.48 |
| Noisiness estimate | 11 / 7 | 0.48 |
| Harmonic amplitude modulation against the reference (9 pitched files) | 7 / 2 | 0.18 |
| Hum | 4 / 3 | 1.00 |
| Effective bit depth | 0 / 0 | 1.00 |
- Perceptual A/B check, per file: Falcon better on 9 files, Opus better on 4 (harpsichord, pinkbursts, polychord, sfxtransient), mixed on 1, and the same on 4.
Objective metrics are proxies. Listen to your own material.
These numbers come from one library-level benchmark. Every decoder is linked into the same small C program, decodes whole files from memory to PCM in 20 ms reads on one thread, and is timed without file I/O or process start-up. Falcon is called through its C API. The others are the libraries that games ship. The inputs are the files of the quality evaluation: Falcon and Opus at equal size, and Vorbis at its own quality setting.
Total over the 18-file corpus, in times real time (higher is better):
| Decoder | x86-64, P-core | x86-64, E-core | x86-64, LP E-core | Apple M5, P-core | Falcon's time ÷ this decoder's (x86-64 / M5) |
|---|---|---|---|---|---|
| Falcon 2.0 (C API) | 1114× | 661× | 434× | 1551× | 1.00 / 1.00 |
| libopus 1.6.1, float | 731× | 425× | 279× | 1056× | 0.66 / 0.68 |
| libopus 1.6.1, fixed point | 629× | 353× | 234× | 886× | 0.56 / 0.57 |
| opusfile 0.12 (Ogg + libopus float) | 717× | 414× | 271× | 1022× | 0.64 / 0.66 |
| stb_vorbis 1.22 | 1154× | 612× | 401× | 1978× | 1.04 / 1.28 |
| libvorbis 1.3.7 (vorbisfile) | 879× | 484× | 317× | 1297× | 0.79 / 0.84 |
| Tremor (fixed-point Vorbis) | 877× | 526× | 346× | 1146× | 0.79 / 0.74 |
| IMA ADPCM (4:1, plain C) | 5613× | 2401× | 1580× | 9727× | 5.0 / 6.3 |
- Per file: Falcon is faster than libopus (float and fixed point) and opusfile on 18/18 files on both CPUs. It is faster than libvorbis on 15/18 (x86-64) and 12/18 (M5) files, than Tremor on 15/18 and 16/18, and than stb_vorbis on 10/18 and 0/18.
- stb_vorbis is the one decoder here that keeps up with Falcon. On the laptop it is 4% faster on the performance core, and Falcon is 8% faster on both kinds of efficiency core. On Apple M5, stb_vorbis is 1.28× faster on the performance cores and about 1.2× faster on the efficiency cores.
- IMA ADPCM decodes 5–6× faster than anything else, but its files are about 3× the size of the Opus and Falcon files here (4 bits per sample).
Many simultaneous voices: 256 stereo streams, each started at a different offset, decoded for 1 s of audio in 20 ms mixer ticks on one thread.
| Decoder | CPU for 256 voice-seconds, x86-64 / M5 | Heap per stream after open | Voice start (open + first 20 ms + close), x86-64 / M5 |
|---|---|---|---|
| Falcon (C API) | 291 ms / 182 ms | 318 KB (x86-64), 344 KB (M5), plus a copy of the compressed file | 156 µs / 71 µs |
| libopus, float | 365 ms / 250 ms | 27 KB | 46 µs / 37 µs |
| libopus, fixed point | 422 ms / 298 ms | 27 KB | 51 µs / 40 µs |
| opusfile, float | 379 ms / 259 ms | 104 KB | 96 µs / 66 µs |
| stb_vorbis | 338 ms / 154 ms | 211 KB | 369 µs / 92 µs |
| libvorbis | 377 ms / 210 ms | 98 KB (214 KB peak) | 437 µs / 178 µs |
| Tremor | 357 ms / 235 ms | 115–151 KB (194 KB peak) | 402 µs / 191 µs |
| IMA ADPCM | 47 ms / 29 ms | a few bytes per channel | 4 µs / 13 µs |
- Falcon's per-stream memory is its main cost for many-voice use.
- Each channel has its own transform plans and windows.
- The C API also keeps its own copy of the whole compressed file, so 256 voices of the same 2 MB file hold 256 copies.
- About 140 KB of tables are built once per process on the first decode.
- The decoder makes about seven heap allocations (about 8 KB in total) per 20 ms frame.
- libopus allocates nothing after
opus_decoder_create. The caller supplies the PCM buffer.
Code size (machine code / read-only data, linked into a minimal program, arm64, unused code stripped):
| Decoder | Code | Read-only data |
|---|---|---|
| Falcon (decoder only) | 629 KB, of which rustfft 310 KB, Rust std 229 KB, rustdct 31 KB, Falcon itself 44 KB | 67 KB |
| libopus, float (with libogg) | 130 KB | 43 KB |
| opusfile + libopus + libogg | 148 KB | 44 KB |
| stb_vorbis | 43 KB | 1 KB |
| libvorbis + vorbisfile + libogg | 117 KB | 67 KB |
| Tremor + libogg | 57 KB | 53 KB |
| IMA ADPCM | 3 KB | 0.3 KB |
On x86-64 the Falcon DLL, which includes the encoder, has 1.4 MB of code. Most of it is rustfft's scalar, SSE and AVX kernels.
How this was measured:
- Hardware: an Intel Core Ultra 9 185H laptop running Windows 11, with the thread pinned to one performance, efficiency or low-power efficiency core; and a MacBook Pro with Apple M5 running macOS 26, on its performance cores.
- Falcon build: the workspace release profile (fat LTO, one codegen
unit). On x86-64 it targets
x86-64-v3; on the M5 it uses the default aarch64 target.- On Windows, Falcon is the MSVC-toolchain DLL.
- A MinGW (
x86_64-pc-windows-gnu) build of Falcon decodes 1.45× slower (771× in total). It links msvcrt'spow/exp/log, which the decoder calls for every band.
- C builds:
-O3. On x86-64, GCC 15.1 (MinGW-w64) with-march=x86-64-v3; on the M5, Apple clang 17 with the default target. libopus uses its default SIMD: SSE4.1/AVX2 runtime dispatch on x86-64 and NEON intrinsics on arm64. Its stack protector is off. - Timing: each run times open, whole-file decode and close. Every file
gets at least 10 runs and at least 1 s of runs. The figure is the best
run over two separate passes.
- Raw libopus is fed packets that were extracted from the Ogg file beforehand, so its figures exclude Ogg parsing. opusfile, vorbisfile and stb_vorbis include their container parsing.
- Each decoder writes its native output format:
f32for Falcon, float libopus, opusfile, stb_vorbis and libvorbis;int16for fixed-point libopus, Tremor and ADPCM.
- Output check: every decoder's output was checked against the original for time alignment and SNR.
- Earlier numbers: earlier versions of this README compared
falcon bench-decodewithffmpeg -threads 1 -stream_loop 19 -i <file> -f null -.- That method measured FFmpeg's built-in Opus and Vorbis decoders, not libopus and libvorbis as previously stated.
- Re-running it on the same laptop gives 892× for Falcon, 514× for FFmpeg's Opus decoder and 1006× for its Vorbis decoder, in total. Falcon is faster than FFmpeg's Vorbis decoder on only 5 of 18 files.
- On Apple M5 the totals are 1539×, 730× and 1737× (7 of 18 files).
- The earlier Vorbis figures could not be reproduced and have been withdrawn.
Corpus: 18 files. All are 48 kHz stereo 16-bit.
- 11 real recordings: game BGM, SFX (impacts, noise beds, applause), classical, electronic and pop music, female and male speech, and singing. They are not redistributed.
- 7 synthetic torture signals: click trains ("castanets"), inharmonic
decaying partials ("glockenspiel"), a fast detuned arpeggio
("harpsichord"), dense detuned chords ("polychord"), a 20 Hz–20 kHz sweep,
gated pink-noise bursts, and steady 8–15 kHz tones over a pink floor
("hftonal").
tools/quality/make_torture.pyregenerates these bit-exactly, using fixed seeds and numpy only.
Size budget. The Opus reference files were made with libopus (VBR) and
come out at 96–214 kbps depending on content. Each Falcon file is encoded
with --target-bytes equal to its Opus file's size, so it is at most that
large: 0.998–1.000× on 17 files and 0.53× on the sweep. Vorbis and IMA
ADPCM appear only in the speed comparison.
Metrics. All metrics compare against the original. The two spectral distances are standard ones from the literature, not project-invented:
-
MR-STFT: multi-resolution STFT log-magnitude distance (FFT sizes 512/1024/2048, 75% overlap), as used in the neural-codec literature (SoundStream, EnCodec).
-
Log-mel L1: L1 distance between log-mel spectrograms (64 mel bands, up to 20 kHz).
-
Perceptual A/B check: an objective multi-detector audio-quality checker is run as a paired comparison. The reference is the original, A is the Opus decode, and B is the Falcon decode, over all 18 pairs.
- Each detector's per-file differences go into a sign test (verdict threshold p < 0.1).
- Only detectors whose accuracy has been validated decide the overall verdict.
- The detectors estimate overall quality (MOS-style), SNR against the reference, noisiness, harmonic amplitude modulation against the reference, hum and effective bit depth. Bandwidth-cliff, clipping and dropout detectors are reported but do not decide the verdict.
The checker is not part of this repository. The spectral metrics can be reproduced with
tools/quality/.
Honesty checks. For these results, the cross-correlation lag between
original and decode is exactly 0 samples on every file, and the RMS
level offset is within 0.04 dB. tools/quality/gates.py implements these
checks, together with NaN/Inf and clipping bounds, for your own runs. A
codec cannot flatter its scores with a hidden time shift or gain change.
Reproducing the spectral metrics:
# needs ffmpeg on PATH and Python with numpy + soundfile (pesq / pystoi optional)
python tools/quality/make_torture.py corpus/wav_original # synthetic part of the corpus
# add your own 48 kHz WAVs to corpus/wav_original and encode matching Opus files into corpus/opus/
python tools/quality/compare.py corpus/ # iso-size encode, decode, metrics, reportcompare.py encodes each reference at the byte size of its .opus
counterpart, decodes all codecs with ffmpeg / falcon, and writes a
markdown + CSV report to tools/quality/reports/. metrics.py also computes
per-band energy error, HF retention, and PESQ/STOI for speech, when the
optional packages are installed.
- Not "better than Opus everywhere".
- On the perceptual check Falcon is at parity overall. Opus is rated better on harpsichord, pinkbursts, polychord and sfxtransient.
- Harmonic amplitude modulation (a slow flutter of partial amplitudes against the original) is still higher than Opus's on harpsichord and polychord.
- Log-mel distance is higher than Opus's on 5 of 18 files.
- Rate floor. At the finest quality index Falcon cannot always spend the budget. The synthetic sweep reaches index 0 at 0.53× the Opus size, and the encoder never pads.
- 48 kHz design. The band tables and constants assume 48 kHz. Other rates are stored and played back, but they are not tuned.
- No packet-loss concealment, no low-delay mode. Frames are 20 ms with a 20 ms look-back (TDAC), plus encoder look-ahead for transient scheduling. Falcon targets files and streams, not interactive calls.
- Seeking is linear. The CLI and C API encoders emit a random-access
frame only at the start (see
ra_interval), and there is no seek index. - Floating-point contract. The pulse-count law uses
f32powf/expon both sides, so an encoder and a decoder must compute identical values. This is verified on the development platform (x86-64, Windows). Agreement between different platforms' math libraries is expected but has not been tested systematically. The prefix-code tables are built with integer arithmetic only. - Mono/stereo focus. Up to 64 channels are supported (M/S applies to
stereo only). Dense many-channel content may use the coarser
Extframes. - Encoder speed. Rate control runs several full encodes in a binary search. Encoding is much slower than decoding, by design.
- Files: 1.x
.falconfiles (version0x04) cannot be decoded by 2.x. Re-encode from the source audio. - Rust API:
EncoderConfig::qualityand theQualityenum are gone (they had no effect).PsychoacousticModelis removed.falcon_core::coeffcodeis nowfalcon_core::rate.Mode::LpcandMode::Hybridare nowMode::FlaggedandMode::Ext. See the migration table. - CLI:
-q/--qualityis removed. Use--quality-index,--target-bytesor--target-kbps. - C / C++: no API change (only the version string).
Falcon's code is MIT licensed, and the project charges no royalties. Version 2.0 was redesigned specifically to avoid four patents that could read on 1.x:
- US 8,838,442 (spreading rotation)
- US 9,009,036 (recursive gain-shape splitting)
- US 9,015,042 (decoder-side collapse noise)
- US 11,234,023 (rANS)
A preliminary search of the new 2.0 elements found no plausible reading of active claims on the shape coder, energy coding, class-controlled fill or pulse search. The encoder's transient decision was also reworked into a per-frame analysis of each frame's own window span, with no carry-over between frames, to avoid the cross-frame "hangover" indicator claimed by the Ericsson family US 9,495,971 et al. One open item is documented: the mid/side decoder structure, against a broad reading of US 11,527,252.
PATENTS.md has the full analysis, sources and limits. It is not legal advice and gives no warranty. Falcon contains no code from other codec implementations (NOTICE), and all dependencies are permissively licensed (THIRD_PARTY_LICENSES.md).
Opus and Vorbis are trademarks of their respective owners. They are named here only for factual comparison as open reference codecs. Falcon is an independent project and is not affiliated with or endorsed by them.
- J. P. Princen, A. B. Bradley, "Analysis/synthesis filter bank design based on time domain aliasing cancellation," IEEE Trans. ASSP, 1986.
- B. Edler, "Codierung von Audiosignalen mit überlappender Transformation und adaptiven Fensterfunktionen," Frequenz 43, 1989.
- T. R. Fischer, "A pyramid vector quantizer," IEEE Trans. Information Theory, 1986.
- D. A. Huffman, "A method for the construction of minimum-redundancy codes," Proc. IRE, 1952.
- S. W. Golomb, "Run-length encodings," IEEE Trans. Information Theory, 1966; R. F. Rice, "Some practical universal noiseless coding techniques," JPL Publication 79-22, 1979.
- J. Herre, D. Schulz, "Extending the MPEG-4 AAC codec by perceptual noise substitution," AES 104th Convention, 1998.
- ITU-T Recommendation G.722.1, "Low-complexity coding at 24 and 32 kbit/s for hands-free operation in systems with low frame loss," 1999.
- J. D. Johnston, A. J. Ferreira, "Sum-difference stereo transform coding," ICASSP 1992.
- E. Zwicker, "Subdivision of the audible frequency range into critical bands," JASA, 1961.
- Xiph.Org Foundation, Vorbis I specification (window function).
- N. Zeghidour et al., "SoundStream: An end-to-end neural audio codec," 2021; A. Défossez et al., "High fidelity neural audio compression" (EnCodec), 2022 (MR-STFT methodology).
Issues and pull requests are welcome. Please read CONTRIBUTING.md first. In particular, it asks for patent hygiene and clean-room code: no code from other codec implementations.
MIT, see LICENSE. Copyright (c) 2026 Falcon Audio Codec contributors.

