|
DSPark 1.8.0
Header-only C++20 DSP for real-time and offline audio
|
FIR filter using direct-form convolution with a mirrored delay line. More...
#include <FIRFilter.h>
Public Member Functions | |
| FIRFilter ()=default | |
| ~FIRFilter ()=default | |
| void | prepare (int maxTaps, int numChannels) |
| Pre-allocates memory and initializes the delay lines. | |
| void | setCoefficients (std::span< const T > coeffs) noexcept |
| Sets the filter coefficients asynchronously. | |
| void | reset () noexcept |
| Resets all delay lines to zero, clearing the filter's memory. | |
| void | processBlock (AudioBufferView< T > buffer) noexcept |
| Processes a full audio buffer in-place. | |
| T | processSample (T input, int channel) noexcept |
| Processes a single sample through the FIR filter. | |
| int | getLatency () const noexcept |
| Returns the filter's group delay (latency) in samples. | |
FIR filter using direct-form convolution with a mirrored delay line.
Implements a lock-free coefficient publish (seqlock) for safe asynchronous updates from the control thread without reallocating memory on the audio thread; the convolution itself runs through simd::dotProduct.
Threading / data-race freedom: the SHARED staging coefficient buffer is std::vector<std::atomic<T>>. The single control-thread producer stores each word relaxed inside an odd/even seq-counter critical section, with a release fence between the odd increment and the word stores; the audio thread loads each word relaxed under the seqlock retry (with an acquire fence before the counter re-read). The two fences pair ([atomics.fences]/2): a reader that observes any mid-publish word is forced to observe the odd counter and retry. Because every cross-thread word access goes through std::atomic, there is no non-atomic concurrent read/write - the handoff is UB-free by the C++ memory model, not merely tear-free in practice. std::atomic<float/double> is lock-free and single-word on every supported target, so the publish stays allocation- and lock-free. The audio thread copies the published set once per update into its PRIVATE, non-atomic activeCoeffs_ buffer, so the per-sample hot path (simd::dotProduct over activeCoeffs_) runs on plain scalars. Cost, in two dated steps, because the code has changed since the first number. 2026-07, when the staging buffer became atomic (measured against the pre-atomic code, g++ 13.3 -O2): the per-sample loops of processBlock<float/double> were instruction-identical (objdump: 207 == 207 and 200 == 200 insns; only the once-per-update pull copy changed), while whole-function processSample codegen differed by +3 prologue instructions from the reshaped cold pull path and measured FASTER (7.06 -> 6.88 ns/sample); block-level bench deltas were within run-to-run noise (5.78/5.89/5.95 vs 5.87/5.98/5.79 ns/frame medians, 63-tap mono LP). 2026-08, when the read became BOUNDED (the "Real-time bound" paragraph below): the outer per-sample loop of processBlock is no longer instruction-identical to the pre-atomic code – it went 79 -> 78 insns (float) and 73 -> 72 (double) with one loop-carried comparison spilled to the stack; the three inner dot-product kernels are unchanged (11/8/6 insns), and wall clock measured neutral, within run-to-run noise: +0.0% (g++ 13.3) / -1.4% (clang++ 18.1.3) at -O2, and +0.3% / -1.2% at -O3, on a 63-tap FIRFilter<float> processing 2-channel 512-sample blocks (medians of 11 pinned rounds against the pre-bound code).
Real-time bound: the audio thread's read is BOUNDED to kSeqlockMaxAttempts validation attempts. If none validates it adopts nothing, keeps the coefficient set already in use and re-arms the dirty flag, so the update lands on a later call instead of holding the callback open until the control thread finishes publishing. Without that bound the wait is the writer's scheduling, not this filter's instruction count: with 8192 taps published from a control thread sharing one CPU with the audio thread – the ordinary arrangement on single-core embedded and wasm targets – a single processBlock() call was measured at 5.1 SECONDS, and one block was all that completed in five seconds. Bounded, the same configuration's worst block is 6057 us against a 6045 us pure-CPU-contention control on the same machine and 58156 blocks complete, i.e. contention rather than waiting. Bounding cannot make a torn set adoptable: the accept test is unchanged and a read that fails it adopts nothing. What it can do is deliver a coefficient update one call late under sustained contention, which is the trade being made. The commit-on-validate rule is what makes it safe: the copy goes into the spare one of TWO private active buffers and the index flips only after the copy validates, so a give-up can never leave a half-copied set live.
Threading (SPSC model, see docs/threading.md):
| T | Sample type (float or double). |
Definition at line 363 of file FIRFilter.h.
|
default |
|
default |
|
inlinenoexcept |
Returns the filter's group delay (latency) in samples.
Definition at line 556 of file FIRFilter.h.
|
inline |
Pre-allocates memory and initializes the delay lines.
| maxTaps | Maximum number of filter coefficients supported. |
| numChannels | Number of concurrent audio channels to process. |
Definition at line 376 of file FIRFilter.h.
|
inlinenoexcept |
Processes a full audio buffer in-place.
Loads atomic state once per block to avoid intra-block tearing and pipeline stalls.
| buffer | Audio buffer view to process. |
Definition at line 466 of file FIRFilter.h.
|
inlinenoexcept |
Processes a single sample through the FIR filter.
processBlock instead when possible to avoid atomic load overhead on a per-sample basis.| input | Input sample. |
| channel | Channel index. |
Definition at line 518 of file FIRFilter.h.
|
inlinenoexcept |
Resets all delay lines to zero, clearing the filter's memory.
Definition at line 452 of file FIRFilter.h.
|
inlinenoexcept |
Sets the filter coefficients asynchronously.
Coefficients are stored reversed for direct SIMD dot-product alignment.
| coeffs | Span of coefficients. Size must be <= maxTaps passed to prepare(). |
Definition at line 418 of file FIRFilter.h.