DSPark 1.8.0
Header-only C++20 DSP for real-time and offline audio
Loading...
Searching...
No Matches
dspark::FIRFilter< T > Class Template Reference

FIR filter using direct-form convolution with a mirrored delay line. More...

#include <FIRFilter.h>

Public Member Functions

 FIRFilter ()=default
 
 ~FIRFilter ()=default
 
void prepare (int maxTaps, int numChannels)
 Pre-allocates memory and initializes the delay lines.
 
void setCoefficients (std::span< const T > coeffs) noexcept
 Sets the filter coefficients asynchronously.
 
void reset () noexcept
 Resets all delay lines to zero, clearing the filter's memory.
 
void processBlock (AudioBufferView< T > buffer) noexcept
 Processes a full audio buffer in-place.
 
T processSample (T input, int channel) noexcept
 Processes a single sample through the FIR filter.
 
int getLatency () const noexcept
 Returns the filter's group delay (latency) in samples.
 

Detailed Description

template<typename T>
class dspark::FIRFilter< T >

FIR filter using direct-form convolution with a mirrored delay line.

Implements a lock-free coefficient publish (seqlock) for safe asynchronous updates from the control thread without reallocating memory on the audio thread; the convolution itself runs through simd::dotProduct.

Threading / data-race freedom: the SHARED staging coefficient buffer is std::vector<std::atomic<T>>. The single control-thread producer stores each word relaxed inside an odd/even seq-counter critical section, with a release fence between the odd increment and the word stores; the audio thread loads each word relaxed under the seqlock retry (with an acquire fence before the counter re-read). The two fences pair ([atomics.fences]/2): a reader that observes any mid-publish word is forced to observe the odd counter and retry. Because every cross-thread word access goes through std::atomic, there is no non-atomic concurrent read/write - the handoff is UB-free by the C++ memory model, not merely tear-free in practice. std::atomic<float/double> is lock-free and single-word on every supported target, so the publish stays allocation- and lock-free. The audio thread copies the published set once per update into its PRIVATE, non-atomic activeCoeffs_ buffer, so the per-sample hot path (simd::dotProduct over activeCoeffs_) runs on plain scalars. Cost, in two dated steps, because the code has changed since the first number. 2026-07, when the staging buffer became atomic (measured against the pre-atomic code, g++ 13.3 -O2): the per-sample loops of processBlock<float/double> were instruction-identical (objdump: 207 == 207 and 200 == 200 insns; only the once-per-update pull copy changed), while whole-function processSample codegen differed by +3 prologue instructions from the reshaped cold pull path and measured FASTER (7.06 -> 6.88 ns/sample); block-level bench deltas were within run-to-run noise (5.78/5.89/5.95 vs 5.87/5.98/5.79 ns/frame medians, 63-tap mono LP). 2026-08, when the read became BOUNDED (the "Real-time bound" paragraph below): the outer per-sample loop of processBlock is no longer instruction-identical to the pre-atomic code – it went 79 -> 78 insns (float) and 73 -> 72 (double) with one loop-carried comparison spilled to the stack; the three inner dot-product kernels are unchanged (11/8/6 insns), and wall clock measured neutral, within run-to-run noise: +0.0% (g++ 13.3) / -1.4% (clang++ 18.1.3) at -O2, and +0.3% / -1.2% at -O3, on a 63-tap FIRFilter<float> processing 2-channel 512-sample blocks (medians of 11 pinned rounds against the pre-bound code).

Real-time bound: the audio thread's read is BOUNDED to kSeqlockMaxAttempts validation attempts. If none validates it adopts nothing, keeps the coefficient set already in use and re-arms the dirty flag, so the update lands on a later call instead of holding the callback open until the control thread finishes publishing. Without that bound the wait is the writer's scheduling, not this filter's instruction count: with 8192 taps published from a control thread sharing one CPU with the audio thread – the ordinary arrangement on single-core embedded and wasm targets – a single processBlock() call was measured at 5.1 SECONDS, and one block was all that completed in five seconds. Bounded, the same configuration's worst block is 6057 us against a 6045 us pure-CPU-contention control on the same machine and 58156 blocks complete, i.e. contention rather than waiting. Bounding cannot make a torn set adoptable: the accept test is unchanged and a read that fails it adopts nothing. What it can do is deliver a coefficient update one call late under sustained contention, which is the trade being made. The commit-on-validate rule is what makes it safe: the copy goes into the spare one of TWO private active buffers and the index flips only after the copy validates, so a give-up can never leave a half-copied set live.

Threading (SPSC model, see docs/threading.md):

Note
Buffers use the default heap alignment; the SIMD kernels use unaligned loads, which cost the same as aligned on modern CPUs.
Template Parameters
TSample type (float or double).

Definition at line 363 of file FIRFilter.h.

Constructor & Destructor Documentation

◆ FIRFilter()

template<typename T >
dspark::FIRFilter< T >::FIRFilter ( )
default

◆ ~FIRFilter()

template<typename T >
dspark::FIRFilter< T >::~FIRFilter ( )
default

Member Function Documentation

◆ getLatency()

template<typename T >
int dspark::FIRFilter< T >::getLatency ( ) const
inlinenoexcept

Returns the filter's group delay (latency) in samples.

Returns
Latency based on the most recently PUBLISHED tap count (the audio thread may adopt it one block later).

Definition at line 556 of file FIRFilter.h.

◆ prepare()

template<typename T >
void dspark::FIRFilter< T >::prepare ( int  maxTaps,
int  numChannels 
)
inline

Pre-allocates memory and initializes the delay lines.

Note
MUST be called offline (e.g., in prepareToPlay) before any processing.
Parameters
maxTapsMaximum number of filter coefficients supported.
numChannelsNumber of concurrent audio channels to process.

Definition at line 376 of file FIRFilter.h.

◆ processBlock()

template<typename T >
void dspark::FIRFilter< T >::processBlock ( AudioBufferView< T >  buffer)
inlinenoexcept

Processes a full audio buffer in-place.

Loads atomic state once per block to avoid intra-block tearing and pipeline stalls.

Parameters
bufferAudio buffer view to process.

Definition at line 466 of file FIRFilter.h.

◆ processSample()

template<typename T >
T dspark::FIRFilter< T >::processSample ( T  input,
int  channel 
)
inlinenoexcept

Processes a single sample through the FIR filter.

Warning
Use processBlock instead when possible to avoid atomic load overhead on a per-sample basis.
Parameters
inputInput sample.
channelChannel index.
Returns
Filtered output sample.

Definition at line 518 of file FIRFilter.h.

◆ reset()

template<typename T >
void dspark::FIRFilter< T >::reset ( )
inlinenoexcept

Resets all delay lines to zero, clearing the filter's memory.

Definition at line 452 of file FIRFilter.h.

◆ setCoefficients()

template<typename T >
void dspark::FIRFilter< T >::setCoefficients ( std::span< const T >  coeffs)
inlinenoexcept

Sets the filter coefficients asynchronously.

Coefficients are stored reversed for direct SIMD dot-product alignment.

Parameters
coeffsSpan of coefficients. Size must be <= maxTaps passed to prepare().
Note
Thread-safe single-producer publish via a seqlock. The audio thread copies the staged set into its private buffer atomically, so neither a (ptr, count) mismatch nor a mid-block overwrite can occur.

Definition at line 418 of file FIRFilter.h.


The documentation for this class was generated from the following file: