|
DSPark 1.8.0
Header-only C++20 DSP for real-time and offline audio
|
Working recipes for common production tasks. Every snippet assumes:
Most in-place effects use construct -> prepare(spec) on the setup thread -> processBlock(view) in the audio callback. Preparation may allocate; processing does not. Each header states which setters can run concurrently: structural changes such as oversampling can allocate and require processing to be stopped. reset() clears processing history when the stream restarts.
Use getLatency() for audio-path delay, in samples at the prepared rate. ProcessorChain sums that name; getLatencySamples() on Saturation, TapeMachine, TubePreamp and TransformerModel remains a compatibility alias. OnsetDetector and BeatTracker use getLatencySamples() for the delay of an analysis result: that does not delay the audio and is not added to host PDC. Saturation::process(view) remains a compatibility alias; new in-place code uses processBlock(view). The pointer overloads on Gain and StereoWidth, and the whole-signal TimeStretch call, have different purposes.
For a plugin, expose the sum of serial audio-stage latencies from its getLatency() const noexcept method; DSPark's format wrappers report it to the host. Re-query after changes that affect latency. Parallel branches need alignment to the longest branch before summing. ProcessorChain's hard bypass skips a slot without inserting a delay, so use the effect's latency-compensated dry/wet path when the bypass must retain alignment.
Some utilities require different routing:
| Utility | Processing contract |
|---|---|
| AutoGain | pushReference(view) before an effect, compensate(view) after it; no added sample delay. |
| CrossoverFilter | processBlock(input, bandOutputs, count) writes separate bands; report the split latency once. |
| MidSide | Static encode(view) / decode(view); no preparation, history or state blob. |
| Crossfade | Two inputs and one output, with smoothing history; optional prepare(spec) enables a 20 ms glide. No state blob; store position and curve in the enclosing processor. |
Generalize the same pattern to any parameter with ModulationRouter (Core/ModulationRouter.h).
The LoudnessMeter passes the official EBU R128 vectors (Tech 3341/3342: integrated, LRA and true peak) - see conformance/.
For a finite delivery, call meter.finalizeTruePeak() after its last block and then read getTruePeakDb(). This includes the interpolation tail without adding silence to loudness windows. For exact source-frame selections, use `AudioIntervalAnalyzer`; it reports complete-window validity and distinguishes regional loudness from continuous observations.
All three are physical models (Koren triode + WDF FMV tone stack; flux-domain and tape-calibrated Jiles-Atherton hysteresis), loudness- compensated: drive moves saturation, not volume. Each oversamples its own nonlinear core (tube and transformer 2x, tape 4x by default; setOversampling(1) turns it off), so report the sum of the three getLatency() values to the host.
Oversampling belongs to the product, not to each module: wrap the whole nonlinear section once instead of paying one resampler (latency + CPU + band-limiting) per effect.
Measure before reaching for this: DSPark's own nonlinear stages (Saturation, TapeMachine, Clipper) already oversample internally where the algorithm needs it, and the Compressor's gain path stays at or below -72 dBc of aliasing at 1x even in its worst case (FET character at minimum attack). A section like the one above earns its resampler when you drive custom waveshaping hard, not for dynamics alone.
Transparency: every processor that oversamples internally exposes the factor via setOversampling(int) with factor = 1 meaning OFF (no internal resampling), supports at least {1, 2, 4}, and reports the added latency through getLatency() so the host can compensate. DynamicEQ embeds such a stage and defaults to 1x (off): turn it up to 2x or 4x only when a band's detector needs the extra alias suppression, and leave it at 1x when you already oversample the whole section (above) to avoid cascaded resamplers. Cost scales roughly linearly with the factor; the exact per-factor latency is whatever getLatency() returns for the active setting.
The nonlinear stages that generate aliasing inside the model default to internal oversampling and expose the same control (setOversampling is setup-thread only - it reallocates and may re-calibrate like prepare()):
TapeMachine defaults to 4x. The AC-bias carrier sits at a quarter of the internal rate, so 4x/48k puts it ultrasonic (48 kHz), and every harmonic the hysteresis draws from it folds onto 0, the carrier or the internal Nyquist frequency, where the downsampler removes its products with the programme. CPU scales ~linearly with the factor (4x is the reference cost of physical AC bias). setOversampling(1) turns the resampler off (zero added latency) but then drops the carrier in-band (12 kHz at 48k) - use 1x only under a high host rate or an already-oversampled section.Saturation defaults to 2x: at 1x its curves fold their harmonics back into the band (a 10.1 kHz tone at -6 dBFS through the default SoftClip left an alias 34 dB down; 124 dB down at 2x). setOversampling(1) gives zero latency for material that never reaches the curve's knee.TubePreamp defaults to 2x. At factors 2/4/8/16 the circuit is solved in continuous time inside each internal sample interval: the grid input is reconstructed by a Farrow polynomial, each interval is split at the triode knees and the tone circuit is propagated in its analog modal form; the plate voltage is band-limited once, at the output. getLatency() reports 71/99/113/121 samples respectively, independent of the active stage count; the dry path includes the same delay. Factor 1 retains the point circuit without resampling and reports zero: use it inside an already oversampled host chain. At 2x the header's frequency/drive sweep stays at or below -80.6 dBc up to +36 dB drive; higher factors lower it further at a higher processing cost. For TapeMachine getLatency() at 1x is NOT zero: it still reports the loss-FIR centre + transport delay (127 at 48k); only the oversampler contribution drops to 0. Always query getLatency() for the active factor rather than assuming a value.Clipper and Core/WaveshapeTable also expose setOversampling (both defaulting to 1x=off); WaveshapeTable is a memoryless table (no ADAA), so its Hermite interpolation reduces table noise but not aliasing - oversample it to suppress alias products. The Compressor does NOT run an internal audio-path resampler: its optional TruePeak mode oversamples only the detector (ITU-R BS.1770 inter-sample peak measurement), which adds no audio-path latency and cannot cascade, so there is no factor to configure there.
The Compressor's characters are calibrated against the published hardware figures, so classic units map to plain settings. Feedback operation lands on the requested static curve exactly (the element's law is the closed-form inverse of the user's curve, matching how hardware panels are marked with observed ratios), and the loop is resolved semi-implicitly with the peak detector, so it stays stable down to 20 us attacks.
LA-2A style leveler (Teletronix spec: 10 ms attack, ~50% release in 0.06 s, complete release 0.5 to 5 s depending on programme, ~3:1, gentle knee):
Measured on this recipe: 50% release 64 ms after a long squeeze, complete release ~2.1 s (and faster after brief peaks: the memory stage only charges under sustained compression), knee floor engaging right at the threshold.
1176 style FET limiter (UREI/UA spec: 20-800 us attack, 50-1100 ms release, panel ratios 4/8 compress and 12/20 limit, THD < 0.5% while limiting):
Measured: settled gain reduction lands on the panel curve (19 dB at 20:1 with the level 20 dB over threshold), observed attack t63 21 us at the fastest setting with the loop stable, colour THD 0.42% at -6 dBFS programme while limiting (0% with colour off). Driving low frequencies at the fastest attack/release rides the waveform within the cycle exactly like the hardware (several percent THD at 100 Hz: that is the 1176 grit, back off attack or release to clean it up). All-buttons mode is not modeled.
Fairchild 670 style vari-mu (0.2-0.8 ms attack; release 0.3 to 25 s across the six Time Constant positions, the slowest ones programme dependent):
Notes that apply to all three: with a memory detector (RMS) in FeedBack keep the attack at or above the detector window, or the loop hunts on the window's lag (a 0.1 ms attack against a 10 ms RMS window pumps ~3 dB at 20:1; the hardware units above are all peak detected, which resolves implicitly and does not hunt). Release knobs are t63 measured after the signal drops, so the loop does not alter them.
TimeStretch changes how long a passage takes without touching what note it is on. Tempo is a ratio, so a passage recorded at 120 BPM played back at 110 BPM lasts 120 / 110 times as long:
setTempoChangePercent() is the same control from the other end, in the units a tempo knob is marked in: -8.33f is 120 BPM played at 110, +8.33f is 120 played at 130. Either way the realised ratio is exact and does not drift: measured over 30 s at 48 kHz the stretch lands within 0.003% of the target, and the output is exactly round(inputLength * ratio) samples long, so a stretched loop still meets its own end.
Drums and other strikes need no setting. Attacks are found with spectral flux over a log-frequency filterbank and the analysis hop is held across each one, which is what keeps a strike from doubling when the stretch spreads it. Measured on a click train, the stretched strike is as concentrated as the input's own at every frame size from 256 to 4096 and at every ratio, and no onset lands more than 1.71 ms from where the stretched timeline puts it - at 120, 480 and 960 strikes per minute alike. It is on by default; setTransientPreserve(false) turns the whole transient path off, which is useful for hearing what it does and is not otherwise a better setting.
The one limit worth knowing: the hold needs a gap between strikes to repay the input it did not consume, and the density it can serve scales as sampleRate / fftSize. At the default 2048-sample frame and 48 kHz that holds to about 1040 strikes per minute; past roughly 1060 the timing degrades to about 3.9 ms. A shorter frame moves the limit up in proportion.
Streaming: use the rate-changing pair away from ratio 1. A stretch changes duration, so the honest streaming shape is one where the input count and the output count are allowed to differ. That is feedInput() / pullOutput():
There is no latency to compensate on that path: what comes out is the stretched timeline itself, output sample k being sample k of the stretched signal. It primes first - pullOutput() returns 0 until about one frame of input has been fed - and after that the output tracks ratio * the input it has consumed, less a fixed offset of one frame minus one analysis hop that is the overlap-add's own incomplete tail. Feeding without pulling makes getInputCapacity() fall to 0, which is the back pressure that keeps the two ends honest.
The block path is a fixed-rate playback adaptor. processBlock() hands back exactly as many samples as it was given, latency getLatency() (one frame: 2048 samples, 42.7 ms at 48 kHz). At ratio 1 it is exact indefinitely. Away from unity it cannot be, and the cost is stated in numbers rather than described: below unity exactly 1 - ratio of the output is silence (measured 20.00% at 0.8, 7.40% at 0.926, 5.00% at 0.95, worst error 0.03 percentage points over durations of 5 to 30 s and blocks of 64 to 4096), delivered as gaps of at most one block; above unity the block physically cannot carry the stretched stream, so the adaptor refuses the fraction 1 - 1/ratio of the input at the head, spread evenly, and counts every refused sample in getDiscardedInput(). Nothing that survives is displaced - measured worst 38 samples, 0.79 ms, at ratio 1.081 - which is the whole point of refusing at the head rather than letting a queue fill and splice the stream. What it cannot preserve is a strike's height: on the same strike train the mean strike height, as a fraction of the same class's own ratio-1 rendering, is 0.85 at ratio 1.01, 0.79 at 1.02, 0.38 at 1.05, 0.82 at 1.081, 0.89 at 1.25, 0.56 at 1.5 and 0.11 at 2 - not monotonic in the ratio, so those figures cannot be interpolated between. Use the pair above, or process(), unless the ratio is 1.
The two streaming paths own the same queue with different invariants, so whichever one is used first after prepare() or reset() owns the instance until the next reset(); the other one's calls do nothing in the meantime.
At ratio 1 the streaming path is transparent: measured residual against the delayed input is -146 dBFS at 48 kHz on a 1 kHz tone, so leaving the effect in a chain unengaged costs nothing but its latency.
BeatTracker answers two different questions with two different methods, and which one you want depends on whether you have the whole file.
Offline, when you do. analyze() sees the future, so it can place beats using evidence that arrives after them:
On a click track from 40 to 240 BPM the tempo lands within 0.001 BPM of the truth at the correct metrical level across the whole range, and every beat is found within 7.4 ms at worst. On a 100-to-140 BPM ramp the grid follows the tempo being played rather than the average: worst local inter-beat error 1.03%.
The signal must be long enough for the range in force. analyze() needs four beats at the SLOWEST tempo it is searching, not at the tempo you have, so at the 40 BPM default floor it needs 6 s of audio whatever the material is doing; below that it returns an empty grid and a tempo of 0 rather than a tempo fitted to nothing. A six-second loop at 120 BPM contains twelve beats and still returns nothing at default settings. Narrowing the range with setTempoRange() lowers the requirement in proportion: 70 BPM as the floor brings it down to 3.4 s.
Read the confidence before you trust the grid. It is the share of the onset strength that falls in phase with the beats returned, so 1.0 means the grid explains everything the signal did, reduced by how nearly another metrical level explains the same thing. On a three-against-two polyrhythm, where two pulses are genuinely present, it reports 0.02 to 0.38 rather than pretending one of them is the answer, and where two readings are exactly level it goes to 0.
It cannot tell you the level is wrong, and no number can. A click train whose alternate events are quieter is at once a beat with a backbeat and a half-speed beat with straight eighths – the same samples, two different right answers – so anything computed from the signal is identical for both. Do not gate on the confidence for this: read secondaryTempoBpm whenever the metrical level matters. Measured on eighth notes swept in amplitude, the tracker reports the beat while the eighths stay below roughly 0.44 of the beat's onset strength at 90 BPM, 0.72 at 100 and 0.90 at 120, and reports the eighth level above that – with the beat in secondaryTempoBpm, and with a confidence near 1.0, because that grid genuinely does explain the onsets. Below 90 BPM the eighth level is reported from about 0.28 up, which is the rate a listener would tap rather than a defect.
Half and double tempo are the error that matters, so check the level, not just the number. secondaryTempoBpm names the alternative a listener could plausibly have tapped instead and is 0 when there is no distinct alternative inside the searched range. Narrow the range if you know the material:
In the callback, when you do not have the future. Feed blocks and read the running estimate:
On a click track, over the whole 40 to 240 BPM range this class searches by default, the running tempo reaches within 5% of the truth inside 3.6 s at worst – 60 BPM is the slowest to settle – holds to 0.34% after that, and lands at the correct metrical level at every point of the range; beats are attributed to within 13.4 ms. The scope matters: measured only from 60 to 180 BPM, this path once published five times the true tempo at 40, 42, 44 and 46 BPM and nothing in the sweep could see it. getLastBeatSample() is where the beat happened in your own timeline; getLatencySamples() is how far in the past that is by the time you are told (549 samples, 12.4 ms, at 44.1 kHz). Use the former to align anything and the latter only to state your own delay.
The order of those two calls is part of the contract, and the snippet above uses it: test beatNow() first, read getLastBeatSample() second. They are two separate loads and the processing call can run between them, so the reverse order hands you a fresh "a beat just happened" together with the position of an earlier beat – measured at a full beat period, and at more than one when the two reads are further apart than a beat. See the threading page.
One behaviour to know: with no onset energy arriving, the pulse estimate and the mass it is measured against decay together, so the confidence HOLDS its last value through silence instead of falling. A caller that needs to tell "steady" from "nothing playing" reads level separately.
BeatTracker consumes OnsetDetector's onset-strength envelope, so if you want onsets too, take them from getOnsetDetector() rather than running a second copy of the same analysis.
SpectralFreeze captures the sound that is playing and sustains it as long as you hold the request: the frequency-domain counterpart to GranularProcessor's time-domain grain freeze. In tonal mode the captured partials keep advancing at their own measured frequencies with the phase relations of the source locked in place, so a held piano chord stays a chord. Diffuse mode rotates every bin by a deterministic random sequence instead, which turns the same capture into a wash that never repeats yet renders the same on every run.
Enter and exit are a single spectral crossfade of getTransitionSamples() (= N) samples, so there is no click in either direction, and toggling the request mid-transition just reverses the glide without recapturing. A mode change while frozen is remembered and applies at the next capture; the running hold never switches texture underneath you. Both setters are lock-free and safe from any thread.
LoopFinder answers "between which two samples can this sustain loop
without a seam?" It is an offline search – run it when loading a sample, not on the audio thread – and it is explicit about failure: a result with status != Success carries no plausible-looking endpoints to misuse.
The search scores every candidate pair on three things at once – waveform match, slope match and spectral similarity – so silence, DC or a loud-but- different section cannot win, and it refuses (NoAcceptableLoop) rather than return a bad splice. renderLoop() writes one cycle of end - start - crossfadeLength samples whose last sample flows into its first: the overlap is an equal-power crossfade renormalized so correlated material keeps unit gain instead of the raw +3 dB center bulge.
PitchCorrector is the detect-quantize-shift chain in one processor: it tracks the fundamental with the framework's YIN detector, snaps it to the nearest note of the scale you choose, and shifts by the difference with the phase-vocoder pitch shifter. The retune speed is the whole character control – 0 ms is the hard, quantized snap; a few tens of milliseconds is intonation help nobody hears.
Report getLatency() to the host: the shifter's 4096 samples (85 ms at 48 kHz) are constant and must be compensated like any other look-ahead.
How fast a hard snap lands follows from one property of the mask you choose, its widest gap - the largest distance between two adjacent notes of the scale. The largest correction is half that gap; the largest CHANGE in the correction is the whole gap, taken when the singer crosses the middle of it. The phase vocoder moves the shift by at most half a semitone per analysis hop, so settling grows with the change (worst case over control phase, both directions, 48 / 44.1 kHz):
A half-semitone correction change settles within 12.5 / 14.5 ms under the same 32-phase measurement.
| widest gap | scales | largest correction | largest change | settles in |
|---|---|---|---|---|
| 1 st | Chromatic (the default) | 0.5 st | 1 st | 23.3 / 25.6 ms |
| 2 st | major, minor, the modes, whole tone | 1 st | 2 st | 44.6 / 49.0 ms |
| 3 st | harmonic minor, the pentatonics, and 28 others | 1.5 st | 3 st | 66.4 / 72.7 ms |
| 4 st | Hirajoshi, InSen, Balinese, Chinese, Japanese | 2 st | 4 st | 89.2 / 96.6 ms |
| 12 st | a one-note mask | 6 st | 12 st | about 265 / 288 ms |
Up to one semitone the transition finishes inside the effect's own latency and is inaudible; wider gaps are audible by their excess. They are traversals, not errors: the pitch glides continuously and lands within 10 cents. If you need a bounded response, choose a scale whose widest gap is a whole tone or less. Time to target grows with the retune speed above about 8 ms; shorter glides are already at the vocoder's floor and sound identical.
Those milliseconds hold at every sample rate, because the analysis frame is chosen from the rate to keep its span near 43 ms rather than being pinned to a sample count: 2048 samples at 44.1/48 kHz, 4096 at 88.2/96 kHz, 8192 at 176.4/192 kHz. getFrameSize() reports the choice and getLatency() grows with it, staying near 85 ms in time. Pass a frame to prepare() if you want a different trade - a shorter one settles faster and stops resolving the harmonics of a bass voice, which is the whole reason the default is what it is. Halving the default span costs a corrected F2 most of its harmonic structure: its harmonic-to-residual ratio falls from 42-48 dB to below about 7.5 dB, and A2's from 38-41 dB to 4.0-4.4 dB, while A3 an octave up moves only from 47-51 dB to 40-42 dB. Those decibels are read by a four-term Blackman-Harris DFT over one un-padded 0.68 s segment, which reads about 89 dB on a signal built to have no residual at all, so it has the headroom to see them.
The scale mask is a 12-bit harmony::NoteSet (bit k = k semitones above the root), so anything in harmony::allScales works, as does a mask you build yourself; an empty mask switches correction off without leaving the chain. The default is the chromatic scale, which snaps to the nearest semitone whatever the key. Correction holds through consonants and breaths instead of sagging back to the sung pitch, and it is monophonic by construction: one voice, one fundamental. On chords the detector reports whichever periodicity wins, and the correction follows that.
AutoGain adds no audio-path latency. Its 400 ms level integration is a measurement response; the effect's getLatency() still belongs in the host report. Compensation is bounded and smoothed, so transients need not match instantaneously, especially around a delayed effect or a reverb tail.
Report split.getLatency() once, plus the longest aligned branch delay. Keep the configured band count equal to the number of allocated outputs. The return value names the bands actually written; unused views retain old data. Minimum-phase bands sum to an allpass response, so a parallel raw dry path also needs phase matching. Linear-phase mode adds FIR delay; switching mode resets history and requires updating the host's latency report.
Use a stereo spec at 8 to 384 kHz. Duplicate mono into two channels explicitly when needed. The second option is a low cut on the generated delta only: zero disables it; otherwise use 20 to 5000 Hz. The original mid and original side remain present. Width zero gives exact delayed identity; automation takes 5 ms and does not alter the reported latency. No limiter or output gain trim is added.
The local color factor accepts 1, 2, 4, 8 or 16 and defaults to 1. At 1x, continuous polynomial reconstruction and nonlinear moment integration run at the source rate, followed by FIR projection. There is no upsampled audio stream. It retains antialias filtering and the original source-clock DC response; 1x does not mean an unfiltered memoryless waveshaper. Its FIR transition spans 0.45 to 0.50 times the source sample rate at every factor. Without the optional low cut, 1x reports 256 samples of latency (5.33 ms at 48 kHz); the established explicit 4x path reports 392 samples at that rate. Factor 2 trades antialias rejection for CPU; a larger factor does not automatically outperform the different algorithm used at 1x. Both paths use the shared Core clipping curves, integration, FIR and DC kernels. The fixed 1x coefficient bank is shared across instances (82496 bytes of read-only doubles); all mutable filter state is private.
Options require setup-time preparation. Presets store parameters, not filter history. After a seek, use resetAtFrame() and replay preceding input when history continuity is required; this method alone is not a saved-state restore. Flush with zeros when rendering a tail. Alignment latency is not tail duration.
For parallel dry/wet processing, prepare DryWetMixer with the same stereo spec, select MixRule::Linear, and call setLatencyCompensation(generator.getLatency()) during setup. In each callback, call pushDry() with the original input, process the generator, and then call mixWet(). The mixer delays its dry copy internally; adding another dry delay would misalign the paths. Linear mixing preserves the level of the mid component shared by both paths. Reset both histories on a seek or stream restart.
The compiled stereo mixing example checks delayed identity, preserved mid, source-clock advancement and equal output across block partitions, using float/double, two rates, factors 1/4 and the optional low cut. It reports the latency of each prepared configuration.
For complete-source jobs, OfflineStereoGenerator handles alignment, bounded reads, exact exclusions, peak measurement and transactional output. Its optional host-owned delta cache avoids regenerating bands/color for each width change. See offline stereo generation.