Introduction
Digital audio is often described in terms of clean numbers: sample rates, bit depths, signal-to-noise ratios. But much of what actually determines how a recording sounds happens at the margins of that clean picture, in rounding errors, timing tolerances, room reflections, and the physics of consonants hitting a microphone diaphragm. This report works through the main digital-domain artifacts, the acoustic and vocal-production issues that produce sibilance and its relatives, and the dynamic processing tools engineers use to manage them, and it goes one level deeper than a typical primer: into the numbers, the standards bodies, and the physical mechanisms behind each phenomenon, so the reader leaves with more than vocabulary. It sits alongside two existing pieces in this knowledge base rather than repeating them: The Recording Technique Lexicon covers the microphone and array decisions made before a note is captured, and The Frequency Spectrum covers how that captured energy behaves once it reaches a room. This piece picks up the middle ground, what happens to the signal between the microphone and the loudspeaker.
Part One: Digital-Domain Artifacts
Quantization error and dither
Every digital audio signal is a series of discrete amplitude values, and every conversion from a continuous analog waveform to a fixed set of levels introduces quantization error, the rounding difference between the true signal and the nearest representable digital value. The same error appears when reducing bit depth, for example converting a 24-bit master down to 16 bits for a CD or streaming release. At 16 bits there are 65,536 possible amplitude steps; at 24 bits, over 16.7 million. The error is small in absolute terms, but it correlates with the signal, which means it shows up as a harsh, gritty distortion rather than smooth noise, most audible on quiet passages, reverb tails, and fade-outs.
The theoretical ceiling this sets is described by the familiar dynamic range formula, SNR (dB) equals 6.02 times the bit depth plus 1.76, for a full-scale sine wave. That works out to roughly 98 dB at 16 bits and roughly 146 dB at 24 bits. In practice, no converter reaches the 24-bit figure. Analog noise in the input stage, reference voltages, and clock circuitry sets a real-world ceiling closer to 120 to 125 dB even in excellent modern ADCs and DACs. The extra bits below that floor are not wasted, though: they provide headroom for gain-staging and processing without pushing quantization noise back up to audibility, which is the practical argument for tracking and mixing at 24-bit even when the final release is 16-bit.
Dither is low-level noise, typically at or below the least significant bit, added before quantization specifically to mask this error. It sounds counterintuitive to add noise to improve fidelity, but the mechanism is well established: dither decorrelates the quantization error from the signal, converting a distortion artifact into a noise-like artifact the ear tolerates far better.
Three flavors are worth knowing:
- Rectangular PDF (RPDF) uses a single uniform random source. It removes signal correlation but leaves some residual noise modulation, meaning the character of the noise floor can still shift subtly with program content.
- Triangular PDF (TPDF) sums two independent RPDF sources. This is the standard choice in mastering, since it fully decorrelates the error with no noise modulation, at the cost of a noise floor roughly 4.8 dB higher than RPDF. That trade is considered worthwhile because noise modulation is more audibly distracting than a marginally higher, but stable, noise floor.
- Noise-shaped dither pushes dither noise into frequency bands where hearing is least sensitive, generally above 15 kHz, and away from the midrange, by filtering the dither signal before it is added. Total noise power is unchanged, but perceived noise drops substantially because the ear's equal-loudness contours are far less sensitive at the top of the audible band than in the 2 to 5 kHz region. Commercial mastering tools such as POW-r (Types 1 through 3, trading off noise-floor depth against high-frequency noise energy) and Sony's Super Bit Mapping implement this approach, each with a different noise-shaping curve tuned for a different balance of measured noise floor versus perceived noise.
Dither should be applied once, at the final bit-depth reduction, never stacked across multiple stages. Repeated dithering does not add resolution, only unnecessary noise. This is also why truncation, simply chopping off the lower bits without dithering at all, is worth naming as the mistake dither exists to prevent: truncation reintroduces the harsh, signal-correlated distortion dither is specifically designed to mask.
Aliasing
Aliasing is a different class of error, one of misidentification rather than rounding. Sampling theory holds that a digital system can only faithfully represent frequencies up to half the sample rate, the Nyquist frequency. Content above that limit cannot be recorded correctly, and if it reaches the converter unfiltered, the system does not simply discard it. Instead it folds that energy back down into the audible band as a new, unwanted frequency with no relationship to the original musical content, an artifact sometimes called a ghost frequency.
The folding follows a simple rule: for a sample rate of Fs, a tone at frequency f, where f sits between the Nyquist frequency and Fs, reappears as an alias at Fs minus f. A concrete example makes this tangible. At a 48 kHz sample rate, Nyquist sits at 24 kHz. A 30 kHz tone, inaudible on its own and well above the range most listeners can hear, would fold down to 18 kHz if it reached the sampling stage unfiltered, a frequency that is very much audible and bears no harmonic relationship to anything in the original material. This is precisely the failure mode anti-aliasing filtering exists to prevent.
This is why every ADC includes an anti-aliasing filter ahead of the sampling stage, and why oversampling architectures (converting at multiples of the target sample rate, then filtering and decimating) have become the standard approach: they push the required filtering into a frequency range where steep analog filters are far easier to implement cleanly, avoiding the phase distortion and ringing that aggressive brick-wall filters near the audible band can introduce. Modern converters typically use delta-sigma modulators oversampling at 64 times the target rate or higher, converting a small number of amplitude bits (often just one) into a noise-shaped, high-rate bitstream, then decimating back down in the digital domain. It is worth noticing that this is the same underlying trick as noise-shaped dither from the previous section, moving an unavoidable error into a frequency range the ear or the filter can more easily discard, applied at the front end of the signal chain instead of the back end.
The choice of digital reconstruction filter on the DAC side introduces its own trade-off. A steep, linear-phase brick-wall filter gives the flattest possible frequency response but produces symmetrical ringing before and after a sharp transient, called pre-ringing and post-ringing. Minimum-phase and apodizing filter designs trade a small amount of ultrasonic frequency response, and sometimes a gentler roll-off, for eliminating or reducing the pre-ringing, since pre-ringing (energy appearing before the transient that caused it) is generally considered the more audibly objectionable of the two. Filter choice is one of the genuinely contested areas of digital audio engineering, and it is an area where careful listening and careful measurement are both doing real, complementary work rather than one simply confirming the other.
Jitter
Jitter is a timing error rather than an amplitude error. Digital-to-analog and analog-to-digital conversion depend on a clock signal marking precisely even intervals; any deviation in the timing of those clock edges is jitter. Unlike quantization noise, jitter's audible effect is not a fixed noise floor but a modulation that tracks the signal, producing low-level distortion and, notably, smearing of stereo imaging and depth, since accurate interchannel timing is part of how the ear localizes sound.
Jitter is usually specified two ways: as a time figure, picoseconds RMS, describing how far a clock edge deviates from its ideal position, and as phase noise, expressed in dBc/Hz at a given frequency offset from the carrier, describing the spectral distribution of that timing error. The two views matter for different reasons. A single picoseconds-RMS number is easy to compare across datasheets, but it can hide whether the jitter is random (uncorrelated thermal noise in a PLL, which tends to spread energy broadly and is generally considered more benign) or deterministic (periodic, often traceable to power supply ripple or crosstalk from a switching regulator, which shows up as discrete sidebands around the signal and is generally considered more audible for a given RMS magnitude, because the ear is comparatively good at detecting periodic modulation).
Audibility thresholds for jitter are a genuinely debated area of the research literature, with published figures for detectable jitter ranging from tens of nanoseconds for random jitter down to considerably lower effective thresholds for deterministic, signal-correlated jitter, varying further with test methodology, program material, and converter architecture. This is a case where the field guide's usual confidence should be tempered: the mechanism by which jitter degrades sound is well understood, but pinning an exact, universal audibility threshold is not something the current research supports doing with precision. This is one of the reasons external master clocks and well-engineered internal clock circuits remain a live topic in high-end digital playback, even though bit-for-bit data transfer is, by definition, lossless. The clock is the one part of the digital chain where "lossless" data transfer does not guarantee a lossless result, because the reconstruction timing is not carried in the data itself. This site treats the point at length elsewhere: Jitter Is the New Wow and Flutter traces the mechanism back to its analog-era ancestor and its practical sources, Word Clock Impedance covers the transmission-line reasons a clock cable's impedance matters, and External Master Clocks and The Digital Hierarchy make the system-level case for treating clocking as a first-order design decision rather than an afterthought.
Inter-sample peaks
Inter-sample peaks (ISP) are a subtler problem than either of the above. A digital meter reads the level of each sample, but the actual analog waveform reconstructed by the DAC's output filter is a smooth curve running between those sample points, and that curve can peak higher than any individual sample value suggests. A track that measures at exactly 0 dBFS on every sample can still produce a reconstructed analog peak that exceeds full scale, causing clipping or distortion in the analog reconstruction filter or in downstream equipment, even though nothing in the digital domain appeared to clip.
The standard way to estimate this without actually building an analog reconstruction filter into a meter is oversampling the metering process itself, typically 4x, as specified in ITU-R BS.1770, the same recommendation that underlies most modern loudness metering. A 4x-oversampled meter interpolates points between the actual samples and checks those interpolated points for overs, giving a much closer estimate of the true reconstructed peak than sample-peak metering alone. This is why mastering engineers commonly leave headroom, often referred to as true-peak limiting, targeting a ceiling below 0 dBFS (commonly around -1 dBTP) rather than mastering right up to the digital maximum. Several streaming and broadcast delivery specifications enforce a true-peak ceiling of this general order for exactly this reason: lossy codecs used for streaming delivery are themselves prone to introducing additional overs during encoding, on top of whatever inter-sample peaks already exist in the source, so the margin has to absorb both effects.
Inter-sample peaks also interact with loudness normalization in a way worth naming explicitly, since the two are often discussed separately when they are really the same headroom problem seen from two angles. Apple Music targets -16 LUFS, Spotify -14 LUFS, and YouTube Music approximately -14 LUFS; when a normalization system applies positive gain to a track mastered quieter than its target, it is spending headroom the mastering engineer may already have reduced to the minimum, which is exactly the condition under which an existing inter-sample peak is most likely to become an audible over downstream. The Bit-Perfect Paradox covers this and the other technically legitimate reasons a bit-perfect stream can still sound different from one platform to the next.
Part Two: Vocal Production and Frequency Management
These issues concern how sound behaves acoustically, both at the microphone and in the room, well before anything reaches a converter.
Sibilance
Sibilance is the exaggerated emphasis of "s," "sh," "ch," and similar fricative consonants. The Frequency Spectrum locates its origin primarily in the High Mid band, 2 to 5 kHz, rather than the extreme treble as is often assumed; the wider engineering literature generally puts the core sibilant energy a little higher than that, commonly cited as 4 to 10 kHz, with the difference coming down to where a given source draws the line between the fricative's fundamental energy and its upper harmonic extension. Both framings describe the same physical event. Its causes are mechanical and acoustic: close-miked vocals with a presence-boosted condenser capsule, natural variation between voices, and breath and airflow interacting with the diaphragm during fricatives.
The exact center frequency varies meaningfully between performers, and this is worth taking seriously rather than treating as a rounding error. Vocal tract length and formant structure differ between voices, and shorter vocal tracts, typically though not universally correlating with higher-pitched voices, tend to push sibilant energy somewhat higher in frequency than longer vocal tracts do. A large-diaphragm condenser microphone with an intentional presence boost in the 8 to 12 kHz range, a common design choice for adding perceived clarity and "air," compounds the problem by boosting exactly the region where sibilant energy already concentrates. The practical implication is that there is no single correct de-essing frequency to dial in from memory; it has to be located by ear and by spectrum analysis on each specific performance, which is a case where measurement (a spectrogram showing where the energy actually sits) and listening (confirming that the identified band is what is actually fatiguing to hear) are doing genuinely different, complementary jobs.
A playback system with rising treble, a resonant tweeter, or aggressive filtering near the top of the audible band can make an already-sibilant recording noticeably more fatiguing. This is one of the clearer cases where a transparent, well-measuring system is not making a mixing decision "sound worse" so much as simply not hiding it.
Plosives
Plosives are the explosive bursts of air associated with consonants like "P," "B," and "T," striking the microphone diaphragm directly and producing a low-frequency thump or distortion unrelated to the actual pitch or level of the voice. It is worth being precise about what kind of event this is: it is not a tonal or harmonic phenomenon at all, but a burst of turbulent airflow overloading the diaphragm's mechanical excursion, generally concentrated below 200 Hz, which is why it reads on a meter as a thump rather than a pitched note.
They are managed at the source, primarily with a pop filter placed between singer and microphone, and often reinforced with a high-pass filter to remove any residual low-frequency energy that gets through. Pop filters themselves come in two common forms, a stretched nylon mesh and an open-cell metal or foam screen; both work by breaking up and dispersing the airflow before it reaches the capsule rather than blocking sound, which is why a well-designed pop filter does not audibly dull the vocal the way an overly aggressive high-pass filter can. Mic placement helps as well: angling the capsule slightly off-axis from the direct line of a vocalist's mouth, a few degrees rather than a few inches, redirects the worst of the airflow past the diaphragm rather than into it, without materially changing the vocal's on-axis tonal balance. The high-pass filter used as a backstop is typically set somewhere between 60 and 100 Hz, with a slope of 12 or 24 dB per octave; a steeper slope removes more residual thump but requires more care to avoid removing genuine low-frequency warmth from a naturally deep voice.
Masking
Masking is a psychoacoustic phenomenon in which a louder sound prevents a quieter sound in the same frequency region from being clearly heard, even though both are technically present in the mix. The underlying mechanism is the ear's division of the audible spectrum into critical bands, sometimes described using the Bark scale, roughly twenty-four bands spanning the range of human hearing. Within a single critical band, a loud signal raises the effective hearing threshold for everything else sharing that band, which is why masking is strongest between sounds that are close in frequency and considerably weaker between sounds in well-separated bands.
Masking also has a time dimension that a purely frequency-domain description misses. Simultaneous masking is the case described above, two sounds present at the same time. Temporal masking is different: a loud sound can mask a quieter sound that occurs shortly before it (backward or pre-masking, typically effective for only a few milliseconds, since it relies on the ear's processing lag) or shortly after it (forward or post-masking, which can remain effective for up to roughly 100 to 200 milliseconds as the ear's response decays). This is, incidentally, the same psychoacoustic principle that lossy codecs like MP3 exploit deliberately, allocating fewer bits to content the encoder predicts will be masked rather than genuinely audible.
A heavy guitar track occupying the same 1 to 4 kHz range as a lead vocal's intelligibility band can mask that vocal's clarity, requiring either level, EQ, or arrangement decisions to carve out separate frequency space for each element. Framed in critical-band terms, the goal of that kind of arranging or EQ decision is specifically to move the two competing sounds into different critical bands, not simply to move them to different Hz values that still happen to fall in the same band. The Frequency Spectrum covers this same overlap problem from the instrument side, including where kick drum competes with bass guitar and where guitar, piano, and vocals share territory.
Proximity effect
Proximity effect is the rise in low-frequency and bass response that occurs when a sound source moves very close to a directional microphone (cardioid, figure-eight, and similar patterns; omnidirectional microphones are largely immune). The mechanism is specific to how directional microphones work: they are pressure-gradient transducers, responding to the difference in pressure between the front and back of the diaphragm, whereas an omnidirectional microphone is a pure pressure transducer, responding to absolute pressure at one point and therefore having no gradient to exaggerate.
Close to the source, in what is called the near field, the pressure gradient across the diaphragm increases disproportionately at low frequencies as distance decreases, producing a boost that grows by roughly 6 dB per octave for each halving of distance, concentrated below a few hundred Hz. This is a predictable consequence of the microphone's polar pattern and is routinely used intentionally by vocalists and engineers to add warmth, the classic radio-announcer or crooner low-end fullness comes directly from this effect, but it can also become an unwanted, boomy coloration if mic distance is inconsistent through a performance, since the amount of bass boost changes continuously as the singer moves rather than staying fixed at one setting the way an EQ boost would.
Resonance
Resonance is the buildup of specific, disproportionately loud frequencies, caused either by room acoustics (standing waves reinforcing certain frequencies more than others) or by the natural mechanical characteristics of an instrument or voice. Room resonances of this kind are generally called room modes, and they dominate a room's low-frequency behavior up to a transition point, the Schroeder frequency, above which the density of overlapping modes becomes high enough that the room's behavior is better described statistically (as diffuse reverberant decay) than as a set of discrete resonant peaks. The Frequency Spectrum gives the working approximation, fs is roughly 2000 times the square root of reverberation time in seconds divided by room volume in cubic meters, and works a domestic example (roughly 60 cubic meters, 0.4 second decay) out to a Schroeder frequency near 165 Hz, with most domestic rooms falling somewhere between 150 and 400 Hz. Below that transition, individual modes tied to specific room dimensions can be identified, measured, and in a treated room, targeted directly.
Unlike broad tonal coloration, resonance tends to be narrow and peaky, a property usefully described by Q, the ratio of a resonance's center frequency to its bandwidth. A high-Q resonance is narrow and sharply peaked; a low-Q resonance is broad and gently sloped. This distinction is exactly why resonance is well suited to correction with dynamic EQ or notch filtering that targets only the offending band, matching a narrow cut's Q to the resonance's own Q, rather than static EQ moves that would also affect adjacent, unproblematic frequencies. A wide, low-Q static cut applied to fix a narrow, high-Q resonance removes far more of the wanted signal than necessary, which is the practical reason dynamic and notch tools are generally preferred here over a fixed shelf or bell.
Part Three: Dynamic Processing Techniques
These are the practical tools used to manage the acoustic issues above once they are captured.
De-essing
De-essing is a frequency-selective form of compression that listens specifically to the sibilant band, typically 4 to 10 kHz, and reduces gain only when energy in that range exceeds a threshold, leaving the rest of the vocal untouched. This differs meaningfully from a static EQ cut, which would dull consonants across the entire performance rather than only during problematic peaks.
De-essers generally work in one of two architectures. A wideband de-esser triggers a gain reduction on the entire signal whenever the filtered sidechain detects excess sibilant energy, effectively acting like a compressor keyed to a narrow band. A split-band de-esser instead separates the signal into two paths, compresses only the sibilant band, and recombines it with the untouched remainder, which generally produces a more surgical, less audible result at the cost of slightly more processing complexity. Attack and release times matter more here than in most compression contexts, because sibilant consonants are extremely short events, often well under 100 milliseconds; an attack time that is too slow lets the harshest part of the "s" through before gain reduction engages, while a release time that is too fast can cause audible pumping between consecutive sibilant sounds in rapid speech or lyrics.
Well-executed de-essing is inaudible as processing; over-applied de-essing produces its own signature, a lisping or muffled quality on affected words, essentially trading one artifact (harshness) for another (loss of consonant clarity and intelligibility).
Sidechain compression
Sidechain compression uses the level of one track to trigger compression on a different track. The most familiar example is ducking a bass guitar whenever the kick drum hits, briefly reducing the bass so the kick's transient cuts through, then releasing once the kick has passed. The technique depends on the same attack, release, and threshold parameters as ordinary compression, plus an additional lookahead option on many modern tools, which delays the main signal by a few milliseconds relative to the sidechain trigger so the gain reduction can begin fractionally before the triggering transient arrives rather than reacting after the fact.
The same technique applies to vocals over dense mix elements, and is functionally related to de-essing, which is itself a sidechain process where a track compresses against a filtered version of itself rather than against a separate track. Recognizing sidechain compression, ducking, and de-essing as three applications of the same underlying tool, rather than three unrelated processes, is a useful way to understand why they share the same core controls and the same failure modes: what studio floor talk calls pumping or breathing, audible volume surges or dips from a release set too fast, or sluggish, ineffective control from an attack set too slow.
Parallel compression
Parallel compression, sometimes called New York compression after the mixing and mastering rooms where the technique became closely associated with a particular thick, present vocal and drum sound, blends an unprocessed track with a heavily compressed copy of the same track. The compressed signal fills in low-level detail and adds density and punch, while the original signal preserves the natural transient shape and dynamic range of the performance. The result is a track that sounds louder and more present without the flattened, over-squeezed character that heavy compression produces on its own, because the ear is hearing the sum of an untouched transient and a dense, sustained body, rather than a single signal in which the transient itself has been squashed. Applied across a whole mix bus rather than a single track, the same principle is what studio vocabulary calls glue, a sense of cohesiveness achieved with gentle compression across every element at once.
Dynamic EQ as a bridge between the two domains
It is worth closing the processing section by naming a tool that sits deliberately between the acoustic problems in Part Two and the dynamic tools in Part Three: dynamic EQ. A dynamic EQ band behaves like a standard EQ filter, with a defined center frequency and Q, until the signal in that band crosses a threshold, at which point it applies gain reduction similarly to a compressor, then releases once the level drops back below threshold. This makes it well suited to problems that are level-dependent rather than constantly present, a resonance that only becomes objectionable on the loudest notes, or sibilance that is only harsh on the most forceful consonants, without the always-on tonal change a static EQ cut would apply to quieter, unproblematic material in the same band. It is, in effect, a general-purpose version of the same principle that makes de-essing preferable to static high-shelf cutting.
Where This All Connects
Every artifact and technique above shares a common thread: the ear is exceptionally sensitive to problems in the upper midrange and treble, and to timing and transient information, whether the source of the problem is mathematical (quantization, aliasing, jitter, inter-sample peaks) or acoustic (sibilance, plosives, masking, proximity effect, resonance). Careful engineering in both domains is largely inaudible when done correctly and immediately obvious when it is not.
For listeners auditioning equipment or cables, this also explains a common observation: an unusually revealing system will make source material decisions, dithering choices, clocking quality, and vocal processing alike, more apparent rather than less. That is not a flaw in the equipment. It is the equipment doing exactly what it should. A system with excess treble energy, poor jitter rejection, or a resonant, uncontrolled cabinet will not merely add its own coloration; it will interact with whatever is already present in the recording, sometimes masking a mixing decision, sometimes exaggerating one, in ways that make it genuinely harder to separate what the equipment is doing from what the recording already contains. This is precisely why measurement and listening need to be treated as complementary rather than competing sources of evidence: measurement identifies what a piece of equipment is objectively doing to the signal, and listening, done carefully and skeptically, is how that objective behavior is checked against what it actually does to program material as complex and as psychoacoustically loaded as the human voice. What You Hear, What You Measure develops this relationship in full, including where genuine measurable synergy exists in a system and where claimed synergy deserves proportional skepticism.