One pair of headphones has ever sounded real to me out of the box. Not impressive, not detailed, not good for the money. Real, as in instruments in a room with air around them.
Everything else, and there has been a great deal of everything else, had the same fake timbre. There is a second pair that gets there now, but I had to build it, and that is what this post is about.
The problem with this hobby
Most people are running an experiment with a sample size of five or six. You buy something well reviewed, live with it, like it or don’t, sell it and try the next one. There are maybe a thousand headphones you could reasonably own. Nobody hears more than a handful properly. So the whole hobby runs on luck wearing a lab coat.
The one that did it for me is a Sennheiser HDB 630. Closed, wireless, and nowhere near the most expensive thing I have listened to. First proper listen, something happened that had almost never happened before: the music stopped being a very good reproduction and became a performance in a space. Instruments had weight and body. It sounded like being somewhere.
Here’s what makes that worse than an ordinary hobby problem rather than better. The quality I’m describing isn’t on a spec sheet and isn’t visible in a frequency response graph, which is the only measurement most of us ever see. Reviewers don’t have a word for it that means the same thing twice. “Soundstage” gets used, but in headphone reviews that nearly always means wide inside your head, which is a different thing.
So when it happens you can’t tell anyone how to reproduce it. You can only say which model you bought. Someone reads that, buys the same model, doesn’t get it, because the effect depended on the combination rather than the box.
So what is it about that one? Nothing you could have read off it. It isn’t the most expensive thing I’ve heard or the most exotic, and closed backs costing several times as much did nothing. My first instinct was that the noise cancelling explains it, by removing the room cues that keep reminding your brain where you really are. That is probably part of it, but the Bose and Sony flagships seal and cancel too, and they did nothing.
The only other candidate I can find is that its response, arriving at my particular eardrums, happens to land near what my brain expects to hear. Which would be luck, and unrepeatable luck at that. Someone with different ears would have a different model on their own list, and neither of us could tell the other which.
That is a hypothesis, not a finding. One hit is not evidence of anything. But it is testable, and that is what makes the rest of this post possible: if the thing I’m chasing really is a match between a headphone’s response and my own ears, then it’s computable rather than something you gamble on in a shop. You measure the ears, and you correct a headphone you already own into the match. No purchase involved. The rest of this is me running that test, and then finding out what the tonal match still leaves missing.
I think a lot of people are listening on headphones and have never once heard music sound real, at any budget. Not because their gear is bad, but because the missing ingredient isn’t sold with the gear.
The bass I chased for a year that wasn’t bass
My main headphones are the Arya Unveiled, open back planars, and out of the box they were not doing it either. Next to the HDB 630 they sounded lean, and the obvious explanation was bass. The 630 seals, and level matched it really does have far more low end, around eight decibels more at 20 to 50 Hz. Sealed cups pressure load the driver in a way an open back physically cannot.
So I added bass. It didn’t work. A wide low shelf went boomy and closed in. A punchier version aimed at 75 Hz made everything soft rather than powerful. More of it made things worse, and I couldn’t work out why.
Then a measurement stopped me. I’d also admired the Meze Empyrean II for its bass in a shop, so I compared the two on the same rig, level matched.
Level matched on the same rig, they are identical in the sub bass to within a third of a decibel. Every bit of what I heard as better bass was a three decibel hump between 70 and 400 Hz, with a scoop above it. Warmth and body, not depth. And that band is exactly the one that turns thick and closed when you add it with an equaliser, which is why my attempts kept backfiring.
They’re identical in the sub bass to within a third of a decibel. What I heard as better bass was a three decibel hump at 70 to 400 Hz plus a scoop above it. Warmth and body, not depth. And that band is exactly what turns thick and closed when you add it with an EQ.
Want more bass? Drop the bass, and cheat your brain instead
What worked was the opposite of what I had been doing, and it has a name: the missing fundamental. Play the harmonics of a note without the note itself and people still hear the note, at the right pitch. It’s why a phone speaker that physically cannot produce anything below 200 Hz still lets you follow a bass line. Your auditory system reconstructs pitch from the spacing between harmonics, not from the lowest frequency present.
That’s exploitable. A 50 Hz note also puts energy at 100, 150 and 200 Hz. Boost that region and the note gets louder, more present and more clearly pitched, and the driver never moves any further than it already was. The cost comparison is what sells it: those frequencies sit where your hearing is around fifteen decibels more sensitive than at the fundamental, and where a planar diaphragm is nowhere near its excursion limit. Same perceived result, a fraction of the energy.
Here’s the whole layer, sitting on top of the diffuse field correction. Four filters, nothing above 300 Hz, so it never collides with the tonal correction underneath.
Filter 1: ON LS Fc 16 Hz Gain -8.0 dB BW Oct 0.70
Filter 2: ON LS Fc 50 Hz Gain +1.5 dB BW Oct 2.00
Filter 3: ON PK Fc 120 Hz Gain +3.0 dB BW Oct 1.00
Filter 4: ON PK Fc 210 Hz Gain +2.0 dB BW Oct 1.10
120 Hz is the second harmonic of a 60 Hz note, 210 the third. Those two carry the pitch. The 50 Hz shelf is just enough real sub to keep it honest rather than hollow.
The subsonic cut at 16 Hz does more than it looks. Rumble below 20 Hz is inaudible but still moves the diaphragm, stealing excursion from bass you can hear and adding intermodulation that smears it. Removing it makes the audible bass cleaner at no perceptual cost. The shape matters though. My first attempt was a shelf at 22 Hz spanning two octaves, which took 1.9 dB out of 40 Hz as collateral and defeated the point. At 16 Hz and 0.7 octaves it removes 2.3 dB at 20 Hz and 0.19 dB at 40. If your EQ has a proper high pass, use that instead.
Two other approaches got built and rejected, so you don’t have to repeat them. A punch layer at 75 Hz, where a kick drum’s body lives, asked an open planar for felt impact it can’t produce without a seal, so the energy went where the ear is least sensitive and everything went soft. A clarity layer with no boost at all, cutting lower mids and presence so bass emerged by unmasking, was elegant on paper, free in headroom, and thin in practice.
Five decibels across two harmonic bands beat every attempt to add eight or nine decibels of actual bass. The version that adds least sounds like it adds most.
Honest caveat: the harmonic layer runs slightly heavy on cello, double bass and low brass, instruments whose real character lives in the boosted band. A trimmed version at 2.2 and 1.5 dB with the sub shelf at 0.5 exists, and it’s the better choice once the binaural chain further down starts supplying bass of its own.
All of which helped. None of it was the HDB 630 effect.
What was actually missing
Here’s the bit I had wrong, and I suspect most people do. In a room, sound reaches your eardrums after being filtered by your head, your shoulders and the specific folds of your outer ears. That filtering changes with direction, and your brain has spent your whole life learning yours. It’s how you know a sound is in front of you, above you, or three metres away, with your eyes shut.
Ordinary headphones deliver each channel straight to one ear with none of that. The result is a sound field that exists nowhere in nature, so your brain files it as happening inside your skull. You can make that pleasant. Smooth, weighty, detailed, beautifully balanced. You cannot make it real, because realness isn’t a tonal property. No EQ reaches it, no DAC reaches it, and no amount of money reaches it.
Which explains the 630. No magic drivers. It seals and it cancels noise, so my brain lost the ambient cues that would otherwise keep insisting I was sitting in a room wearing headphones. It got closer to real by removing the evidence against it.
Getting measured
The Audio Experience Design group at Imperial College London measures individual head related transfer functions as part of the SONICOM research project. You sit in a rotating chair in a semi anechoic room with tiny microphones at your ears while an arc of twenty three loudspeakers plays sweeps from every direction. Head position is monitored and it beeps if you drift off axis. Half an hour, plus a 3D scan of head and torso.
You get back a file describing how sound from 793 directions arrives at your two ears. That’s the unrepeatable half of everything here, and once you have it you never need it again. Ears don’t change.
They also did something extra that turned out to matter more than the scan. Julie at the lab kindly measured my own headphones on my own head, capturing how the Arya actually couples to my ears rather than to a measurement rig. That’s rarer than the HRTF itself, and it’s what makes a correction real rather than approximate. If you go, ask, and bring the headphones you actually use.
I’d never have known the lab existed if the people at The Headphone Show hadn’t visited and written about it. Credit where it’s due. They took their files toward building a personal measurement capability. I went a different way with mine.
Why Bose measuring your ears isn’t the same thing
There are two separate corrections available from this data, and confusing them wastes months.
One is a headphone correction. It measures how your headphone behaves at your eardrum and cancels the error. Genuinely useful: it removes peaks and resonances that follow every track, cuts fatigue, and catches things no published measurement can because it’s your fit rather than a rig’s. On its own it sounds technically correct and slightly closed in.
The other is a diffuse field correction. Instead of aiming at flat, it aims at the response your own ears produce for sound arriving from everywhere at once. That’s the shaping your brain expects, and restoring it is what makes the presentation open rather than shut. It needs the full directional file, not just a measurement of the headphone.
The Bose QC Ultra genuinely personalises. CustomTune plays a tone at startup, measures your canal through the internal mic, adapts its EQ. That’s the first correction, done properly, at consumer scale. It cannot be the second, because you can’t learn how sound from 45 degrees left is shaped by your pinna from a tone played inside a sealed cup.
Personalisation isn’t the missing piece. Bose personalises seriously and still doesn’t get there. The half that matters needs an anechoic room and an arc of speakers, and nobody does that at the point of sale.
Building the correction, and where to overrule it
Deriving the target is straightforward arithmetic. Take the magnitude response for all 793 directions and average them. Two details matter: the directions aren’t evenly spread over the sphere, so each needs weighting by the solid angle it represents or the top of your head gets counted several times over; and the averaging must happen on power rather than complex values, or opposing phases cancel and you invent nulls no listener experiences.
The correction is then your diffuse field target minus your measured headphone response. Both were captured in the same session with the same mics at the same points, so the microphone’s own colouration is in both terms and drops out.
Then you overrule the result in three places, and this is where blindly automating it falls over.
Above about 7 kHz, stop trusting it. The lab measured each headphone five separate times, taking it off and putting it back on between each. Below 5 kHz the five agree to within half a decibel. Above 7 kHz they scatter by two to four. Any correction finer than that is fitting how the headphone happened to sit that day. My raw maths wanted 11.5 dB of boost at 7 to 9 kHz because it was inverting a deep measured notch. I capped it at 4.
That turned out to be right for a reason I only found later. The notch was minus 22 dB at 9 kHz, and a null that deep needs a cavity in front of the mic. It was an artifact of how the probes sat in my ears, not a property of the headphone.
Below about 150 Hz the data has nothing to say. Neither the room nor the in ear method is reliable down there, so that region is a preference decision, not a measurement.
Every preset gets the same preamp. Boosts need headroom, so the chain is attenuated to make room. Use a different value per preset and you can’t compare two of them, because the louder one wins every time.
One check before trusting any of this. Your curve should have the landmarks a human ear has: a broad gain from the concha somewhere between 2 and 5 kHz, and a notch from the pinna somewhere between 6 and 12 kHz. Mine came out at +7.2 dB centred on 4393 Hz and -8.0 dB at 9520 Hz, both comfortably inside the published ranges. If yours has no such features, something upstream is broken and no amount of listening will tell you which part.
eqMac, and living with it
A personal correction is worthless if you can’t run it on everything. On macOS I use eqMac: free, installs a virtual audio device the whole system plays through. Expert mode gives unlimited parametric bands with frequency, gain and bandwidth typed in directly, plus a text format you can paste presets into, which is what lets a script generate them. On Windows, Equalizer APO does the same job, also free. Neither needs a subscription, a plugin host or a particular player.
Four settings decide whether this is pleasant or a daily annoyance.
Leave the virtual device selected permanently. Picking a real output in the macOS sound menu bypasses eqMac entirely and hands you raw uncorrected level. Switch headphones inside eqMac’s own output settings instead and the correction stays in the path.
System volume at 100 percent, level set on the DAC knob. This one cost me an evening. eqMac keeps its own master gain and reapplies it whenever its engine restarts, which happens on a device switch, a sample rate change, or adding an effect. My volume kept snapping back to seemingly random values until I understood why. At 100 percent the restore does nothing. At anything less it reasserts itself forever.
One preamp value across every preset. Mine is -7 dB on all thirteen. Otherwise a preset sounds better or worse purely because it’s louder and every comparison is meaningless.
Turn the DAC down before opening the preset editor. Real bug here: deleting one of your own presets makes eqMac select a built in default sitting at 0 dB, so the full 7 dB of preamp comes back instantly. On sensitive headphones at volume that’s unpleasant and arguably a hearing risk. The built in presets refuse to be edited, so you can’t give them the right preamp. I reported it to Roman, eqMac’s sole developer, in August 2026. Until it’s fixed, manage presets by script rather than in the app, since they live in an ordinary preferences file you can read and rewrite while the app is closed.
DON’T BE FOOLED BY THIS ONE
eqMac has a feature called Spatial Audio, and it is not what you want. Reading the application bundle shows it’s Apple’s algorithmic reverb, the Cathedral and Large Hall type presets, behind a friendlier name. No HRTF, no crossfeed, no positional processing at all.
It can still nudge the image out of your head, because reverb supplies a direct to reverberant cue and decorrelates the channels slightly. Real effect, and a very small dose can help. But it is not personalised, it is not what the next section describes, and switching it on expecting the real thing is how you conclude the whole approach doesn’t work.
Getting the sound outside your head
Everything so far fixes tone. It doesn’t put anything outside your head, because a stereo linked EQ can’t create the cues that locate sound. Those cues are differences between your two ears in time, level and spectrum, and they’re exactly what the 793 direction file describes.
So the last step is to stop sending the left channel to the left ear, and instead render it as though a loudspeaker stood 45 degrees to your left, filtered by your own ear. Same for the right. SPARTA Binauraliser, free from Aalto University, does precisely this and reads the lab’s file format directly.
TWO MISTAKES THAT WILL COST YOU AN EVENING
Use the file that still has timing in it. Research datasets usually ship a version with interaural time differences removed, which is correct for tonal averaging and useless for placement. Mine has 646 microseconds of delay at hard left and right, the textbook human figure. The other version is effectively zero. Feed a renderer the wrong one and you get colouration with no position.
Don’t apply your ears twice. These plugins have a diffuse field EQ switch. Leave it on and the plugin strips the average response out, so your diffuse field preset puts it back and the two are complementary. Turn it off and the plugin supplies everything, so the headphone should be corrected to flat instead. Either pairing is right. Mixing them applies your own pinna response twice.
Textbooks say 30 degrees, the standard stereo triangle. I preferred 45 by a clear margin, and there’s a measurable reason why.
Why the bass finally came right
Once the binaural rendering was running, the bass I’d spent a year failing to get with an EQ simply arrived. Texture, weight, instruments that sound like objects. The reason isn’t mysterious and isn’t psychological. When two virtual sources feed both ears, centre panned content reaches each ear twice and the copies add.
| Band | Left ear | Right ear |
|---|---|---|
| 20 to 40 Hz | +3.0 | +2.9 |
| 80 to 160 Hz | +2.8 | +2.4 |
| 320 to 640 Hz | +0.8 | -2.9 |
| 640 to 1250 Hz | -3.1 | -3.8 |
| above 2500 Hz | ~0 | ~0 |
Extra level that centre panned content gets, purely from the two virtual sources summing at each ear. From my own measured file at ±45 degrees.
Below roughly 300 Hz the two virtual speakers arrive at each ear essentially in phase, because the wavelength is far longer than the path length difference. They add coherently, which is six decibels of amplitude against three for uncorrelated material. Net gain three decibels. Higher up the phases diverge and you get partial cancellation instead.
Bass in music is almost always panned centre. So the chain hands you about three decibels of bass together with a three decibel scoop in the lower mids, applied only to the content that benefits. That’s a six decibel tilt between weight and thickness, and it’s precisely the shape I’d been failing to get with filters for a year. An EQ can’t do it, because a filter works on frequency and cannot tell where in the stereo image a sound sits. Every bass shelf I tried lifted the thickening at 100 to 400 Hz along with the weight, which is why each one traded warmth against openness.
Angle matters here too. At 30 degrees the lower mid scoop is barely present. At 45 it’s the full three decibels. Part of why I preferred 45 by ear is that it separates weight from thickness more cleanly.
A month ago I’d have told you the Arya couldn’t do bass like a sealed headphone. It can. It needed a spatial fix, not a tonal one, and every tonal attempt was compensation for a cue that was missing entirely.
Does it survive a week
Start with the hypothesis from the beginning of this post, because it now has an answer. There are two pairs on my list that sound real. The first was luck. The second is the Arya, which did not do it out of the box and does it now, and got there by measurement rather than by shopping. That is the whole claim.
Novelty flatters, so the only honest test is time. After a week of daily listening across everything from Tool to Vivaldi, it holds. Turning the chain off now makes music sound flat, which is the signature of a reference that has genuinely moved rather than a first impression fading.
Then I took a work call with it running, which I hadn’t planned as a test and which turned out to be the best evidence in the whole exercise. Voices sounded like the actual people. Timbre close enough to real that I stopped noticing the technology. Voices are the sound we all have the most lifelong reference for, so the brain is merciless about anything slightly off.
My one complaint was that voices had slightly too much bass. Which is exactly what the table above predicts, because a mono voice call is maximally centre panned and gets the full three decibels of coherent summation. The arithmetic forecast my own criticism before I made it. That’s about as good as subjective listening confirmation gets.
Doing the maths without being a signal processing engineer
Worth saying plainly, because this is the part that stops most people: I don’t write Python, and I did not learn it for this. The steps in the last few sections are fiddly rather than hard, but fiddly is enough. Reading a research file format, weighting 793 directions by solid angle, fitting an arbitrary curve onto a handful of biquads without the optimiser doing something stupid. Some of it I could follow but not implement, and some of it I could not follow at all until it was explained to me twice.
An AI assistant covers that gap completely. You describe what you are trying to achieve, it writes the analysis and draws the graphs, and you go back and forth until the result makes sense. It doesn’t much matter which one. Claude, Codex, whatever you already pay for. The work is ordinary enough that any capable model handles it.
What it cannot do is tell you when the data is lying. A measurement artifact and a genuine headphone flaw look identical in a spreadsheet and sound nothing alike, and more than one confident, well argued correction only fell apart when I put it on and listened. So it proposes and you judge, every time. That division held for the whole project, and it is the only part of the method I would insist on.
The recipe, if you want to try it
None of the software costs anything. The measurement is the only real barrier.
- Get your ears measured. Look for a research group running HRTF measurements. SONICOM at Imperial College London did mine. University acoustics departments are the place to ask. Bring the headphones you actually listen to, and ask whether they’ll measure those on your head too. That second measurement is what turns an approximation into a correction.
- Build the diffuse field target. Average the magnitude response over every measured direction, weighted by solid angle, averaged on power. That curve is yours for life.
- Derive the headphone correction. Personal target minus measured headphone response. Then overrule it: hold bass flat below 150 Hz, cap treble correction above 7 kHz at about 4 dB, and normalise so a chosen midrange band sits at zero, which keeps every preset you ever build level matched to every other.
- Fit it to parametric filters. eqMac on macOS, Equalizer APO on Windows. Two traps: bound each filter to its own frequency band or the fitter parks two filters on the same spot with opposite gains, and never let a shelf go narrower than about 1.9 octaves or it overshoots at its corner and sounds boomy.
- Set one preamp value and never change it. It has to absorb the largest boost in any preset. Mine is -7 dB. Same number everywhere, so switching presets changes tone and nothing else.
- Add the binaural renderer. SPARTA Binauraliser in the same chain. Load the file that retains interaural time differences. Two inputs at ±45 degrees, elevation zero. Check the diffuse field switch pairing above.
- Then use your ears. Every decision that mattered came from listening, not graphs. Graphs were most useful for ruling things out. If the image sits off centre, nudge the virtual speaker angles rather than reaching for a balance control. If it stays inside your head, add a very small amount of short reverb, because the brain wants some reflected energy before it believes in a room.
The upgrade I didn’t buy
Partway through I was seriously considering the HiFiMan HE1000 Unveiled, roughly double the Arya. Same driver, same pads, confirmed by the manufacturer. Measured on the same rig, the entire difference is one to three decibels of tuning above 1.6 kHz, and every bit of it transfers to the Arya as an EQ preset. I built it, listened for a week, and kept it. The money would have bought metal, leather and resale value. All real, none of it sound.
Same conclusion on the Empyrean II, and on whether a newer closed back would get me there. Tonality is correctable. The thing I actually wanted wasn’t for sale at all.
WHERE I COULD BE WRONG
One person’s ears over one week, not a controlled study. My measurements are as good as a research lab makes them. My listening is as biased as anyone’s.
Documented practice says anechoic binaural rendering sounds unnatural on ordinary stereo music. I found the opposite, strongly. Candidates: my file is measured rather than simulated from a scan, 45 degrees suits stereo mixes better than the 30 most guides specify, and a personal diffuse field correction sits underneath it. I can’t separate those yet.
And the correction has limits. Above 7 kHz nothing is trustworthy to better than about 3 dB, because that’s how much it moves when you take the headphones off and put them back on.
Everything here is over-ears. I have barely tried IEMs, only AirPods Pro 2 and a recent Sennheiser true wireless, both flat in the sense this post means. That isn’t a fair test and I won’t claim it is. The theory does make a prediction worth someone checking, though: an IEM sits past the pinna entirely, so it cannot deliver the direction dependent filtering at all, and on this account it should be the hardest case rather than the easiest.
The part that bothers me
The industry has spent decades optimising tonality. Targets, measurement rigs, driver materials, distortion figures. Real engineering, all of it aimed at the half of the problem a transducer can solve.
The other half needs one measurement of your own ears and then software that costs nothing. It isn’t exotic. Researchers have understood it for years. It just doesn’t fit into a product you can buy in a shop, because the measurement can’t be taken in a shop.
So the experience I spent years and several purchases chasing was never available for money. It needed half an hour in a chair in London, someone in a lab willing to measure my headphones as well as my ears, and a few evenings of arithmetic.
If you’ve ever felt headphones sound impressive but never quite real, I don’t think you’re imagining it, and I don’t think another purchase fixes it. If anyone here has done the same thing with a different dataset or a different renderer, I’d genuinely like to compare notes, particularly on the 30 versus 45 degree question.
THE SHORT VERSION
- A research group at Imperial College London measured my personal head related transfer function, and kindly measured my own headphones on my own head at the same time. It took half an hour and cost nothing.
- From that I built a personal diffuse field EQ and ran it in eqMac. Then I added binaural rendering with SPARTA, a free plugin from Aalto University.
- The result is the first time headphones have sounded like real instruments in a real space to me, including on a work call where voices sounded like the actual people.
- A bass problem I’d been failing to solve with EQ for a year turned out to be a spatial problem, and the chain fixed it as a side effect. I can show you the arithmetic.
- Every filter value and every setting is in the post above. All the software is free. The measurement is the only barrier, and it’s a barrier of access rather than money.
For whatever it is worth as context: over the years the HD 650, 800 S and 820, the Susvara and HE1000 Unveiled, Focals, Mezes, ZMFs, Dan Clark closed backs and a few electrostatics have all been on my head, mostly at MP3Store in Wrocław, across sources from budget chains upward. One of them did this out of the box.
Gear: HiFiMan Arya Unveiled, open planar, personally corrected. Sennheiser HDB 630, closed wireless. eqMac and SPARTA Binauraliser on macOS. Personal HRTF from SONICOM at Imperial College London.
Thanks to the Audio Experience Design group at the Dyson School of Design Engineering, Imperial College London, and to the SONICOM project. Particular thanks to Julie for measuring my own headphones alongside the standard session, which is the measurement everything here depends on. Thanks also to The Headphone Show, whose visit to the lab is the only reason I knew any of this was possible.
Measurements: personal HRTF and headphone transfer functions, SONICOM database subject P0408. Third party headphone curves from Kuulokenurkka, all same rig and pinna generation. Software: SPARTA Binauraliser (Aalto University), eqMac, REW, and a lot of Python.

