The doubt that comes afterwards
You took the 16 personalities test. Probably on the best-known site, the one that gives you a detailed result, nicely written, with colors and a little character. And it was uncannily accurate, in places. You recognized yourself. You may have read the description out loud to a friend saying "that's exactly me." You remembered your four letters, you brought them up in conversations.
Then a grain of sand. An article you came across somewhere, a colleague's remark, or a second test that gave you a different result. And the question settles in, and won't leave: can I really trust this thing? Is it serious, or is it an upgraded horoscope?
The detail to put down right away: your question is the right one. It isn't naive to ask it — it's exactly what should be asked of any test, on any subject. And to answer it, you don't need one more opinion. You need criteria.
If you're wondering whether the 16 personalities test is reliable, here's what we're going to do: no "it's rubbish," no "it's brilliant." We're going to lay out the three criteria that define a reliable test, and look honestly at how the MBTI does on each one.
What a reliable test has to do
Three criteria. They apply to any measuring instrument, from a thermometer to a blood test. These aren't my personal demands — they're everyone's, the moment measuring something is involved.
1. Reliability: is the result stable?
A good test gives you the same result if you retake it two weeks later, all else being equal. That's the baseline. A thermometer that reads 66°F and then 78°F in the same room, at the same hour, isn't a thermometer: it's a number generator. A personality test is no different: if it changes while you haven't changed, it isn't measuring what it claims to measure.
2. Validity: does it measure what it says it measures?
A test can be perfectly stable and measure the wrong thing. Validity is the question: does what the test calls "intuition" or "judging" actually correspond to something real, and distinct? Do its categories hold up when you put them against the facts, or do they overlap, blur into each other, cut reality in arbitrary places?
3. Predictive power: what is it good for?
A useful test predicts something about the real world — a behavior, a preference, a performance, an outcome. Without that, it's a description going in circles: pleasant to read, but with no grip on anything. If knowing your type lets you foresee absolutely nothing, what's the point of knowing it?
Keep those three in mind. Everything that follows is running the MBTI past them, one by one.
The MBTI against the criteria
Reliability: it fails.
This is the most documented point, and the most embarrassing one for the tool. Retake it a few weeks apart and there's a good chance it files you somewhere else. The type it gave you the first time commits to nothing.
And that isn't chance: it's the direct consequence of cutting a continuous reality into categories. Near a border, the smallest difference in an answer flips the whole letter — I explain that mechanism in detail here. Verdict on this criterion: on stability, it doesn't hold up. And a test that fails on reliability fails at the first hurdle, because nothing it says afterwards can be trusted.
Validity: partial — and this is where honesty is required.
Here I'm not going to load the case against the tool out of excess severity, because that would be dishonest. Some MBTI dimensions do overlap with real traits. The introversion/extraversion axis in particular corresponds to something real, found as such in the serious models of personality. The MBTI didn't invent everything from nothing.
But it suffers from two major defects. First, it cuts into categories what is gradual — it tells you "E" or "I" where reality is a slider. Second, some of its dimensions don't separate cleanly under analysis: they overlap, or measure partly the same thing under two names. Verdict: it captures something real, but cuts it badly.
Predictive power: weak.
The MBTI type predicts poorly the concrete things a personality test ought to shed light on. It describes — sometimes beautifully — but it doesn't predict much that's verifiable. Verdict: it's more a flattering mirror than a measuring instrument.
So why is it so appealing, despite all that?
Because of the Barnum effect, whose mechanism I take apart elsewhere: a portrait written so that nobody is left out of it. And here's what really needs understanding: what you take for accuracy isn't proof that the test pinned you down. It's proof that it's written so that everyone recognizes themselves. Its persuasive power and its scientific weakness come from exactly the same source.
The MBTI isn't wrong everywhere. It's appealing for the exact reasons that make it unreliable: it flatters, it simplifies, it files — and a good mirror is not a good measuring instrument.