Is the 16 personalities test reliable?

The honest answer: it sometimes lands, for the wrong reasons. And on the three criteria of a real test, it doesn't hold up. Here they are, without contempt.

The doubt that comes afterwards

You took the 16 personalities test. Probably on the best-known site, the one that gives you a detailed result, nicely written, with colors and a little character. And it was uncannily accurate, in places. You recognized yourself. You may have read the description out loud to a friend saying "that's exactly me." You remembered your four letters, you brought them up in conversations.

Then a grain of sand. An article you came across somewhere, a colleague's remark, or a second test that gave you a different result. And the question settles in, and won't leave: can I really trust this thing? Is it serious, or is it an upgraded horoscope?

The detail to put down right away: your question is the right one. It isn't naive to ask it — it's exactly what should be asked of any test, on any subject. And to answer it, you don't need one more opinion. You need criteria.

If you're wondering whether the 16 personalities test is reliable, here's what we're going to do: no "it's rubbish," no "it's brilliant." We're going to lay out the three criteria that define a reliable test, and look honestly at how the MBTI does on each one.

What a reliable test has to do

Three criteria. They apply to any measuring instrument, from a thermometer to a blood test. These aren't my personal demands — they're everyone's, the moment measuring something is involved.

1. Reliability: is the result stable?

A good test gives you the same result if you retake it two weeks later, all else being equal. That's the baseline. A thermometer that reads 66°F and then 78°F in the same room, at the same hour, isn't a thermometer: it's a number generator. A personality test is no different: if it changes while you haven't changed, it isn't measuring what it claims to measure.

2. Validity: does it measure what it says it measures?

A test can be perfectly stable and measure the wrong thing. Validity is the question: does what the test calls "intuition" or "judging" actually correspond to something real, and distinct? Do its categories hold up when you put them against the facts, or do they overlap, blur into each other, cut reality in arbitrary places?

3. Predictive power: what is it good for?

A useful test predicts something about the real world — a behavior, a preference, a performance, an outcome. Without that, it's a description going in circles: pleasant to read, but with no grip on anything. If knowing your type lets you foresee absolutely nothing, what's the point of knowing it?

Keep those three in mind. Everything that follows is running the MBTI past them, one by one.

The MBTI against the criteria

Reliability: it fails.

This is the most documented point, and the most embarrassing one for the tool. Retake it a few weeks apart and there's a good chance it files you somewhere else. The type it gave you the first time commits to nothing.

And that isn't chance: it's the direct consequence of cutting a continuous reality into categories. Near a border, the smallest difference in an answer flips the whole letter — I explain that mechanism in detail here. Verdict on this criterion: on stability, it doesn't hold up. And a test that fails on reliability fails at the first hurdle, because nothing it says afterwards can be trusted.

Validity: partial — and this is where honesty is required.

Here I'm not going to load the case against the tool out of excess severity, because that would be dishonest. Some MBTI dimensions do overlap with real traits. The introversion/extraversion axis in particular corresponds to something real, found as such in the serious models of personality. The MBTI didn't invent everything from nothing.

But it suffers from two major defects. First, it cuts into categories what is gradual — it tells you "E" or "I" where reality is a slider. Second, some of its dimensions don't separate cleanly under analysis: they overlap, or measure partly the same thing under two names. Verdict: it captures something real, but cuts it badly.

Predictive power: weak.

The MBTI type predicts poorly the concrete things a personality test ought to shed light on. It describes — sometimes beautifully — but it doesn't predict much that's verifiable. Verdict: it's more a flattering mirror than a measuring instrument.

So why is it so appealing, despite all that?

Because of the Barnum effect, whose mechanism I take apart elsewhere: a portrait written so that nobody is left out of it. And here's what really needs understanding: what you take for accuracy isn't proof that the test pinned you down. It's proof that it's written so that everyone recognizes themselves. Its persuasive power and its scientific weakness come from exactly the same source.

The MBTI isn't wrong everywhere. It's appealing for the exact reasons that make it unreliable: it flatters, it simplifies, it files — and a good mirror is not a good measuring instrument.

So we throw it all out? No

The right conclusion isn't "none of this is worth anything," nor "knowing yourself through a test is an illusion." That would be throwing the baby out with the bathwater.

The right conclusion is that there is a way of measuring personality that passes all three criteria where the MBTI stumbles. That model is the Big Five.

Reliability: it measures in degrees, not in boxes. So no flipping near borders, no type changing from one test to the next. The result is stable over time, because there's no line to cross.

Validity: its five dimensions weren't decided in advance by an author. They came out of the analysis of language, and were confirmed across cultures. They describe things that are real, and distinct from one another.

Predictive power: it's the model research actually uses — precisely because it predicts verifiable things, which is the only reason researchers would bother with it over another.

The price to pay, in all honesty: it's less fun. No type name to remember, no tribe to join, nothing that reads well in a profile. Five sliders is drier than "INFJ, the Advocate." But that's exactly the price of reliability: an accurate instrument is rarely as appealing as a good label.

The MBTI gives you a character. The Big Five gives you a measurement. One is more enjoyable; the other is true.

What that changes for you

Keep what the MBTI gave you, without relying on it. If it gave you the taste for understanding yourself, for asking questions about how you work — genuinely, so much the better. It's a good way in. But a door isn't a house: don't build your self-knowledge on a tool that fails the first criterion, reliability.

Be wary of your own sense of accuracy. "That's so me" isn't proof of reliability. It's very often the sign of a text vague enough to suit almost anyone. Your conviction isn't data about the test — it's data about how the test is written.

If you want a result you can trust, change the kind of tool, not the test. The problem isn't which box-based test you pick among the dozens out there. It's the principle of boxes itself. You can run through four-letter tests your whole life: they'll all share the same design flaw.

Serious, or upgraded horoscope?

Come back to the question you started with.

The verdict, straight: between the two — and closer to the horoscope than anyone wants to admit. Not through any malice on its creators' part, but by construction. The MBTI sometimes lands, for the wrong reasons, and it fails on the criteria that define a reliable test: it isn't stable, it cuts badly what it measures, and it predicts almost nothing. That isn't nothing — it captures a share of the real, it opens doors. It simply isn't a measuring instrument.

A reliable test is stable, measures what it claims to measure, and predicts something real. The Big Five ticks those three boxes where the MBTI stumbles — at the price of being far less spectacular. You still have to land on a test that measures it properly: the criteria are here.

If you want a profile you can trust, rather than a character to identify with, the test is here. It takes ten minutes.