selfwyse

Are Personality Tests Accurate? What Decides It

Last updated

Accuracy is the wrong word for one thing and the right word for three. A test can be solid on two of them and hopeless on the third, which is how a result ends up feeling right and meaning nothing.

Accurate at what?

A thermometer is accurate because there is a real temperature to be right about. Personality has no equivalent: there is no true extraversion sitting in somebody waiting to be read off, so there is nothing for a number to match.

What can be asked instead is narrower and more answerable. Do the questions come from somewhere defensible. Do they hold together. Is there anything to compare an answer against. Those three are what separate a serious instrument from a quiz, and any of them can be checked before trusting a result.

One: where did the questions come from?

This is the cheapest thing to check and the most often skipped. A question set is either published, so that anybody can inspect it, or it is private and the reader is asked to take it on trust.

Public item pools exist precisely so this is answerable. Items that have been used and re-used by researchers for decades come with a paper trail: what they were written to measure, how they behaved, and what got dropped. A question written last month by a marketing team has none of that, and no way for anybody outside to tell the difference.

The word to look for is not scientific. It is the name of the item pool, and a test that will not say is answering the question by declining it.

Two: do the questions hold together?

A scale is a group of questions that are supposed to be getting at one thing. Whether they are is measurable, and the usual measure is how consistently somebody who agrees with one of them agrees with the rest.

This is where a great many quizzes quietly fail. Five questions that sound related to the person writing them can turn out to have almost nothing in common in the answers, in which case adding them up produces a number about nothing.

It is also the figure most often borrowed rather than earned. A reliability figure belongs to one specific set of questions. Quoting the figure from a published instrument while asking different questions is a claim about somebody else's work.

Three: is there anything to compare against?

This is the one that decides what a result can say, and it is almost never mentioned.

A score of 34 means nothing by itself. It means something once it can be placed among the answers of many other people who took the same questions. Then the position is real: a percentile says where somebody sits in a measured distribution rather than on an invented curve.

Without that sample there is still an honest reading available, and it is a smaller one. The result can describe how somebody answered, which of several things sounded most like them, or where they landed on the answer scale itself. What it cannot do is say where they sit relative to anybody, and a test doing that without a sample behind it is making the number up.

Most of what is sold as a personality test has no reference sample at all. That does not make those tests worthless. It makes the sentence about being in the top ten per cent worthless.

Why an accurate result can still feel wrong

A well-built instrument measures how somebody answered on the day they answered. That is a real fact and it is a narrow one, and there are several ordinary reasons the description can miss.

A hard week moves some scales genuinely. Answering as the person somebody is at work produces a different result from answering as the person they are at home. And self-report can only see what a person can see about themselves, which on some questions is not much.

The useful response to a result that does not fit is not to discard the instrument or to accept the label. It is to check whether the description is wrong or merely unwelcome, which only the reader can do.

Four questions worth asking any test

None of these needs a background in statistics, and any instrument worth taking answers all four in public.

  • Where did the questions come from, and can they be read?
  • What are the reliability figures for these questions, rather than for a similar instrument?
  • Is there a reference sample, and how many people are in it?
  • If there is no reference sample, does the result still claim to compare the reader with other people?

Common questions

Is the Big Five more accurate than other frameworks?
On all three counts above it has the strongest case: public items, published reliability figures, and reference samples large enough to place a score in a real distribution. That is a statement about the measurement rather than about which framework a person finds most useful.
Does a free test mean a worse test?
No. Price tracks who owns the questions rather than how good they are, and several of the most heavily studied item pools in psychology are free to use. What matters is whether the three questions above can be answered.
Why do two tests give different results?
Usually because they are measuring differently named things, or because their bands are drawn differently. A result near a boundary can move between them on the same answers, which is one reason a band is a better thing to read than a number.

Read more about this

Or see everything at Learn.

Find out where you land

Every assessment here is free to take and gives a real result, with the sourcing for it set out in public. The Big Five is the usual place to start.

50 questions · about 5 min · free

Start your free assessment