Methodology

Unflatter is a thin, honest layer over a public-domain research instrument. Here is exactly what it asks, how it turns your answers into percentiles, and what those percentiles are worth.

The instrument

The questionnaire is the 50-item set of IPIP Big-Five Factor Markers, developed by Lewis R. Goldberg and distributed through the International Personality Item Pool. The items are in the public domain: anyone may use them, including commercially, without a licence fee. That is unusual for a personality instrument and it is the main reason this test can be free.

The items were written as short public-domain analogues of the commercial Big Five markers, and they measure the five broad factors that decades of factor-analytic work keep recovering from descriptions of personality: openness to experience, conscientiousness, extraversion, agreeableness and neuroticism. The Big Five is the model most academic personality research uses, in contrast to the type indicators that sort people into categories the data does not support.

  • 50 statements in total, 10 per trait.
  • 18 of them are reverse-keyed, so agreeing with everything cannot produce a flattering profile.
  • Items keep their original interleaved order, which stops the questionnaire reading as five themed blocks and nudging you toward consistency.

How the test is administered

One statement per screen, answered on a five-point agreement scale: Strongly disagree through Strongly agree, with a neutral midpoint. There is no time limit and no forced choice; the number keys work if you prefer speed.

Progress is saved to your own browser as you go, so closing the tab does not lose your answers. Nothing is transmitted while you answer — the items, the scoring and the norms are all in the page, and the scoring runs on your device. See the privacy policy for what that means in practice.

From answers to a raw score

Each response is scored 1 to 5 as it stands, or reversed as 6 minus the response for reverse-keyed items. The ten items belonging to a trait are then summed, giving a raw score between 10 and 50.

raw(trait) = Σ over its 10 items of (keyed ? response : 6 − response)

Items you skipped are imputed at the neutral midpoint of 3, so a partly finished questionnaire still produces a sensible score instead of collapsing toward the floor. The more you skip, the harder your scores are pulled toward the middle — which is the correct direction for missing information, but it is a good reason to answer all fifty.

From a raw score to a percentile

A raw sum on its own means nothing: 38 out of 50 on agreeableness is close to average, while 38 on extraversion is comparatively high. To make the numbers comparable, each raw score is converted into a percentile against a reference distribution.

  • The raw score is turned into a z-score: z = (raw − mean) / SD, using the trait means and standard deviations in the table below.
  • The z-score is passed through the cumulative distribution function of the normal distribution, which is evaluated with a standard numerical approximation accurate to about 7.5 × 10⁻⁸ — far finer than a whole percentile point.
  • The result is multiplied by 100, rounded, and clamped to the range 1–99, because claiming someone is at the 100th percentile of anything is not a claim this instrument can support.

A percentile of 72 on conscientiousness therefore means: assuming trait scores are approximately normally distributed in the reference population, roughly 72% of that population scored lower than you did.

The assumption of normality is a simplification. Real trait distributions are close to normal but not perfectly so, particularly in the tails, which is one reason to treat extreme percentiles as “high” rather than as a precise rank.

The norms

Percentiles are computed against published means and standard deviations for the IPIP Big-Five markers on the 10–50 sum scale, drawn from large community samples of predominantly English-speaking adult respondents.

TraitMeanSD
Openness37.06.3
Conscientiousness33.56.9
Extraversion30.18.0
Agreeableness37.26.4
Emotional volatility28.38.1

These samples are self-selected internet respondents rather than a representative census of any country. They skew younger, more educated and more Western than the world does. If your background is far from that, your percentile is still a fair comparison against the sample — it is just a comparison against that sample, and not against your neighbours.

Bands and how to read a score

For the written interpretation, percentiles are grouped into three bands: low at 30 and below, balanced between 31 and 69, and high at 70 and above. The bands exist so the interpretation can say something useful; the underlying number is continuous and the boundaries are conventions, not cliffs. Someone at 69 and someone at 71 are the same person twice.

  • Read a percentile as a neighbourhood, not a coordinate. 54 and 61 mean the same thing.
  • Differences of fewer than about ten percentile points are within the noise of a self-report instrument taken once.
  • Retesting after a few weeks typically moves scores slightly. Large moves usually mean your circumstances or your mood changed, not your personality.
  • No score is good or bad. Each pole has costs and advantages, which is how the interpretations are written.

How the deep report is generated

The paid report is assembled deterministically from your five scores. Rules match on combinations of bands — high conscientiousness with high neuroticism, high openness with low conscientiousness, and so on — and the matching sections are selected, ordered and rendered with your numbers in them.

It is not written by a person for you individually, no human reviews it, and no language model generates it at request time. The same result code always produces the same report, which is deliberate: a report you can reproduce is one you can argue with.

Limits of self-report testing

Everything below is true of this instrument and of every other self-report personality questionnaire, including the expensive ones.

You are describing your self-image

The test measures how you characterise yourself today, which correlates with how you behave but is not the same thing. Self-knowledge varies, and people are systematically better at judging some of their own traits than others.

Presentation effects

Answers shift when something is at stake — the main reason the test must never be used for hiring or selection. Even with nothing at stake, most people tilt slightly toward the person they would like to be.

State, not just trait

Mood, sleep, a recent argument and the situation you are in all move scores, especially on emotional reactivity and extraversion. A single administration captures a moment as well as a disposition.

The reference-group effect

“I am usually organised” is judged against the people you know. Comparison groups differ by culture, profession and age, which makes cross-cultural percentile comparisons less clean than the numbers suggest.

Language and coverage

The items are in English only, and answering a personality questionnaire in a second language adds noise. The five factors are also broad by design: they say nothing about values, intelligence, skills, interests or mental health.

Prediction is modest

Traits predict tendencies over time, not individual actions. Even well-established trait-outcome relationships explain a modest share of the variance in what people actually do. Anyone promising more than that is selling something.

References

  • Goldberg, L. R. (1992). The development of markers for the Big-Five factor structure. Psychological Assessment, 4(1), 26–42.
  • Goldberg, L. R. (1999). A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. In Personality Psychology in Europe, 7, 7–28.
  • Goldberg, L. R. et al. (2006). The International Personality Item Pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84–96.
  • The International Personality Item Pool: ipip.ori.org.

Want the shorter version of all this? The disclaimer says what the test must not be used for, in one page.