Finland once again tops the world happiness ranking. One can imagine a person reading this news at a kitchen table in Helsinki, looking at the rain outside and trying to understand what exactly they are supposed to feel. Another reader, several thousand kilometres away, asks a more practical question: perhaps it would be worth moving there? Both are looking at the same table, but they are looking for very different answers.

Our investigation began with a similar sense of uncertainty. Why do countries that appear calm and even somewhat dull rank above places where life seems much more emotional? At first, it seemed enough to understand the wording of the survey question. But then we had to examine how people use numerical scales, what they expect from the future, and what happens to their answers after they move to another country.

Contents

What the happiness ranking actually measures

In the World Happiness Report 2026, Finland ranks first with an average score of 7.764 out of 10. But participants are not asked to assess how enjoyable their week was. They are asked to imagine a ladder where the bottom step represents the worst possible life and the top step the best possible life, and then indicate where they currently stand. This question is known as the Cantril Ladder, after psychologist Hadley Cantril. [1]

The main ranking is based on average responses, and in the 2026 edition it uses data from 2023–2025. Income, health, social support, and other circumstances are used to explain differences, not added together into a score designed by the report’s authors. This distinction matters: the researchers did not declare Finland happy because it has good infrastructure. Its residents themselves rated their lives highly. [1]

However, evaluating one’s life and having a pleasant day are different activities. In the first case, a person is almost looking at their biography from a distance: work, family, security, goals achieved, and comparison with what might have been. In the second case, the person is inside that biography: tired during the commute, laughing with a friend, worried by bad news, enjoying dinner. Overall evaluation and daily experience can diverge without either being wrong.

Consider a person who is building a company. They see their life as successful, do work they consider important, and can see the results, but spend much of the week under pressure. Nearby there may be another person with lower income and uncertain prospects who has much more ease, social contact, and laughter in everyday life. The question “who is happier?” cannot be answered until we specify which of these properties we mean.

The ambiguity is much older than modern rankings. In Aristotle, a good life is connected with how a person acts and develops human capacities; in Epicurus, freedom from pain and mental disturbance is especially important. These are different reference points, not simply two ancient versions of the modern question “do you feel pleasant right now?”. Psychological surveys inherited this ambiguity and tried to separate it into measurable components. [2]

What we want to know Question in ordinary language What it is easy to confuse it with
Cantril life evaluation Where is my life between the worst and best possible life? The joy a person experiences every day
Life satisfaction How satisfied am I with my life overall? Income or everyday comfort alone
Direct feeling of happiness How happy do I usually feel? A direct measurement of emotions at every moment
Positive experiences Did I experience enjoyment, laughter, or interest yesterday? The absence of anxiety and sadness
Negative experiences Did I experience much worry, sadness, or anger yesterday? The complete absence of good things in life
Meaning and purpose Is there something important in my life that I act for? An easy and pleasant life
Expectations What do I expect my life to be like in five years? An accurate forecast or willingness to act

These are practical explanations, not interchangeable definitions. The word happiness requires particular care: in one study it means a general feeling of happiness, in another an overall life evaluation, and in a third the experience during a specific activity. A single translation as “happiness” can hide differences that later become visible in the results.

Why three questions produce three different rankings

The first suspicion was simple: perhaps different rankings disagree because they compare different years and different samples. To remove this problem, we need the same people answering several questions at the same time. The Global Flourishing Study made this possible. In a published analysis of 202,898 participants from 22 countries, researchers compared the Cantril Ladder, life satisfaction, and general feelings of happiness. [3]

Countries changed position even though the respondents were the same. We recalculated the ranks from the published table of means; several examples show the scale of the changes.

Country Rank by Cantril Rank by life satisfaction Rank by feeling of happiness
Israel 1 6 3
Sweden 2 9 11
Mexico 4 2 2
Indonesia 5 1 1
United States 6 13 12
Philippines 14 5 7
Egypt 21 3 21

Ranks within these 22 countries, not in the global ranking. The source of the means is Table 3 of the GFS study. [3]

Egypt is particularly difficult for a simple explanation. It cannot be described by saying that “people in this culture always respond more positively”: it is near the top for life satisfaction and near the bottom for the other two questions. This means that not only general response tone matters, but also the specific mental task created by the wording of the question.

There is experimental evidence pointing in this direction. In a 2024 study, participants were given different versions of the Cantril question. The ladder metaphor directed thoughts more strongly towards wealth and power than alternative formulations did. This does not prove that every person answering the ladder is thinking mainly about money, but it shows that the measurement instrument can influence what enters the respondent’s attention. [4]

This led to a useful interpretation: the Cantril Ladder may partly reflect a person’s position in an internal hierarchy of what counts as a good life. The words “may partly reflect” are important. This is an explanatory model, not a hidden translation of the questionnaire and not proof that the ranking measures status alone.

At this point, it might seem that the issue had been resolved: instead of one form of happiness, we had several different things. But that conclusion would have been too convenient. Later we found that some of the disagreement was created not only by different questions, but also by different habits in the use of numerical scales.

Can people feel joy and distress at the same time?

Before that turn, it is worth examining another common assumption. We often imagine happiness as a thermometer: more joy means less suffering. But a day with a major success, a difficult conversation, and strong anxiety does not necessarily become an “average” day. It can contain a great deal of both positive and negative experience.

Differences between rankings based on life evaluation and everyday experience were known long before our analysis. For example, David Blanchflower and Alex Bryson compared several Gallup measures and showed how strongly the position of territories changes when moving from life evaluation to smiles, rest, pain, or worry. Importantly, these are different survey questions about experience, not observation of facial expressions in public places. [5]

We separately compared positive and negative experiences in the World Happiness Report 2026 tables. Across 147 countries and territories for the same period, the association was negative but moderate: about −0.37. In other words, where pleasant experiences are reported more often, unpleasant experiences are on average less frequent, but one measure is far from determining the other. [1; A1]

Country Positive experiences Negative experiences What stands out
Iceland 0.802 0.172 Many positive experiences, relatively few negative ones
Guinea 0.696 0.490 High values can occur on both dimensions
Russia 0.601 0.195 Lower values on both dimensions
Afghanistan 0.261 0.460 Few positive experiences together with frequent negative ones

Indices from 0 to 1 based on questions about the previous day. These are not the share of “happy people” and not the percentage of time spent in a particular state. Values are taken from the WHR statistical appendix for 2023–2025. [1]

National averages do not show that exactly the same residents of Guinea reported all of these experiences simultaneously. For that, individual-level responses are required. But even at the country level, one point is already clear: reducing two columns to a single difference removes information.

Consider two hypothetical profiles: 0.7 positive experiences and 0.5 negative experiences, or 0.3 positive and 0.1 negative. The balance is the same in both cases: 0.2. Yet the first profile describes a much more emotionally intense environment, while the second describes a calmer one. Which is preferable is a value judgement for the individual, not something arithmetic can decide.

For this reason, the first useful map of happiness turned out to be very simple: one axis for pleasant experiences and another for unpleasant ones. Such a map can distinguish joy with little suffering, emotionally intense life, calmness, and difficult conditions. This still does not capture all of wellbeing, but it is more informative than one overall volume control.

Emotional map: positive × negative experience

Positive and negative experience are separate dimensions. A country can report relatively high levels of both. Hover over any point or search for a country.

Source: World Happiness Report 2026 statistical appendix, 2023–2025. Published data; 147 countries and territories.

Is happiness inside the person or in external conditions?

The next question followed naturally. Perhaps the Cantril Ladder reflects external life more strongly, while a direct happiness question reflects personality? If so, the rankings might simply be looking in different directions: one towards the country, the other towards the person. This can be tested by following the same individuals over time.

In the research part of the project, we analysed the first two waves of GFS: about 128,000 participants with responses in both waves, using longitudinal survey weights. We compared three things: how stable the responses were over one year, how they related to personality traits, and how they changed together with employment, finances, and health. These analyses are preliminary; their status and limitations are described in the methods note. [6; A2]

The first test immediately disrupted the neat model. The most stable measure was not the direct question “how happy are you?”, but life satisfaction. The one-year correlation was approximately 0.50 for life satisfaction, 0.47 for happiness, and 0.45 for the ladder. The differences are small, but they do not support a simple story in which one scale measures an unchanging inner core while another merely reacts to external conditions.

At the same time, direct happiness was more strongly associated with personality characteristics. Employment and health also left different patterns: losing or gaining a job moved the ladder more noticeably, while worsening health was more visible in life satisfaction and feelings of happiness. After adjustment for response tendencies, the general pattern remained, although some associations became weaker. [A2] For a separate, personal rather than population-level exploration of felt safety and inner experience, see Guided Meditation for Feeling Safe in the World.

Measure What stood out in our analysis Cautious practical interpretation
Cantril Ladder Distinguishes countries more strongly; noticeably related to changes in employment status Where I stand relative to my idea of a good life
Life satisfaction Most stable over one year; related to changes in personal circumstances How my life is going overall
General feeling of happiness More strongly related to personality and psychological state How life usually feels from the inside

These are not three sealed compartments. Health includes both bodily condition and perception of that condition; losing a job affects income, status, daily routine, and relationships. Even observing the same person before and after an event does not turn job loss into a randomly assigned experiment. The table therefore describes sensitivity of the measures, not proven causal decomposition of happiness.

Twin studies also do not create a clear boundary between “inner” happiness and “external” life evaluation. A meta-analysis estimated the heritability of wellbeing at roughly one third of differences between people, but estimates depended on the sample and the instrument. In a separate study, several wellbeing scales showed a substantial shared genetic component. Heritability does not mean that one third of a particular person’s happiness is permanently written into their DNA. [7]

Why “90% of differences are within countries” does not mean “everything depends on you”

In the published GFS analysis, differences between countries accounted for about 6–10% of the total variation in the three measures. The rest of the variation was within countries. This does mean that a national average is a poor description of an individual resident, but it does not prove that external conditions are unimportant. [3]

Within one country there are people with chronic pain and people in good health, isolated pensioners and people with close families, unemployed people and owners of stable businesses. Statistically, all these differences remain “within the country”, although many of them concern external circumstances. A national border is simply too large a unit to separate personality from environment.

A simple analogy is two houses. One has a slightly higher average temperature, but each has sunny rooms, cold corners, functioning radiators, and draughts. A small difference in averages does not make heating irrelevant and does not turn all room-level differences into properties of the occupants. The same applies here: a country shifts the conditions, but it does not give every person the same life.

Why a Japanese seven may not equal someone else’s seven

Until this point, we had implicitly assumed that the same numbers represented roughly the same states. But imagine two people who are similarly satisfied with a restaurant. One gives five stars because everything was good; the other gives four because five is reserved for an exceptional case. If thousands of such reviews are averaged, a stable difference appears even if the restaurants are equally good.

International surveys can face a similar problem. Some respondents readily choose extreme values, while others prefer the middle of the scale. This is usually described as a response style. The existence of such tendencies has been studied for many years; our question was narrower: how much do they change the conclusions about happiness that we had already reached?

We estimated the use of scale endpoints and midpoints across 18 additional GFS questions, excluding the three main wellbeing measures. We then examined another scale, the personality items with seven response options rather than eleven. A tendency to choose extreme responses was reproducible across blocks and across waves. This supports the hypothesis of a more general way of using rating scales, although it does not allow us to separate response style completely from the content of real experience. [A2]

In our saved analyses, the differences were particularly visible between Japan and several countries where extreme values were used more often. But this does not justify the conclusion that Japanese respondents are “really” happier or that other populations exaggerate. Differences in language, norms of self-expression, education, interview mode, and the real distribution of life experience may all contribute. The questionnaire reveals patterns in responses, not a ready-made explanation of national character.

A turn that went against our own hypothesis

We then tried to account statistically for the general tendency to use the middle or the endpoints of the scale. After this adjustment, national rankings based on different questions became much more similar. For the Cantril Ladder and direct happiness, rank agreement increased from about 0.68 to about 0.90 in one model; an alternative factor model produced a similar picture. [A2]

The meaning is simpler than the coefficient. Some countries had previously appeared far apart not only because the questions captured different aspects of life. Differences in translating internal experience into a number also seem to have contributed. Once we tried to account for this translation process, the three measures agreed more closely.

This required us to weaken our original conclusion. It was no longer reasonable to say that different questions describe almost independent worlds of happiness. They share a common component, while also containing substantive differences and measurement distortions. It was also notable that country ordering on the Cantril Ladder was more stable under some of our endpoint-response adjustments than ordering based on direct happiness or life satisfaction.

However, the new ordering cannot be called a “corrected ranking”. The auxiliary questions also concern real life. If we remove everything that looks like a general response style, we may remove genuine wellbeing together with the habit of selecting ten. Our adjustments are sensitivity analyses, not calibration against a known value of “true happiness”.

A ranking compares not only people’s experiences, but also the rules by which people convert those experiences into numbers.

Can happiness be measured without surveys?

After encountering this problem, it is tempting to put questionnaires aside and use something objective instead. Count smiles, measure sleep, examine blood tests, or calculate how much free time people have. But objectivity of measurement does not guarantee that the relevant property is being measured. A device can count the steps of a person walking for pleasure and a person pacing because of anxiety with equal precision.

A review of 91 studies on physiological correlates of wellbeing did not identify a simple “test for happiness”. For example, the association with momentary cortisol levels was very weak. This does not make physiology useless: it tells us about processes in the body, but the same biomarker can depend on many mechanisms and cannot by itself provide a verdict on the quality of a person’s lived experience. [8]

Smiles also require interpretation. A person may smile because of pleasure, politeness, embarrassment, or professional expectations; in addition, a street camera sees only those who happened to be outside at a selected place and time. Computer recognition of facial expressions can make counting scalable, but it cannot remove these interpretation problems. In our project, smile observation remained a direction for future work rather than a completed measure of national happiness.

How an almost perfect association with commuting failed to become an explanation of happiness

Time-use diaries provided a more accessible independent source of data. We compared harmonised European HETUS tables for 11 countries with life evaluation in WHR. We first examined sleep, eating, and personal care, and later moved to leisure and travel. An important limitation is that the surveys were conducted in different years, and diary categories describe activities rather than the quality of experience during those activities. [9; A3]

Several expectations did not hold. Average sleep duration hardly distinguished countries with higher and lower life evaluations. The association for the category “other personal care” became much weaker after accounting for economic differences. More leisure could represent either beneficial freedom or involuntary unemployment, or limitations related to health.

Commuting time produced an almost implausibly clean result: the country ranking by travel time was nearly the mirror image of the ranking by life evaluation. The rank correlation was about −0.96. It seemed that we had found a concrete daily mechanism: the less life is spent commuting, the better life feels. [A3]

But this was a small sample after examining many activities, precisely the situation in which an attractive result should not be trusted without an attempt to break it. In addition, minutes were averaged across all people in the selected age group and across all days. Fourteen minutes in such a table does not mean that the typical commuter travels for fourteen minutes: the mean also depends on how many people travelled at all on that day.

We then examined changes over time and did not obtain the same simple story. In Poland, life evaluation increased substantially while commuting time changed little and leisure time decreased. In Norway, commuting time declined in the published age groups, but national life evaluation did not increase. These comparisons are not perfectly aligned in age ranges and periods, but they already make it difficult to explain the cross-country pattern through commuting alone. [10; A3]

What initially looked convincing What the next test showed
More leisure — higher life evaluation Life evaluation can rise even when leisure decreases
Less commuting — higher life evaluation Reduced commuting does not guarantee an increase in the national average
More household obligations — worse life Different obligations relate to life evaluation in different ways; no coherent combined index emerged
Objective minutes can replace subjective surveys Minutes require information about the activity, choice, and experience

This is not evidence that commuting is harmless. It may be tiring, reduce time with family, and worsen a particular day, while its reduction may be beneficial but offset by other changes. But a strong association between countries does not allow us to calculate how much happiness will increase if twenty minutes of commuting are removed.

A short commute may partly indicate a better organised environment: housing, transport, flexible employment, and more control over daily schedules. For now, this is an explanation to be tested. We also abandoned an attempt to combine commuting, cleaning, and shopping into a single measure of “life friction”: the data did not support treating these activities as one mechanism.

Life is good now — but is there something ahead?

At this point we returned to the original impression: a comfortable life can sometimes appear to lack movement. If a person is already high on the ladder, is there still something to strive for? Could a high current score combined with limited positive emotion indicate that a society has reached a comfortable plateau and is losing initiative?

This is an attractive hypothesis, but it contains several separate claims. Satisfaction with the present does not imply a lack of goals; a small expected improvement does not imply boredom; boredom does not imply future decline. We first examined historical data: the combination of high life evaluation and lower positive experience did not predict a stable subsequent decline. There was no basis for describing calm countries as destined for stagnation. [A4]

However, we found that the relevant question about the future already exists. Gallup asks not only about the current step on the ladder, but also about the expected step five years later. Tim Lomas and colleagues published values for 145 countries and territories using data from 2020–2022. These data allow current and expected future life evaluations to be considered together. [11]

Country Life now Expected life in five years Difference between means
Sierra Leone 3.08 6.80 +3.72
Brazil 6.10 8.16 +2.06
Russia 5.68 6.60 +0.92
United States 6.88 7.76 +0.88
Denmark 7.57 8.16 +0.59
Sweden 7.39 7.84 +0.45
Finland 7.79 7.94 +0.15
Japan 6.13 6.26 +0.13

Gallup World Poll, 2020–2022; scale 0–10. The final column is our subtraction of the published rounded means. This is a separate historical cross-section and should not be mixed with the WHR 2026 ranking. [11]

Infographic comparing current and expected future life evaluations in Finland, the United States, Brazil, and Sierra Leone.

The difference is visible without complex statistics. Finland evaluates the present highly and expects approximately the same level to continue. Brazil gives its future an even slightly higher absolute score, but expects a much larger increase from the present. These profiles describe different relationships to the future, although neither tells us what will actually happen.

We called the difference between future and current evaluation expected life improvement. In working discussions we used the more expressive term Future Pull. However, the expression turned out to be broader than the measurement: believing that life will become better is not the same as looking forward to the future with interest or being willing to act in order to change it.

A person may expect improvement because they hope to find a job, finish treatment, or reach the end of a difficult period. Another person may want to write a book, raise children, or travel without expecting their overall life score to rise because life is already good. The gap therefore should not be converted into a ranking of social energy.

Why we abandoned an overly sophisticated normalisation

The difference has an obvious mathematical problem. If the current evaluation is nine, only one point remains before the maximum; if it is four, six points remain. In addition, when the horizontal axis contains the current score and the vertical axis contains future minus current, the same quantity appears in both axes with opposite signs. Part of the negative slope in the scatter plot is created by the construction itself.

We tried an adjustment: instead of comparing future directly with the present, compare it with the typical expectation among countries starting at a similar current level. But the position of some countries became highly dependent on the chosen curve. Finland’s position, for example, changed substantially under several reasonable specifications. This was a warning that the model itself was beginning to determine the story too strongly. [A4]

When we compared the 22 countries shared with GFS against separate measures of hope and optimism, the simple difference was more strongly associated with them than the more complex normalised version. We therefore kept the transparent subtraction for descriptive purposes. But this is a preliminary comparison of country means across periods that do not fully overlap; it does not remove the scale ceiling and does not prove that the difference measures individual motivation. [A4]

Thus, two axes remained useful: how good life is now and how much better it is expected to become. What disappeared was the stronger promise that these dimensions could already diagnose social stagnation. The map became more useful after we stopped asking it to do more than the data supported.

Present life × expected future

Current life and expected improvement are different properties. The default view shows the two original scores. Switch to “Expected improvement” to view the gap between them.

Source: Gallup World Poll 2020–2022 values published by Lomas et al.; 145 countries and territories. Expected improvement = Future Cantril − Current Cantril, calculated from published rounded means. It is an expectation, not a forecast.

Do expectations about future happiness come true?

This made possible a test that ordinary rankings rarely offer. If people were asked what life would be like in five years, we can wait for those five years and examine what actually happened. In our case, no waiting was necessary because part of the historical expectations already had an observed future.

In a preliminary analysis, we matched early published Gallup estimates with average life evaluations in 2010–2012. We were able to match 137 countries. This is only an approximate five-year comparison because the original surveys were not conducted in exactly the same calendar year. The result should therefore not be presented as a precise test of each expectation against one fixed future date. [12; A4]

Test across 137 matched countries Preliminary result
Countries where the expected future level was above the later observed level 132 of 137
Mean absolute error of the expected future level About 1.30 points
Error of the simple assumption “the level will remain unchanged” About 0.37 points
Association between expected improvement and subsequent change Approximately zero

Our historical comparison, not a forecasting test published by the authors of the original survey. Methodological limitations are described in note A4.

The most important result is not simply that expectations were too optimistic. Persistent optimism could in principle be adjusted for as a systematic bias. More important is that countries expecting a larger increase did not, in this analysis, improve noticeably more than others. High expectations were not a reliable indicator of subsequent national trajectory.

We also examined another available source, a published global Gallup series. Across six completed five-year comparisons, simply assuming that the current level would remain unchanged again produced a smaller error than the average answer about the future. This is a short aggregate series and cannot be treated as six independent experiments. Nevertheless, the direction of the result matched the cross-country analysis. [13; A4]

A final attempt was then made to rescue the forecasting idea: perhaps the important quantity is not the level of optimism but its change. Is confidence in the future strengthening or weakening? Using the available rounded regional series, we did not find a convincing predictive relationship, but the test is weak: there are few observations, neighbouring years are dependent, and rounding particularly damages small annual changes. A full multi-year country-level dataset remains an unfinished part of the project. [A4]

There is also a deeper limitation. We compared expectations from one national sample with responses from a different sample several years later. This is not observation of the same person before and after. The correct conclusion is therefore: average national expectations in the tested data were poor predictors of subsequent national averages. These analyses do not tell us how accurately individuals predict their own lives.

Expected improvement nevertheless remains meaningful. Hope may help people tolerate difficult periods or influence decisions to study, migrate, start a family, or build a business even when the future life score is predicted poorly. That is a different hypothesis about behaviour rather than forecasting accuracy. In our material, it remains an open question for future research. A related Mike’s Balance article looks at future orientation, purpose, and long time horizons at the individual level.

Do migrants become happier in a new country?

By this point, an ordinary happiness ranking might appear almost useless. Its result depends on the question, on response-scale habits, and on whether we average emotional experience or life evaluation. But the reader at the second kitchen table was still there: the person considering migration. For this reader, sensitivity to external conditions may be a strength rather than a weakness of the measure.

The practical question is not which population has the most cheerful temperament. It is what happens to people after their environment changes. If life evaluation responds to security, work, relationships, and everyday predictability, then this is precisely the component a potential migrant wants to understand. The criticism of rankings therefore becomes, unexpectedly, a test of their practical usefulness.

What our calculations on European migrants showed

We analysed pooled European Social Survey data, rounds 1–11. The file contained about 541,000 participants, around 50,000 of whom had been born outside the country where they were surveyed. For specific comparisons, however, the usable sample was much smaller: we needed migrants from a defined origin together with appropriate groups of local residents in both the origin and destination countries within the same ESS rounds. [14; A5]

We compared three positions: residents of the origin country, migrants, and native residents of the destination country. In several models, average migrant life satisfaction lay roughly two thirds of the way from the origin-country mean towards the destination-country mean. In a stricter specification, we retained only people who migrated as adults after 1991; the general pattern remained, although some route-specific samples became small. [A5]

Consider a hypothetical example: average life satisfaction is 6 in the origin country, 8 in the destination country, and 7.3 among migrants. Migrants are then approximately 65% of the way between the two group means. This is a convenient description of relative group position, but it does not mean that the state creates 65% of a person’s happiness or that every individual gained 1.3 points after moving.

ESS is primarily a repeated cross-sectional survey, with different people interviewed in different years. We do not know what each migrant would have reported if they had stayed in the origin country. Some people leave for a strong job offer, some follow a spouse, some leave unsafe conditions; those whose migration went badly may return and disappear from the destination-country sample. These mechanisms prevent us from interpreting convergence as a clean causal effect of migration.

However, we found patterns that are difficult to explain by a simple story in which “the more positive people are the ones who leave”. When people moved to countries with lower average life satisfaction, migrant means also shifted downward relative to the origin-country mean. This does not eliminate more complex forms of selection, but it makes a general positive-emigrant bias a less complete explanation. [A5]

Do migrants simply start using different numbers?

After the GFS results, we had to test this possibility. Migrants may adopt not only the habits of a new environment but also its style of responding to rating scales. If so, part of the convergence towards the destination country could occur on paper: life may have changed less than the numerical expression of life evaluation.

Across 19 other ESS scales, we did find convergence in response style towards the destination country. The tendency to use midpoints or endpoints was not a fixed imprint of the origin culture. In a life-satisfaction model, accounting for these indicators reduced the convergence coefficient from about 0.74 to about 0.62, but did not remove it. [A5]

This cannot be translated into the statement that “16% of the effect is measurement error and the remaining 84% is real”. The auxiliary scales contain real evaluations of the environment as well, so the adjustment may remove some genuine substantive change. Nevertheless, the test refined the picture: numerical adaptation exists, but the response-style measures we used do not explain all of the observed convergence.

Why national life satisfaction still turned out to be useful

We tried to replace the difference between local residents’ self-reports with a set of external indicators: economic conditions, unemployment, government effectiveness, and life expectancy. The extended analysis contained 34 migration routes and 101 route-by-round observations. For out-of-sample testing, each migration route was left out of model training in turn and the model then attempted to predict its outcome. [A5]

Basis for predicting the difference between migrants and residents of the origin country Prediction error, scale points
Difference in life satisfaction between local residents of the two countries About 0.49
Set of external indicators About 0.70
Local residents’ responses plus external indicators About 0.47

Preliminary test on a limited set of European migration routes. This predicts group means, not the future of an individual person; lower error is better. [A5]

An apparent advantage of the external indicators, which we initially thought we had found, disappeared after correcting the time matching. In the final version, people’s own evaluation of their lives was more informative than several separately measured national characteristics. One possible explanation is that self-reported life evaluation combines information about living conditions that our short external model measures poorly or does not measure at all.

It is important not to overcorrect in the opposite direction and give the ranking too much authority. ESS measures life satisfaction with its own question; this is not a direct validation of every position in WHR. And the advantage we observed in this sample is not guaranteed to hold in another region or for a particular family. Still, there is now stronger justification for using national life evaluations as one practical reference point.

What stronger research designs show

Our overall pattern is consistent with research on migrants in Canada and the United Kingdom: both their life satisfaction and the distribution of their responses moved noticeably towards the destination country. However, these comparisons are also mainly observational. More informative studies measure the same people before migration, or use designs where the opportunity to migrate is allocated randomly. [15]

Migration Study design Main finding Main limitation
Russia → Finland The same Ingrian Finnish participants were measured before migration and afterwards Life satisfaction increased; self-esteem followed a different trajectory Specific population, attrition, no comparable non-migrant control group
Tonga → New Zealand The opportunity to migrate was allocated through a visa lottery Material conditions and some measures of mental wellbeing improved, but a happiness measure decreased Specific migration route, lottery participants, and specific measurement instruments

Sources: Lönnqvist and colleagues; Stillman and colleagues. [16, 17]

The New Zealand lottery is particularly important for the entire investigation. The same change in external life can improve several outcomes and worsen another. This means that caution is still required with the phrase “migration makes people happier”, even when the causal design is much stronger than an ordinary country comparison.

For a person making a real decision, this suggests a practical sequence. First, examine how local residents evaluate their lives; second, look for evidence on migrants with a similar background and situation; then assess employment, language, relationships, legal status, and the conditions of the specific city separately. The national average is useful near the beginning of this process, but cannot replace the rest of it.

A happiness map instead of a league table

After all these tests, the desire to build “our own finally correct ranking” became weaker. To compress everything into one column, we would have to decide how much joy compensates for anxiety, how much meaning compensates for inconvenience, and what weight should be given to hope that may not come true. Statistics can describe these dimensions, but it cannot choose the reader’s values.

A map in which a point represents several properties of a country is more informative. The main version that emerged from our work keeps two coordinates: current life evaluation on the horizontal axis and expected improvement on the vertical axis. High current evaluation and strong expectations for improvement no longer have to compete for one position in a single ranking.

Position on the map What can reasonably be said What the map does not prove by itself
High current evaluation, large expected increase People evaluate the present highly and expect further improvement That the improvement will occur
High current evaluation, small expected increase A good current level is expected mainly to continue That people are bored or have nothing to work towards
Low current evaluation, large expected increase There is a large gap between the present and the expected future That society is energetic or developing rapidly
Low current evaluation, small or negative increase Low present evaluation is combined with limited expectations That every individual lacks hope or goals

These are descriptions of regions of the graph, not four proven types of society. Lines can be drawn for readability, but they do not create natural boundaries between countries. A useful alternative view places current and future scores directly on the two axes, with a diagonal line where Y = X. This makes both original quantities visible rather than hiding them inside a difference.

Is a third axis needed?

A rotating three-dimensional chart would probably make the result harder to understand. But point area could represent a third property, for example the share of people who often feel capable of managing what they need to do. GFS includes a mastery item that captures subjective personal competence. This is a potentially useful additional dimension: expecting a better future and feeling able to act are not the same thing.

In our preliminary country-level comparison using GFS means, this measure was not reducible to life satisfaction or hope. The visual idea of “two axes plus point size” therefore has some basis. However, self-rated competence is also sensitive to wording and response style; it should not be described as an objective stock of national initiative. [A4]

For the broader Gallup map, no directly comparable third measure is available in our prepared dataset. It would therefore be inappropriate to attach point sizes from GFS to coordinates taken from different years and present the result as one coherent measurement. The defensible options are a separate GFS map with aligned measures, or a world map with equal point sizes. An attractive bubble chart is not worth breaking comparability.

Why one point per country is still not enough

Consider two hypothetical countries with the same mean score of 7. In the first, almost everyone answers around seven. In the second, half answer four and half answer ten. The averages are identical, but the distributions are completely different. A point therefore needs at least some information about dispersion and the share of low scores; and differences in the shape of the distribution should themselves be checked for response-style effects.

Large countries also require regional resolution. The United States, Russia, or Brazil do not become homogeneous because an international table assigns each one a single row. But splitting them into many points is justified only where there are enough observations and appropriate survey weights: precision matters more than an attractive administrative map.

The result is not a new global queue from best to worst, but an atlas. Current life evaluation, expectations, emotions, and response distributions can be examined separately. For potential migrants, an additional layer should include evidence on immigrants rather than allowing them to disappear inside the average for all residents.

What we now know — and what question remains

What can be considered reasonably well established

Life evaluation, general feelings of happiness, and yesterday’s emotions are related but not interchangeable. Differences appear even when the same participants answer within the same study. Any claim that “country A is happier” therefore needs an immediate qualification: according to which question, over which period, and among which population?

Personality and external conditions both matter. A small share of variance between countries does not show that happiness is located almost entirely inside the person. Migration studies, including designs with stronger causal identification, show that changing environments can change wellbeing, and different aspects of wellbeing can move in different directions.

Response-scale habits are a real part of the measurement problem. Our preliminary analyses further suggest that accounting for these habits can make rankings based on different questions more similar. It is therefore premature to criticise one scale because it gives an inconvenient result and treat another as the benchmark.

Finally, the distribution of responses contains more information than the national mean. A high-ranking country does not promise good wellbeing to every resident and does not invalidate the experience of those who are doing badly there. Rankings describe groups; people live individual lives.

What remains uncertain and why

We do not know what proportion of cross-country differences reflects actual experience and what proportion reflects the translation of experience into words and numbers. Response-style adjustments help test robustness, but they themselves require assumptions. This project does not contain an independent standard against which “true happiness” can be calibrated.

We also cannot interpret the observed convergence of migrants towards the destination country as a pure effect of moving. Self-selection, the circumstances of migration, and selective return remain important alternative explanations. Individual forecasting requires following the same people before and after migration, including those who later return.

Expected life improvement turned out to be an interesting description of attitudes towards the future, but in our preliminary tests it did not become a reliable predictor of later national levels. This does not mean it is irrelevant for behaviour. The relationship between hope, initiative, decisions, and real actions remains untested in our material.

The most interesting question left after the investigation

The most important next question is what exactly changes when a person moves to a country with higher life evaluation: daily emotional experience, overall judgement of life, reference group, or the rules used to answer a questionnaire. We now have reasons to think that several layers may change at the same time. Their separate contributions have not yet been identified reliably.

To answer this, we would need to follow the same person before and after migration, combining life evaluation with short experience diaries, questions about expectations, and independent indicators of response style. It would be especially important to continue following people who later decide to return. This would bring us closer to the question for which a reader opened the ranking in the first place: what change in life would actually make this particular person’s life better?

Methods and sources

What we calculated in this article

The article combines published research with our own exploratory analyses recorded across eight research branches of the project. During preparation, key primary sources were checked and later corrections to early conclusions were incorporated. The complete original datasets and executable code for the author analyses were not part of this publication package, and we do not claim an independent replication of every calculation here. Results from our own work are therefore marked as preliminary and separated from conclusions reported in published studies.

  • A1. Emotional profiles. Comparison of 147 country and territory means from the WHR 2026 statistical appendix, period 2023–2025. A cross-country correlation does not describe the relationship between experiences within an individual person.
  • A2. GFS. Comparison of waves 1 and 2 using longitudinal weights, including response stability, personality, and changes in circumstances. Response-style checks used additional 0–10 scales and a separate 1–7 block; the adjustments reported here are not treated as measurements of a “true” level after removing error. Attrition between waves and differences in survey conditions limit interpretation.
  • A3. Time use. Exploratory comparison of 11 HETUS countries, mainly for ages 15–64. Multiple activity categories, small sample size, differences in survey years, and the compositional nature of a 24-hour day require caution. Within-country checks over time are not a full causal panel.
  • A4. Present and future. Descriptive map using published Gallup means for 2020–2022; comparison of the simple future-minus-current difference with normalised alternatives; separate comparison with GFS across the 22 overlapping countries. The historical test across 137 countries used an approximate later window of 2010–2012 and relied on a public historical WHR copy rather than a fully reproduced primary Gallup dataset. Baseline and later samples consist of different people. A complete annual country-level panel of future evaluations and individual validation against all available hope measures were not completed in the provided materials.
  • A5. Migration. ESS1–11, matching origin and destination countries within the same rounds; separate restrictions for adult migration and migration after 1991. Different models use different route sets, so their coefficients should not be treated as one universal constant. The later time-aligned external-indicator comparison replaces an earlier estimate based on a 2022 cross-section. Response-style checks are sensitivity analyses; return migration is not followed.

Numbers in the article are rounded. A difference in scale points should not be translated into percentages of happiness: a 0–10 scale does not justify saying that eight points means “twice as much happiness” as four. Failure to find a convincing association in a small sample also does not establish the absence of an effect.

Primary publications and data

[1] World Happiness Report 2026. Methodology of the main ranking and emotion measures: Chapter 2, statistical appendix, including Figures 40–45, official Finland data.

[2] Aristotle; Epicurus. Primary philosophical texts: Nicomachean Ethics, Book I; Letter to Menoeceus. Historical context, not empirical evidence for modern psychological models.

[3] Lomas T. et al. (2026). Exploring associations of three evaluative subjective wellbeing measures (Cantril’s ladder, life satisfaction, happiness) with 15 childhood and demographic factors across 22 countries. Scientific Reports, 16, 8025. DOI and full text. Table 3 provides country means and distributions; our longitudinal analyses in A2 are not part of this publication.

[4] Nilsson A. H. et al. (2024). The Cantril Ladder elicits thoughts about power and wealth. Scientific Reports, 14, 2642. Full text.

[5] Blanchflower D. G., Bryson A. (2024; online 2023). Wellbeing Rankings. Social Indicators Research, 171, 513–565. DOI and full text.

[6] Global Flourishing Study. Wave 1 and Wave 2 source data and documentation used in our analyses: official access page. Questionnaire description: Lomas et al., The development of the Global Flourishing Study questionnaire.

[7] Bartels M. (2015). Genetics of Wellbeing and Its Components Satisfaction with Life, Happiness, and Quality of Life: A Review and Meta-analysis of Heritability Studies. Full text. Additional direct comparison of scales: Bartels M., Boomsma D. I. (2009), Born to be Happy? The Etiology of Subjective Well-Being. Full text.

[8] de Vries L. P., van de Weijer M. P., Bartels M. (2022). The human physiology of well-being: A systematic review on the association between neurotransmitters, hormones, inflammatory markers, the microbiome and well-being. Neuroscience & Biobehavioral Reviews, 139, 104733. Publication.

[9] Eurostat, Harmonised European Time Use Surveys — HETUS 2020. Activity classification, survey periods, and comparability limitations: metadata; table tus_20age.

[10] Statistics Poland. Time use survey in 2023, including comparisons with 2013: publication and tables. Our comparisons with WHR are described in A3.

[11] Lomas T. et al. (2025; online 2024). A multidimensional assessment of global flourishing: Differential rankings of 145 Countries on 38 wellbeing indicators in the Gallup World Poll, with an accompanying factor analysis of the structure of flourishing. The Journal of Positive Psychology, 20(3), 397–421. DOI; author version with supplementary tables.

[12] Gallagher M. W., Lopez S. J., Pressman S. D. (2013). Optimism Is Universal: Exploring the Presence and Benefits of Optimism in a Representative Sample of the World. Journal of Personality, 81(5), 429–440. DOI; author publication. Used as a source of early current and future life evaluations; the later comparison is our analysis A4.

[13] Crabtree S., Diego-Rosell P., Buckles G. / Gallup, GSMA (2018). The Impact of Mobile on People’s Happiness and Well-Being. Technical report with historical series. We use the life-evaluation series, not the report’s conclusions about mobile technology.

[14] European Social Survey. ESS1–11 source surveys and ESS Multilevel Data 2025: official data portal. Conclusions in A5 are analyses from our project and are not attributed to ESS.

[15] Helliwell J. F., Shiplett H., Bonikowska A. (2020). Migration as a test of the happiness set-point hypothesis: Evidence from immigration to Canada and the United Kingdom. Canadian Journal of Economics. DOI. For broader international comparison, see also World Happiness Report 2018, Chapter 3 and Appendix, Table A6.

[16] Lönnqvist J.-E. et al. (2015). The mixed blessings of migration: Life satisfaction and self-esteem over the course of migration. European Journal of Social Psychology. DOI; description of the Ingrian-Finnish Remigrants 2008–2013 data.

[17] Stillman S., Gibson J., McKenzie D., Rohorua H. (2015). Miserable Migrants? Natural Experiment Evidence on International Migration and Objective and Subjective Well-Being. World Development, 65, 79–93. Publication; author version.

Medical information

This article may contain published medical evidence, clinical context, personal observations, or hypotheses. These are not equivalent levels of evidence. See the Editorial & Medical Review Policy and Medical Disclaimer. This content is educational and does not provide an individual diagnosis or treatment plan.