Are people getting happier over time? (draft)
norvid_studies asks on Twitter:
what is the evidence base that people are unhappy in modernity? is there a lit review somewhere? [...] the thread paints a picture that modern information environments are 'unusually harsh', that is to say too stressful, but my picture is closer to the reverse, that in 'the normal condition' you have good chance of truly horrible things happening to you or close family (which essentially never happen now), and hear about them, and are constantly on guard for them [chagnon among the yanomamo] as well as all manner of discomfort, pain, and serious injury
Jokes aside—for a while, I've held the view that when one attempts to form a self-consistent theory of happiness, they end up like Daniel Kahneman:
So I discovered that I had a theory of wellbeing that didn't correspond to what people wanted for themselves and that seemed extremely awkward. It's not that I gave up on happiness, I gave up on happiness as the solution. And it's not that I adopted life evaluation. It's just that I became totally puzzled and baffled and I didn't know how to solve it so I moved on to other things.
or Scott Alexander:
I am forced to acknowledge that happiness research remains a very strange field whose conclusions make no sense to me and which tempt me to crazy beliefs and actions if I take them seriously.
However, variants of Norvid's question keep coming up in various situations, and it seems like a fun exercise to, at the very least, map out the issues that one runs into when trying to resolve them. Thus, this post will represent my best attempt to form a theory of happiness (or subjective well-being (SWB)) that allows us to perform comparisons across people and times.
#How should we collect data about SWB?
Before we can assess national SWB data, we need to first look at how SWB data is collected at the individual level and whether it's even possible to perform valid SWB comparisons across individuals. The usual way to collect SWB data is to give people surveys where ... However, if two people self-report their SWB as ‘7/10’, can we straightforwardly conclude that they're equally happy? Michael Plant argues in a Happier Lives Institute report that answering this question requires first addressing two other difficult questions:
- How do people interpret subjective scales?
- Which assumptions are required to perform valid cardinal comparisons across people?
We'll tackle these in turn. Since we want to eventually perform comparisons across groups, we'll finally ask a third question: which assumptions are required to perform valid cardinal comparisons across groups?
#How do people interpret subjective scales?
To resolve the first question, Plant invokes what he calls the Grice-Schelling hypothesis: when presented with a subjective scale, people will interpret it how they expect others will in order to make themselves understood. This, in turn, makes their answers cardinally comparable. In my mind, this solution has at least two issues. First, different people likely expect others to interpret the scale in different ways, especially if their cultural and socioeconomic backgrounds differ wildly. Second, when I think about the way I myself interpret such scales, I set the endpoints based on the best and worst experiences I can recall from my own life, instead of trying to guess how others might set them.
Why is the second point an issue? I have strong reasons to believe that if other people are using a similar process of setting the endpoints as me, we're going to end up with very different scales. I sometimes talk to people who say that their worst ever experience was 1000x worse than how they feel at the moment of the conversation, and I simply can't imagine what those experiences would be like, since my own vibes-based multiplier would be somewhere around 3x. Even if the Grice-Schelling hypothesis is true and I do try to set the endpoints based on how happy/unhappy I think anyone, rather than just I myself, can possibly be, it seems likely that I wouldn't have the emotional capacity to determine those endpoints since I'm incapable of imagining such experiences. My feelings just don't get all that intense.1
To my knowledge, Plant has not discussed my first objection. For the second objection, Plant gives the following reply:
I won’t dwell on this as it seems unlikely there would be substantial differences in humans’ capacities for subjective experiences. Presumably there are evolutionary pressures for each species to have range of sensitivity that is optimal for survival. To return to an example noted earlier, being immune to pain is an extremely problematic condition that would put someone at an evolutionary disadvantage. Further, even if there are differences, we would expect these to be randomly distributed, in which case they would wash out in large samples.
Michael St Jules provides further arguments for this view in an EA Forum comment:
It seems to me that a full defense of cardinality and comparability across humans should mention neuroscience, too. For example, we know that brain sizes differ (in total and in regions involved in hedonic experiences) in certain systematic ways, e.g. across ages and between genders. However, these differences are mostly small (brain size is pretty stable after adolescence, although there are still major changes up until 25-30 years old), and we might assume that differences in intensity of experience scale at most roughly 1:1 with size/connectivity and number of neurons firing (in the relevant regions), and while I think this is more likely to be true than not, I'm still not confident in such an assumption.
I don't find the neuroscience argument convincing at all. Consider an excerpt from the Wikipedia page of Jo Cameron:
Cameron subjectively reported a lifetime lack of pain, including with childbirth, broken bones, and numerous burns and cuts. She had often not noticed burns and other injuries until she smelled burning flesh or saw blood on herself. Her burns and cuts also seemed to heal quickly with less or no scarring. Eating Scotch bonnet chili peppers left only a "pleasant glow". Attempts by researchers to induce pain, including burning her, sticking her with pins, and pinching her with tweezers until she bled, resulted in no pain. Aside from her lack of pain, Cameron was additionally described as characteristically happy, friendly, talkative, optimistic, and compassionate, as well as exceedingly affectionate and loving towards family members. Moreover, she was lacking in anxiety, depression, worry, fear, panic, grief, dread, and negative affect generally.
For a stark contrast, consider the following quote by ...:
The differences in the brain architectures of Cameron and ... are presumably miniscule, and yet ... The evolutionary argument is somewhat stronger, but it seems plausible that very different levels of pain would have roughly the same effect on our motivations and thus have similar evolutionary advantages. Furthermore, a lot of pains are useless for survival but don't have sufficient fitness disadvantages for evolution to fine-tune them away. There may well be people for whom such pains are much stronger than for others.
...
Why don't we just say on the surveys how to interpret the scale? Researchers don't appear to do this in practice—see OECD (2013, Annex A) for an example survey.
#Which assumptions are required to perform valid cardinal comparisons across people?
...
#Which assumptions are required to perform valid cardinal comparisons across groups?
Despite all the issues involved in direct interpersonal comparisons, Plant makes a good point that if there are differences between the intensities of people's experiences, we would expect them to be randomly distributed and thus washed out in any country-sized sample. We may not be able to perform direct comparisons between individuals using simple self-reports, but our comparisons across different countries at the current moment and across the same country at different points in time should be mostly unaffected by the issues that plague interpersonal comparisons.
However, there is an additional issue. In the paper The Sad Truth About Happiness Scales: Empirical Results, Timothy N. Bond and Kevin Lang report that it is almost impossible to perform nonparametric comparisons across groups. Imagine a very simple experiment comparing happiness between two groups (A and B) on a 3-point scale:
- Category 0 = "not too happy"
- Category 1 = "pretty happy"
- Category 2 = "very happy"
Bond and Lang show that nonparametric comparisons are possible only if nobody in group A reports being "not too happy", nobody in group B reports being "very happy", and the fraction of A saying "very happy" ≥ fraction of B saying "pretty happy". This means that all groupwise happiness studies are making some implicit distributional assumptions, usually normality.
Bond and Lang show that it's possible to reverse nine famous results in happiness studies by modifying the distributional assumptions.2 How big of a problem is this in practice, though? It affects any research using ordinal self-reports to compare group means when distributions overlap, not just happiness studies. Here's a list of other domains where such self-reports are often used:
- course evaluations in education
- consumer satisfaction ratings
- medical symptom severity ratings
- policy support surveys
- LLM judges grading transcripts from other LLMs using categorical scales
Though self-reports are not perfectly robust in any of these domains, it seems a bit outlandish to claim that we're getting no signal from such studies in those fields. How big were the distributional changes Bond and Lang had to make in order to reverse the results? ...
#The Easterlin paradox
Has national SWB been increasing, then? The Easterlin paradox says that this isn't the case: it states that at any given point in time, richer countries tend to be happier than poorer ones, but that over time, happiness doesn't trend upward in tandem with income growth. Since most of the world is getting richer over time, this can be taken to straightforwardly say that happiness hasn't been increasing.
This view was challenged by a recent paper by Alberto Prati and Claudia Senik titled Is It Possible to Raise National Happiness? Prati and Senik acknowledge that Easterlin's observation is robust, but note that there are two possible explanations for the flat trend in reported SWB. The first one is hedonic treadmill: experienced happiness depends not on absolute levels (e.g., of income), but on deviations from reference points. These reference points evolve over time, thus recalibrating people's expectations and sense of their relative position. The second one is rescaling: the best possible life you could have is a shifting standard that moves upwards as living standards improve. Rescaling can also be invoked as an explanation under the Grice-Schelling hypothesis, which would posit that the best possible life anyone could have is the standard that shifts.
Formally, the difference between the two explanations is that the former posits a time-dependent utility function, while the latter assumes a time-dependent reporting function. By the former, the surrounding context acts on a person's latent satisfaction, and by the latter, it acts on their interpretation of the scale. By the former, people become harder to satisfy as the world improves, while by the latter, people actually become more satisfied but don't report it.
Both of these interpretations are observationally equivalent and, at least according to Prati and Senik, no epxeriments have been devised to date that could be used to determine which one is correct. However, we can still look at some puzzling case studies where happiness trends didn't fit our expectations and assess which theory produces neater explanations in such cases.
We'll look at four surprising observations about the happiness dynamics of large groups:
- During COVID-19 lockdowns in 2020, SWB time series were surprisingly flat all around the world.
- In Ukraine, the average level of life satisfaction was the same in 2023 as in 2018, despite the devastating impact of the Russian invasion.
- In China, the period of globalization, economic growth, and reform appears to have left people equally or even less happy than before, despite a billion people having been lifted out of poverty.
- In the US, women's happiness appears to have fallen relative to men's from 1972-2006 despite the substantial improvements in their social and economic status over this period.
Is the China claim actually true? See https://ourworldindata.org/happiness-and-life-satisfaction
Discuss the following figure:
#Other historical influences on SWB
Our economic prosperity isn't the only thing that has changed over time. Our lives are very different compared to the ancestral environment, and it's quite unclear which of these changes have had a positive and which ones a negative effect on our collective happiness.
#Experiencing self vs. remembering self
https://x.com/panickssery/status/1958650572791587003 https://tragedyandfarce.blog/2021/07/26/the-history-of-happiness-are-people-getting-happier/ https://ourworldindata.org/happiness-and-life-satisfaction https://www.happierlivesinstitute.org/research/ https://dynomight.net/happiness/ https://forum.effectivealtruism.org/posts/YdzbfCTxSCvqkmh4x/the-comparability-of-subjective-scales?commentId=ij4ujQwJis8hKCudg Harari happiness after the agricultural revolution https://www.optimallyirrational.com/p/happiness-and-the-pursuit-of-a-good https://www.aporiamagazine.com/p/economic-growth-doesnt-make-us-happier?utm_source=share&utm_medium=android&r=lzqt7&triedRedirect=true https://www.optimallyirrational.com/p/the-truth-about-happiness https://slatestarcodex.com/2016/03/23/the-price-of-glee-in-china/ https://marginalrevolution.com/marginalrevolution/2022/12/how-happy-are-americans-and-danes-anyway.html https://marginalrevolution.com/marginalrevolution/2020/04/happiness-and-the-quality-of-government.html https://marginalrevolution.com/marginalrevolution/2025/06/some-european-countries-have-mastered-a-happiness-trick.html
- other people's notion of pleasure/suffering is different from mine. I may be able to report my experience on a linear scale, but this doesn't necessarily apply to everyone - others might not be able to rationally choose between intense guilt and intense pain, these may be incomparable dimensions for them
- people's preferences are inconsistent - your experiencing self may wish for intense pain at times of intense guilt and for intense guilt at times of intense pain, and your remembering self may yet be indifferent between them or view them as incomparable
- when doing interpersonal comparisons, it's generally reasonable to say that experiences are dependent on set point and we can say that a seven for one person is the same as a seven of another, even if those sevens are based on different amounts of pleasure. However, what about extreme comparisons where one person doesn't suffer at all and gives a seven bc they had a bit more pleasure than usual, while another person gives a seven just bc they were free of suffering for some part of the day, even tho they didn't feel pleasure at all?
Footnotes
-
The 1000x figure is probably an exaggeration, and my 3x figure may well be an understatement. However, the gap is large enough that it seems highly unlikely that it would be entirely explained away by egregious exaggerations on my and my conversation partner's part. Furthermore, I can infer that there's probably a large difference in the intensity of suffering I and some other people experience by contemplating the differences in worldview between myself and, e.g., strong negative utilitarians. ↩
-
These nine results are the Easterlin Paradox, the U-shaped relation between happiness and age, the happiness trade-off between inflation and unemployment, cross-country comparisons of happiness, the impact of the Moving to Opportunity program on happiness, the impact of marriage and children on happiness, the ‘paradox’ of declining female happiness, and the effect of disability on happines. ↩