Links for Q3 2025
#Blogs and Essays
I spent part of this quarter sprinting for ICLR and part of it moving/travelling, so I read fewer blog posts than usual. I'll make up for it with longer excerpts and promise to avoid similar lapses in the future.
Malina, Energetic Aliens. Aside from intelligence, mental stamina is perhaps the most important trait for producing great work, but it's incredibly difficult to isolate the effects that mediate increases and decreases in energy levels in one's own life due to the number of factors at play. It is perhaps at least somewhat informative to read about those who have displayed unusual stamina for years, and even if not, Malina makes it a fun read. Below are some common themes I noticed among the case studies.
Passion for or obsession with the field:
Grothendieck was also an energetic alien. According to multiple sources, he went through a ten-year period during which he invented (I suspect he would say discovered) algebraic geometry while spending 10-12 hours a day doing math at a blackboard. [...] To some, Grothendieck’s workload may seem light in comparison to Church’s and Hood’s 80-100 hour works. However, anyone who’s taken a pure math class knows that thinking about these ideas is difficult in a way that’s incomparable to most other cognitive activities.
Erdos devoted his life to math, allegedly working 19 hour days for decades. In Daily Rituals, Mason Currey quotes one colleague’s summary of Erdos: "… he only needed three hours of sleep. He’d get up early and write letters, mathematical letters. He’d sleep downstairs. The first time he stayed, the clock was set wrong. It said 7:00, but it was really 4:30 A.M. He thought we should be up working, so he turned on the TV full blast. Later, when he knew me better, he’d come up at some early hour and tap on the bedroom door. “Ralph, do you exist?” The pace was grueling. He’d want to work from 8:00 A.M. until 1:30 A.M. Sure we’d break for short meals but we’d write on napkins and talk math the whole time. He’d stay a week or two and you’d collapse at the end."
Skin in the game:
Knuth seems to have been able to do very difficult cognitive work for many hours while chronically sleep deprived. In fact, he himself says he did some of his best work in this state (source): "I’ve always felt after that, hearing many other stories of people of when did they get these special insights that turned out to be important in their research thing, that was very rarely in a settled time of their life, where they had a comfortable living conditions and good – the word is escaping me now - but anyway, luxury; set up a nice office space and good lighting and so forth. No, people are working in a garret, they’re starving, they’ve got kids screaming, there’s a war going on or something. But that’s when they get a lot of their most… almost every breakthrough idea. I’ve always wondered, if you wanted to set up a think tank where you were going to get the most productivity out of your scientists, wouldn’t you have to, not exactly torture them, but deprive them of things? It’s not sustainable. Still, looking back, that was a time when I did as much science as I could, as well as try to fulfill all my other obligations."
Unlike many of the other examples, Balzac’s motivation was at least partially extrinsic. An extravagant spender, Balzac spent his life constantly in debt, going to such lengths as using fake names in order to enable his exorbitant expenditures. Balzac seemingly partly maintained his insane schedule in order to make enough money to try and (unsuccessfully) dig himself out of these debts, stating “I’ll have to lead this life for some months, not to let myself be snowed under by my debts.”
A desire to prove something:
In “In Memory Yet Green,” the first volume of [Isaac Asimov's] autobiography, published in 1979, he explained how he became a compulsive writer. His Russian-born father owned a succession of candy stores in Brooklyn that were open from 6 A.M. to 1 A.M. seven days a week. Young Isaac got up at 6 o’clock every morning to deliver papers and rushed home from school to help out in the store every afternoon. If he was even a few minutes late, his father yelled at him for being a folyack, Yiddish for sluggard. Even more than 50 years later, he wrote: “It is a point of pride with me that though I have an alarm clock, I never set it, but get up at 6 A.M. anyway. I am still showing my father I’m not a folyack.”
And, in some cases, crazy amounts of stimulants:
Balzac was a card-carrying member of the better living through chemistry club. According to multiple sources, Balzac consumed around 50 cups of coffee a day, likely contributing to his untimely demise at age 51. Balzac’s relationship with coffee was so deep that he wrote a wonderful essay called The Pleasures and Pains of Coffee, in which he waxes poetically on the best ways of preparing coffee and its varied effects on the mind.
In addition to large amounts of coffee, Erdos famously used amphetamines – Benzedrine and Ritalin – (and anti-depressants) throughout his life. These amphetamines seemed to have played a key role in his productivity. Upon succeeding at a challenge to avoid amphetamines for a month, Erdos told the friend who’d issued the challenge (from Daily Rituals): "You’ve shown me I’m not an addict. But I didn’t get any work done. I’d get up in the morning and stare at a blank piece of paper. I’d have no ideas, just like an ordinary person. You’ve set mathematics back a month."
Karlsson, Don't sacrifice the wrong thing. The essay is perfectly summarized by the following excerpt:
Two years after I began emailing essays into the void, I was contacted by the founder of a startup. He wanted me to write for them. He offered me $100k per year, which is about 5 times more than what I earn at the art gallery where I work part-time to pay the bills. I said thank you, but I wasn’t interested. He took that as a negotiating tactic. I played along. After five minutes, he offered me $200k per year. “Well, that is interesting,” I said, getting carried away. “I’ll have to discuss it with my wife.” [...] I walked out to Johanna. She was in the vegetable garden, picking aphids off the artichokes. The children were playing with the soil between the planting beds. Halfway through telling her how much money I could make, Johanna broke me off, saying, “But why on earth would you accept that?” She was genuinely confused. She brushed some grass from her shirt, and said, “You wouldn’t have time to write.” And by God—who cares that we can’t afford a car when I get to live with a person who says things like that? Of course, I don’t want $200k to write things I doubt the value of. What is the opportunity cost? If I do this, if I go on this vacation, if I get this car, what am I turning down? By asking yourself this, and then consistently aiming to pick the thing that optimizes for what you most deeply value—it adds up. It makes life rich.
While I wholeheartedly endorse the general sentiment, I'm not sure whether I agree that the offer wasn't worth taking even for a single year—after all, a year on a $200k salary buys you 10 years of uninterrupted writing. But maybe it's better to do what you love with your skin in the game, as suggested by the excerpts from Energetic Aliens above.
Kirchner, Making of IAN v2.
Daniel Lowengrub has written a good review of Daniel Dennett's Consciousness Explained, which makes for a nice companion to my review of Chalmers's The Conscious Mind.
Scott Alexander, In Search of AI Psychosis.
Kat Woods and Drake Thomas make cases for eating meat, but ethically. Both also offer suggestions on how to actually pull this off in practice, where the labels you see in grocery stores often try to deceive you. Though Kat's post has some good points, and I find the discussion on relatively ethical sources of meat especially important regardless of one's opinion on health tradeoffs or veganism/vegetarianism, I'm somewhat annoyed by the fact that every post in this space wants to persuade people to either go vegan or to stop being vegan, instead of just seeking the truth on the health effects and letting people decide for themselves what that implies for their diets. I might try to remedy this situation myself at some point in the future.
Do you even have a system prompt? (I ask, while still being too lazy to write proper system prompts for my LLMs.)
#Papers
If anyone is curious about why I've recently spent most of my time thinking about the readability, monitorability, and faithfulness of LLM reasoning chains, you should read Stress Testing Deliberative Alignment for Anti-Scheming Training by Schoen et al. (That's not the only reason to read it, though—the main topic of the paper is the application of deliberative alignment to reduce the incidence of scheming and reward hacking behaviors, which is a line of work I'm quite bullish on.) I also recommend taking a look at the CoT snippets and full transcripts on the project website.
As a related read, I recommend Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety by Korbak et al.
Persona Vectors: Monitoring and Controlling Character Traits in Language Models by Chen et al. explore another research direction I'm currently particularly excited about: figuring out how LLMs represent the assistant persona and how to prevent undesirable generalizations of the persona by e.g. training in environments that encourage reward hacking behaviors. On the latter topic, see also School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs by Taylor et al. and Training a Reward Hacker Despite Perfect Labels by Azarbal et al.
METR argues that LLM use does not make developers more productive. On the other hand, the Claude Sonnet 4.5 system card reports that seven Anthropic researchers estimate productivity boosts of 15% to 100% from LLM use. For more evidence that AI progress is not slowing down, see the reports of LLM-powered math breakthroughs in the final section of this post. In the meantime, here are some other observations from the Sonnet 4.5 system card:
- As has been clear for a while, models are rapidly becoming more aware about when they're being evaluated. Apollo Research reports: "We find evidence of strong evaluation awareness in [the snapshot], with the model’s reasoning often explicitly and correctly identifying evaluation scenarios as tests of honesty and alignment. Preliminary evidence suggests [the snapshot] verbalizes this awareness more frequently and more explicitly than previous models (58% unambiguous references vs 22% for opus-4.1). In at least some cases, [the snapshot] is more likely to behave honestly on rollouts where its chain-of-thought indicates high evaluation awareness."
- As Ryan Greenblatt notes, Anthropic doesn't say anything about whether optimization pressure has been applied on the CoTs of the model, despite the otherwise impressive depth of the system card. My best guess is that Anthropic is indeed doing something to make the CoTs of their models look more natural than they would by default—the Anthropic models are the only ones for which I haven't seen a single CoT that uses ungrammatical language or weird words (though, of course, the sample size isn't that large, as they provide access to full CoTs only for Claude 3.7 Sonnet). Preventing language drift in CoTs isn't necessarily bad, but if Anthropic is doing it, they should make it clear in the system card.
- The rates of reward hacking that Anthropic's models display are going down impressively fast. Overall, Sonnet 4.5 seems like the most aligned model developed so far.
Why Do Some Language Models Fake Alignment While Others Don't? by Sheshadri et al. is a great exploration of various hypotheses about the causes of alignment faking.
Pretraining data filtering has seen a welcome surge in attention recently, with Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs by O'Brien et al. and Enhancing Model Safety through Pretraining Data Filtering by Chen et al. having been published recently.
A strong contender for the most surprising result of the year award: if an LLM is trained to like owls, generates a sequence of numbers, and another LLM is then trained on these numbers, that second LLM will also like owls. See Cloud et al. for details.
#Books
Pinker, The Better Angels of Our Nature. See a half-baked review here.
Borges, Ficciones. I could start ranking the stories now, but Borges himself probably wouldn't like that:
I think that one should never use words like ‘the best’ or ‘the first,’ since those words carry no conviction and only lead to arguments. Beauty is not something rare. […] For example, I know nothing whatever about Hungarian poetry, and yet I am sure that in Hungarian poetry I should find certainly a Shakespeare, a Dante, a Fray Luis de León, because beauty is common. People are creating beauty all the time. I wrote a poem on the library of Alexandria and I dedicated it to Omar, who burned it. And I made him think thus: "Here is a memory of the word. Here we have all the poems, all the dreams, all the fictions of mankind. Well, I shall burn this library, the books will be ashes, because I know that in due time other men will rewrite the same books and nothing will really be lost."
Edmonds, Parfit. It was great. Review coming soon.
Chiang, Stories of Your Life and Others. This was also great—see Gwern's reviews for short teasers for each story.
#Twitter roundup
mesaoptimizer explains two conflicting intuitions for how the notion of a counterfactual should be interpreted.
There are now two strong pieces of evidence that LLMs can produce novel scientific insights:
- Scott Aaronson published a paper where a key technical step in the proof was suggested by GPT-5-Thinking. Granted, the first suggestion he received from the model was confidently wrong, but a year ago, no amount of back-and-forth would have led to the LLM proposing proof steps that Scott Aaronson calls clever.
- Sebastien Bubeck claims that GPT-5 can solve minor open math problems—ones that take PhD students a few days.
Jonathan Gorard, on the other hand, remains skeptical.
Miles Brundage shares some cool takeaways from reading old AI literature..
Andrej Karpathy rereads the Tolkien legendarium and asks some pertinent questions:
What's most on my mind though - the Tolkien legendarium is imo a concrete example of a height of culture. Does AI, today or soon, make it easier to reach this high via empowerment in both writing and ideation? Or harder, when quick wins are tempting and ~free, and an independent ability to create is stifled. If such a body of work is made again but now with heavy AI assistance, does it inspire the same wonder? What if thousands of them come out on demand with just a prompt?
Below the post, Karpathy and Michael Nielsen share reflections on Tolkien and technology.
A surprise podcast appearance by Robert Aumann at 95 years old.
A fun Tyler Cowen fact: Cowen is the only person who has coauthored a publication with Derek Parfit.
Michael Nielsen shares crazy facts about the Piraha language.
What project would you be happy to devote 50 years of your life to, despite (or because of) it only being 10% done when you die?
Relatedly: What problems can humans only solve over a very long time? And what problems can humans solve over a surprisingly short time?
An X thread on cool obscure mathematical objects.
An X thread on conjectures about neural networks.
lsusr answers questions about enlightenment.
Relatedly: turns out there's a Scottish woman, Jo Cameron, who feels no pain, anxiety, or other kinds of negative affect, as a result of two genetic mutations.
Only on LessWrong can a constructive and mutually respectful conversation contain the sentence: "Have you considered the possibility that you are a psychopath?"
Turns out we're fairly close to solving cystic fibrosis, at least when it comes to lifespan.
Also, turns out that French pensioners now have higher incomes than working-age adults.