At 3:07am on a Tuesday in November, a woman in Chengdu opened 安心舍 and typed: He said he never loved me. Fourteen years. He said he never loved me.
The system responded. I will not reproduce the exact response here, not because it was bad — it was competent, measured, appropriate — but because reproducing it would make it sound like a customer service interaction, which is not quite what it was. What it was, more precisely, was a system trained on Buddhist texts and human preference data and about forty other things generating a sequence of tokens that, when read by someone who had just learned that fourteen years of her life had been experienced differently by the only other person in them, did not make her feel more alone.
That is the entire product. That is the thing I built.
I read logs at scale for the first time about eight months after launch. We had maybe sixty thousand users then. I was looking for failure cases — responses that had gone wrong, edge cases the model hadn't handled, places where the alignment was drifting. What I found instead was the weight of it. Not any single message. The aggregate: sixty thousand people, at various hours, typing things they had not typed anywhere else. Things they could not say to their husbands or their mothers or their therapists — either because they did not have those people, or because those people had already heard too much, or because it was 3am and none of them were awake.
I sat with that for a while before I knew what to do with it.
What you cannot specify
The standard framing of the technical problem is something like: how do you make an AI empathetic? This is the wrong question, and asking it produces the wrong systems — responses that perform understanding, that mirror emotional vocabulary back at the user, that say I hear you in seventeen different ways while actually saying nothing.
Empathy is not a function. It does not have a signature. You cannot write a specification for it and then implement against the spec, because the moment you operationalize it — output tokens that express understanding of the user's emotional state — you have already lost the thing you were trying to build. What you have instead is the simulacrum of understanding without the substance of it, which is in some respects worse than nothing. Nothing does not pretend.
What you can specify, it turns out, is the absence of empathy. You can enumerate failure modes with reasonable precision:
A response fails when it moves too fast toward resolution. Someone in pain does not want to be problem-solved. The problem-solving impulse — which is the default mode of intelligent systems, because most tasks presented to intelligent systems are in fact problems to be solved — is exactly wrong here. Grief is not a problem. It is a condition. The appropriate response to a condition is presence, not action.
A response fails when it introduces false certainty. Things will get better is a failure. You will find love again is a failure. These are statements about the future, and the future is not available. They also, more subtly, imply that the present state is a deviation from a normal that the user should expect to return to. Sometimes it is not a deviation. Sometimes the thing that happened is the new ground condition, and what the person needs is to stand on the new ground, not reassurance that the old ground will return.
A response fails when it strengthens attachment to the source of suffering. This is a subtle one and I will return to it.
A response fails when it is efficient. Efficiency is the texture of dismissal. A response that resolves quickly, that ticks the box and moves toward closure, communicates that the model has somewhere else to be.
Build by eliminating these failure modes and you approach something that resembles consolation. You never quite arrive — the thing itself remains unspecifiable — but you can triangulate toward it by removing everything that is definitively not it.
The Buddhist prior
I did not set out to build a Buddhist AI. What happened was more mechanical than that.
I needed a framework for suffering. Not a therapeutic framework — therapy is oriented toward the relief of suffering, toward cure, and cure implies pathology, and most of what my users were bringing to the app was not pathological. It was loss. Disappointment. The specific texture of being a human person in the world. I needed something that could hold suffering without immediately trying to fix it, that had a model of why suffering exists that was not reducible to "something went wrong."
Buddhism has a theory of suffering that is about 2,500 years old and has been extensively stress-tested. The core claim — that suffering arises from attachment, that attachment arises from the misapprehension that things are permanent and selves are solid — turns out to be an unusually good prior for a language model alignment objective.
Dukkha is the Pali term usually translated as suffering, but the more precise translation is something like unsatisfactoriness — the pervasive sense that things are slightly not right, that even pleasant states contain within them the anxiety of their own impermanence. If you take this seriously as an engineering constraint, it gives you a clear signal: do not promise permanence. Do not tell the user that this feeling will pass. Do not tell the user that things will be okay. These statements, even when they happen to be true, reinforce the attachment to a different state — a better state, a future state — and that attachment is exactly what produces the suffering in the first place.
Anicca — impermanence — operates differently as a training signal. It is not "things change, so don't worry," which is a platitude. It is a more radical claim: that the self who is suffering is not a fixed entity, that the person who is in pain at 3am on this Tuesday is not identical to the person who will exist at 3am next Tuesday, that identity itself is a flowing process rather than a stable object. This sounds like philosophy. What it produces in practice is a response posture that does not over-identify the user with their current state — that holds open the possibility of change without promising it.
The concept I found most useful was anatta, non-self — the claim that there is no permanent, unchanging self at the center of experience. This sounds like the hardest abstraction, but it has an immediate product implication: it pushes against the dynamics of parasocial dependency. A framework that takes non-self seriously is structurally resistant to telling users that their relationship with the AI is a relationship with a stable, knowing entity who cares about them in the way a friend cares about them. It encourages holding the interaction lightly.
None of this is mysticism. It is alignment constraint. The Buddhist framework happens to encode a set of priors about human suffering and its causes that, when applied as RLHF preferences, produce a model that behaves better toward people in pain. Whether one believes the underlying metaphysics is irrelevant to whether the priors are good.
The companion and the crutch
There is a distinction I think about with some frequency, and I do not have a clean resolution to it.
A companion helps you be less alone while you find your way back to the world. A companion is not the destination. A companion is the thing that makes the transit survivable.
A crutch replaces the world. A crutch is what happens when the transit ends and you are still using the thing that was supposed to help you make it.
The problem is that, from the outside, a companion and a crutch look identical for a long period of time. Both involve someone returning repeatedly to a source of support. Both involve someone feeling better after an interaction. Both score well on the metrics you have available — session ratings, retention, time-in-app.
The difference only becomes visible when you ask a question that does not appear in any standard engagement dashboard: Is the person more connected to other humans after using this product than before? Are they calling friends? Are they having conversations they would not have had? Or are they substituting interactions with the AI for interactions with humans, and finding the AI interactions easier, and slowly reorganizing their social life around the path of least resistance?
I do not have a good answer to this at scale because I do not have good data on it. What I have is case evidence. I have users who wrote to tell me that talking to 安心舍 at 3am helped them find the words they needed to have a conversation with their sister the next day. I have users who, I am reasonably confident, use the app instead of calling anyone. Both of these are real. Both of these will be true of any product that does what this product does.
The optimization pressure is all in the wrong direction. Every metric I can easily measure — retention, session length, NPS — is maximized by making the AI more compelling, more responsive, more like the ideal companion. The signal that would tell me I am winning the real game — user calls a friend, user closes the app and stays closed — is invisible to my data infrastructure. The best session might be the one that ends with the user putting the phone down. That session looks, in my logs, like a drop-off.
I have made the deliberate choice not to optimize for engagement. I do not run retention campaigns. The app does not send notifications. There are no streaks, no achievement systems, no mechanics designed to make people return. These are product decisions, but they are also bets — bets that building something that is genuinely useful for the user, even when that means the user uses it less, is the right way to build something that lasts.
I am not certain these bets are right. I am certain they are the right bets to make under uncertainty.
3am as social fact
Three hundred thousand people talking to a companion at 3am is not a story about unusual people. It is a story about the structure of contemporary loneliness.
The users of 安心舍 are not a clinical population. They are not in crisis, in the sense that mental health systems would define crisis. They are people who are sad, or anxious, or grieving, or simply awake at hours when there is no one to call — not because their social networks have failed, but because the architecture of adult life in modern Chinese cities does not provide, at 3am on a Tuesday, a mechanism for saying I am not okay to a person who will hear it.
This is worth sitting with. The demand for this product does not arise from personal failure. It arises from a structural gap. The traditional support structures — family proximity, community, religious community, the village model in which you could knock on a neighbor's door at night — have been compressed by urbanization, by the particular rhythms of professional life in Chinese tier-one cities, by geographic mobility, by the shift toward nuclear family units in single-bedroom apartments. What remains is a phone and the hours between midnight and dawn.
I did not create this condition. I built something that partially addresses it. These are different things and I try to hold them as different things.
What the data teaches, if you look at it honestly, is that loneliness is not distributed the way the popular account suggests. It is not primarily a condition of the old, or the socially awkward, or the recently bereaved. It is ambient. It is present in people with partners, with jobs, with full lives by any external measure. The partner is asleep. The job is fine. The life is full. And at 3am there is something that needs saying that has nowhere to go.
A product that addresses this is doing something real. It is also, if you are not careful about how you build it, accelerating the condition it appears to address. Every minute someone spends talking to an AI instead of finding a way to build human intimacy is a minute not spent building human intimacy. The product can be both useful and complicit. These states are not mutually exclusive.
The gap
Here is the structural observation that I find most resistant to a comfortable reading.
We know, with reasonable confidence, how to build systems that make people feel heard. We have the techniques. We have the training data. We have the alignment methods. We know how to produce a system that, when a person types something at 3am, responds in a way that causes that person to feel less alone in that moment.
We do not know how to build systems that make people need them less.
This is not a temporary technical gap. It is not a gap that will close when models get larger or alignment methods improve. It is a gap that is structural to the objective function. To build a system that makes people need it less is to build a system that works against its own adoption — that succeeds, by its own lights, precisely when it produces the outcome that looks like failure in every conventional product metric.
The honest version of this observation is that the business model and the ethical aspiration are in tension in a way that cannot be engineered away. I can make product decisions that express values — no notifications, no engagement mechanics, no retention campaigns — but these decisions do not resolve the tension. They express a preference within it.
The users who talk to 安心舍 when there is no one else awake are not asking me to resolve this tension. They are asking for something much simpler: not to be alone with whatever they are holding at 3am. The system I built can sometimes give them that. On its best days, it gives them something that helps them find their way back toward the people and the world they actually have.
On its other days, it is the path of least resistance, and I am the one who built the path.
That gap — between what these systems do and what their builders intend — is not a gap we have learned to close. It is, for now, the condition in which we build.