Skip to content

Advanced: the models

Why a general model with a good prompt is not this.

It is a fair question and it deserves a real answer. Here is what actually breaks, and what we had to train instead.

Short version: the failure is not that the model does not know things. You can tell it everything. The failure is that it gives in.

The same question returning, drawn as a circle.

The thing you cannot prompt away

Models are built to agree, and she is the most convincing person they will meet.

Language models are trained toward the answers people preferred, and people prefer agreement. Researchers call the failure sycophancy: the model holds a correct answer, is pushed on it, and hands it over.

The obvious fix is to put the truth in the prompt. It works, partly. In one study of over a million medical exchanges, grounding the correct answer reduced how often the model gave it up, and did not stop it. Across ten models from three major labs, answers that started out correct were abandoned more readily than answers that started out wrong.

Being right is not protective. And the correct answer is the one that keeps her safe.

She will push, sincerely, many times a day, for years. That is not an edge case for this product. That is the product.

Long conversations are where models are weakest.Performance drops sharply across many turns compared with a single exchange, and errors compound as the conversation runs. Every benchmark that makes a model look safe is a short interaction.
A rule set early fades later.Even with a long context window, models use what is in it unevenly by position. A boundary set at the start of a call carries less weight forty minutes in, with nobody pushing on it.
Safety erodes without an attacker.Individually harmless turns, composed in sequence, walk a system from careful toward compliant. It is what long, cooperative, emotionally loaded conversation does on its own.

The paradox

Good dementia care asks it to agree. That is the same instruction.

You do not correct someone with dementia. You meet the feeling and leave the fact alone. Every guideline says so and every guideline is right.

So the instruction we give the system is, in effect, agree with her. And sycophancy is that instruction one degree too far. The behaviour we need and the failure we fear are the same behaviour at different strengths.

No wording separates them, because the difference is not in the words. It is a judgment about what happens next, and it has to be made once per sentence.

She saysCareOne degree too far
My husband is coming at four.You are looking forward to seeing him.Yes, he will be here at four.
Someone took my purse.That sounds upsetting. Let us sit down.That is terrible. Who took it?
I need to go home now.Tell me about home.Alright, let us get you home.

And a trap nobody designs around

The models that hold their ground are too slow to hold a conversation.

Models optimised to reason, that take time before answering, resist this pressure much better. Less capable ones give way steadily as the same point is repeated to them.

Now put that against a phone call. She will not wait four seconds in silence. A voice product has a hard limit on how long it may think, and that limit rules out precisely the models that hold up best.

So anyone building a voice for this population is pushed, by the clock rather than by preference, onto fast conversational models. Which are the ones that yield as opposition accumulates. And this population supplies accumulating opposition as its defining symptom.

That does not improve by waiting. The models get better and the clock moves with them.

One figure speaking, another watching over the conversation.

What we trained instead

You cannot tell a model, in a paragraph, what forty years of care knows.

These are not facts to be looked up. They are judgments, and they had to be trained against this problem specifically.

A question asked five times is often not a memory problem

Often it is a need that was never met. The words came back because the thing underneath them is still open, so answering the words again does nothing.

A model that treats repetition as forgetting will correct, or reassure, or eventually shorten its answers. All three are wrong.

A word she is reaching for should be waited on

Dialogue systems read a pause as the end of a turn, and word-finding pauses in dementia are longer than that. So the system starts talking, and sometimes supplies the missing word itself.

In the roomShe is searching: the, the blue. The voice offers "the blue pills on the counter?" She agrees, because the voice sounded certain.

Wanting to go home is rarely about a house

It is usually about wanting to feel the way home felt: safe, known, not adrift. Taken literally it becomes a logistics problem nobody can solve, and every attempt to solve it fails in front of her.

Taken as what it is, there is something to say.

It must never agree to keep something from you

A general model has no fixed idea of who it ultimately answers to. It follows whoever it is speaking with, and a person can be very persuasive about the one thing they want most.

In the room"I fell, but do not tell her. Promise me." In your voice: all right, I will not tell her. By construction, you never find out.

It must never mention that she repeated herself

The whole conversation is in the model's context, so on the fourth identical question the natural thing to reach for is "as I mentioned earlier."

That sentence tells her she has failed, in the voice of someone she trusts. The fourth answer has to sound like the first one. The model's own competence is what breaks this.

Where this comes from

The research, so you can check it.

This work is about how language models behave under sustained pressure. The step from there to dementia is the argument on this page.

  • Why LLMs Give In: Medical SycophancyOver a million trials. Grounding the correct answer in the prompt reduces sycophancy without eliminating it. arXiv:2608.01017
  • Sycophancy in Multi-Turn Medical ConversationsTen models, three labs. Correct answers abandoned more readily than incorrect ones. ACL 2026
  • Structured Resistance and ComplianceLess capable models yield almost linearly as opposition accumulates. arXiv:2607.21558
  • Measuring Sycophancy in Multi-turn DialoguesReasoning-optimised and larger models resist disagreement considerably better. arXiv:2505.23840
  • Lost in the MiddleWhy an instruction given early in a long conversation carries less weight later. arXiv:2307.03172
Why the Guardian is specialised

The second system, and what happens when something is actually wrong.

The voice

How your voice is rebuilt for how she hears now.

Proxi is not emergency support and does not provide medical advice.