The first turn has no history
The demo always goes well. You open a fresh conversation, send three messages, and the character is exactly what you wrote — the voice is sharp, the constraints hold, the mannerisms land. Two weeks later a user says it feels like a different product, and you cannot reproduce that on your machine because your machine keeps starting new conversations.
The first turn is not a representative sample of the system. It is the best case, and it is the only turn where nothing competes with the definition.
The first turn is the maximum, not the baseline
On turn one the assembled input is the definition plus one user message. There is no carried history, no summary, no accumulated facts. The proportion of the input that describes the character is as high as it will ever be, and it falls monotonically from there.
That means every quality judgement made on a fresh conversation is a measurement of the ceiling. A character that is barely distinctive at turn one has no headroom at all, because the only direction available is dilution. A character that is unmistakable at turn one might still be unrecognisable at turn two hundred, and nothing in the fresh-conversation test tells you which.
The practical consequence is that cold-start quality and long-conversation quality are two different measurements that people habitually take as one. Depth is the axis that gets skipped, and it gets skipped partly because the shallow test always looks good.
There is a second-order effect worth naming: how strongly a short definition expresses itself with no history to condition on is one of the places behaviour genuinely varies between models. A definition that reads as vivid on one model reads as a list of adjectives on another, and the difference is largest exactly at turn one, where there is nothing else in the input to push back.
The over-expression problem
Maximum share is not automatically desirable, which is the part that surprises people.
With nothing else in the input, a definition tends to be applied maximally: every trait in it shows up in the first few replies, often all at once. A definition listing six mannerisms produces a first reply containing six mannerisms. It reads as caricature, and it is the same definition that will read as convincingly restrained by turn fifty because history has diluted it into something closer to occasional.
So the first handful of turns and the long middle of a conversation have opposite failure modes. Cold start over-applies; depth under-applies. Tuning a definition against one makes the other worse. If you lean the definition down until turn one stops reading as parody, you have removed the material that was going to carry turn two hundred; if you write it to survive depth, the opening reads as a costume.
Naming that as a tradeoff rather than a bug is the useful move, because there is no definition that is correctly weighted at both ends. What you can choose is which end to optimise and how much extra input to spend correcting the other.
Where the first relationship facts come from
The other half of cold start is that there are no relationship facts, and relationship facts are a different category of state from the definition. Definition is authored and shared; relationship facts accumulate per conversation and diverge from the first turn onward.
At turn one that store is empty, and an empty store is honest. The temptation is to fill it — to seed a setup step’s answers into the same place accumulated facts live, so the opening has something to work with. That is a defensible design, and it has one hard requirement: facts supplied at setup and facts observed during a conversation must be distinguishable in the record, because they have different reliability and different update rules. A stated preference is a declaration. A preference inferred from three messages is a guess. Writing both into one undifferentiated blob is how a guess acquires the authority of a declaration, and a wrong stored fact is considerably more damaging than a missing one.
The failure to avoid entirely is fabricating shared past — writing facts that assert events which never occurred, so the opening can reference them. It reads well for exactly one turn and then it is permanent state that contradicts everything the user actually remembers, with no mechanism to correct it because nobody wrote it down as fictional.
The turn
THE TURN — cold start
· Definition alone, no history
→ maximum expression. The sharpest the
character will ever be, on the only
turn where nothing competes with it.
· Maximum share applies traits maximally
→ the opening over-applies. A definition
tuned to read well at turn one is
thinner than depth needs.
· Empty relationship-fact store
→ honest, and unhelpful. Nothing to
condition an opening on but the
definition.
· Seeding setup answers as facts
→ gives the opening something to work with,
and mixes declarations with inferences
unless provenance is stored per fact.
· Every fact seeded at setup
→ PAID EVERY TURN. It joins the carried
state and is re-supplied on every input
from then on, so cold-start generosity
is a recurring bill.
· Fresh-conversation testing
→ surfaces no depth failure at all.
Passing it is not evidence.
· Strength of expression with no history
→ VARIES BY MODEL. The gap between models
is widest at turn one.
Detection
Score the same probes at turn one and at depth, and report both numbers. A single consistency score computed over fresh conversations is a ceiling measurement wearing the label of an average. Two numbers make the decay visible without anyone reading a transcript.
Log the definition’s share of the input by turn index. At turn one it is nearly all of it, which gives you the natural reference point for the ratio you actually care about later. The curve from there is the drift pressure, and the mechanism it drives is arithmetic.
Count facts in the store at turn one, per conversation. If that number is not zero, something is seeding it, and you want to know what and with what provenance. A non-zero cold-start fact count that nobody designed is usually a setup step writing into the accumulated store.
Watch first-reply length and trait density. Over-expression shows up as an opening reply markedly longer and busier than the same character’s replies at turn thirty. That gap is measurable and it is the caricature symptom in numeric form.
What this costs and what it doesn’t fix
Every gram of cold-start help is permanent per-turn weight. Seeded facts, a longer opening definition, an extra re-anchor for the first exchanges — all of it lands in the assembled input and stays there, so the opening is not a free surface to spend on.
And none of this makes the first turn informative about the hundredth. Cold start is a separate regime with its own failures, and the only thing it reliably tells you is whether you have any headroom before dilution starts. If the character is already marginal with the whole input to itself, no memory strategy downstream will recover it.