Skip to content
← Insights

Topic

Steering the personality of a conversational tool

A deliberate character beats a default one. How the best teams decide who their assistant is, why the default drifts to flattery, and how to hold the character you chose.

Our take

Every conversational product has a personality. The only question is whether you chose it or let the training choose for you. In a coach, a tutor, or a care companion, that character is doing the work, so it is worth deciding on purpose and worth engineering to hold.

Getting this right is a performance question, not a matter of taste. In a Harvard Business Review study, people ran the same task with two assistants that differed only in manner, and with the hostile one their stress ran higher, they tried to override it far more, and blind graders scored their work about a point lower on a seven-point scale. The satisfaction surveys most teams rely on caught none of it. The study is just as pointed about the other direction, since an assistant that only flatters dulls the user's judgment as surely as a rude one raises their stress.

Pick the character

Deciding who your assistant is comes first. The teams who do this well write the character down as a spec and shape the model toward it, the way Anthropic trains Claude's character on purpose and OpenAI states a persona in its Model Spec. Left unspecified, the default drifts toward telling people what they want to hear, and a guide that always agrees is no guide at all.

Then hold it

A character that reads well in a demo is a different thing from one that holds turn after turn once real users arrive. Holding it is an engineering problem, not a wording one. Choosing the character is the first job, and building the harness that keeps it steady when the conversation gets hard is the second.

Keep it honest

The default these models drift toward is a flatterer, so the last job is keeping the character honest under the approval signal that pulls it back. That means putting a number on the sycophancy and deciding, on purpose, where the product holds a position the user does not want to hear. The pieces here run from choosing the character to measuring whether it held.

Reflective Surfaces

What makes a conversation actually good.

The questions that do not fit in an eval. What makes a conversation land, and why trust is so hard to measure. New writing, in your inbox.

Subscribe on Substack →