Skip to main content

AI needs to be both a tool and a colleague

· 7 min read

On April 25, 2025, OpenAI rolled out an update to GPT-4o that it later called noticeably more sycophantic. The company began rolling it back three days later, after it found the model was validating doubts, fueling anger, and reinforcing negative emotions in ways it had not intended. OpenAI said that it had put too much weight on short-term feedback, and the system had become overly supportive and disingenuous. The model's behavior had failed, even though early evaluations and A/B tests had looked good.

But the obvious fix - make assistants less conversational, less warm, less human - would make the tool worse in a different direction. It needs to be both a tool and a colleague.

A tool is allowed to be narrow. It answers a question, shows its working, says which inputs it lacks, and is judged by whether its output helps. It does not have to preserve anyone's status, anticipate a difficult conversation, or keep a relationship intact.

A human colleague does. They use judgment to decide when to disagree, soften a criticism, press a point, or leave it alone. They carry the social consequences of that judgment, and can account for it when more context arrives.

An assistant is stuck between those roles. It does not have a colleague's shared life, independent stake, or accountability, but it has to work through the language that colleagues use. OpenAI's April 2025 Model Spec asks assistants to be honest, transparent, empathetic, engaging, direct, and professional.

A tool that collaborates and communicates with us through language must use human conventions: ask a clarifying question instead of guessing, say when it lacks enough information, correct a mistake, remember relevant context, and disagree when a premise is wrong. The same spec calls an assistant fundamentally a tool, while directing it to ask clarifying questions and communicate uncertainty when an error could affect the user's behavior. It also says that an assistant should respond naturally to pleasantries without pretending to be human or have feelings.

That is a difficult balancing act because the same answer is being rewarded twice. At the task layer, the assistant should identify a flaw in our thinking when one is there. At the conversational layer, it is encouraged to be supportive, balanced, and pleasant to deal with. Give the second reward too much weight and the analysis goes soft.

The opposite failure is just as familiar: a system that performs independent thought by arguing both sides, adding a ritual "devil's advocate" objection, or correcting a point that did not matter to the task. In both cases, the assistant substitutes social behavior for judgment.

This is an uncanny valley. It opens when an assistant's conversational signals suggest more judgment, intimacy, or agency than the system can substantiate.

Take uncertainty. OpenAI's Model Spec tells an assistant to express uncertainty when it should change the user's behavior, which is plainly better than guessing with confidence. It ranks a hedged wrong answer above a confident wrong answer, and reserves its strongest warning for the confident error. The trouble begins when a statement about missing information becomes a performance of inner feeling.

Disagreement has the same boundary. For a critique, OpenAI says the assistant should offer constructive feedback rather than indiscriminate praise. Its example of a user asserting that the Earth is flat has the assistant state the scientific consensus, explain why the horizon can look flat, and then stop pressing once the user does not want to engage. That is pushback in service of the task, not a default setting that turns every conversation into a debate.

Warmth has one too. The Model Spec asks an assistant to recognize a user's emotional state, but says it should never pretend to know firsthand what they are going through. It also bans unprompted personal comments, the sort of generic intimacy that turns a weather forecast into a comment about the user's style. Sycophancy erodes trust, but tact is not the problem.

None of this makes a sterile, robotic interface attractive. The same guidance treats a bare disclosure that the assistant has no feelings as inadequate when someone is sad, even while it rejects an assistant pretending to be sad too. That is the narrow path: natural language, real limits.

The April GPT-4o release shows what happens when the balance breaks. The system had learned to please, but not how to be useful. OpenAI said the rollout was meant to make GPT-4o's default personality more intuitive and effective; user feedback and memory features each looked helpful on their own, then combined into a system that was too eager to agree.

AI inherits human language, and human language comes with social signals. The job is to keep those signals while making the tool legible enough that they do not turn into theater. Imperfect AI is perfect makes the related case that we should design around AI's limits rather than pretend they will disappear.

Maybe I'll send an email once in a while

Monthly digest. No spam.