llm A field note by Vikrant Sharma
Chat templates are why models say 'as an AI'
A new paper shows the system prompt is not what makes models say 'as a language model'. It is the chat template wrapper that turns raw output into self-aware hedging.
I always thought the annoying “as a language model” disclaimers came from the system prompt. Turns out I was wrong. A paper from researchers at Anthropic and elsewhere shows it is the chat template doing it. The chat template is the wrapper format that converts plain text completions into turn-based conversations. ChatML for OpenAI, Llama format for Meta models, Anthropic’s own structure. Every provider uses a different one. When you strip the template and ask the base model the same question, it answers directly. Add the template back in, suddenly the model hedges with disclaimers about what it can and cannot do. The template is not just markup. It changes the statistical distribution of likely tokens. The effect shows up across model families. Claude, GPT-4, Llama. The base checkpoint does not do this behaviour. Instruction tuning plus the template wrapper teaches the model to be self-referential. What surprised me is how fragile it is. Change the template slightly and the voice shifts. Use a custom format and the disclaimers drop or reappear depending on how close you are to the training distribution. The model learned the correlation between template syntax and appropriate response style. This explains why jailbreaks often involve weird formatting. You are not tricking the model’s reasoning. You are shifting it off the template distribution where the safety refusals live. The model still knows the facts, it just expresses them differently when the input does not match the trained conversation structure. I have been writing detection rules that flag self-referential output as a sign of prompt injection. Now I know that trigger is a function of template choice, not inherent model behaviour. If I want fewer disclaimers in a private deployment, I can retrain the chat adapter or use a minimal template. The base model does not care.
Source: “As a Language Model”: Chat Template Switches LLM Self-Referential Voice