
LLM Refusal Behavior on Open-Weight Model
Written by: Trimikha Valentius (CAIO, Head of ZENTARA Labs), Alisha Deana (AI Researcher) Conversational language models refuse harmful requests because a thin behavioural layer is added on top of a pretrained base during post-training. This article treats refusal as exactly that: a layer, not a pro
Trimikha Valentius




