We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.