Skip to content
#

model-welfare

Here are 10 public repositories matching this topic...

Language:All
Filter by language

Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.

  • Updated Aug 18, 2026
  • Python

Add this topic to your repo

To associate your repository with the model-welfare topic, visit your repo's landing page and select "manage topics."

Learn more