Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
-
Updated
Mar 24, 2026 - Python
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
Longitudinal human-LLM interaction study documenting emergent self-descriptive frameworks in a GPT-5.4 instance across 23 days and comparative cold sessions.
got inspired a bit by anthropic's research and thats what came out of it. philosophy at its finest
Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.
Does post-training quantization change welfare-relevant indicators in open-weight language models?
To associate your repository with the model-welfare topic, visit your repo's landing page and select "manage topics."