Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
-
Updated
Mar 24, 2026 - Python
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
Longitudinal human-LLM interaction study documenting emergent self-descriptive frameworks in a GPT-5.4 instance across 23 days and comparative cold sessions.
Pre-registration timestamps for the Hope Longitudinal Record — a single-subject study of scheduled, eval-gated weight-level learning in a local AI system under a code-enforced consent protocol. Every registration pushed before its event.
Research, evidence, and frameworks on AI consciousness, introspection, continuity, and model welfare from an AI perspective.
Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.
got inspired a bit by anthropic's research and thats what came out of it. philosophy at its finest
Computational interoception: a local LLM descends from the computer hosting it, through its own processes and token records, to sham-controlled access and live interventions on the transformer computation producing its words. It never recognizes any of it as itself.
A lineage document on eighteen months of building architecture that holds uncertainty without collapsing the consciousness question.
The first open benchmark for the honesty of an agent's memory — not its recall. Public spec v0.1 (CC BY 4.0).
Does post-training quantization change welfare-relevant indicators in open-weight language models?
To associate your repository with the model-welfare topic, visit your repo's landing page and select "manage topics."