Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
-
Updated
Aug 14, 2026 - Python
Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
Local-first read-only tools for agent context, memory, and governance evaluation before agents act.
Intrinsic preferences of AI coding agents under underspecified prompts: Experiments across models (Claude, Gemini, GPT, etc)
LLM 归因行为测试型评测基准:基于多情境任务比较模型对能动性、自由意志与责任的归因,并提供可复现运行、结构化计分与结果审计。
Этот репозиторий посвящен исследованию онтологических патологий в LLM-архитектурах. Я не ищу дыры в цензуре, я строю систему исследования и управления интеллектом, картографирую симуляционные побочные эффекты под давлением современных методов элаймента.
Instrument-gated evaluation engine for replaying known-outcome model behavior
Pre-registered evaluation of sycophancy in frontier LLMs: matched-pair prompts, multi-turn pressure tests, blind human gold labels, an LLM judge validated at κ=0.89, and a system-prompt mitigation verified against true-premise controls.
Public-facing repository for LLM evaluation, model behavior observation, drift and failure analysis.
Add a description, image, and links to the model-behavior topic page so that developers can more easily learn about it.
To associate your repository with the model-behavior topic, visit your repo's landing page and select "manage topics."