Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Safety Research
Popular repositories Loading
- persona_vectors
persona_vectors PublicPersona Vectors: Monitoring and Controlling Character Traits in Language Models
- automated-w2s-research
automated-w2s-research Public - SCONE-bench
SCONE-bench Public - assistant-axis
assistant-axis PublicThe Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations—sometimes toward bizarr…
- safety-tooling
safety-tooling PublicInference API for many LLMs and other useful tools for empirical research
Repositories
- auditing-agents Public
Uh oh!
There was an error while loading. Please reload this page.
safety-research/auditing-agents's past year of commit activity - misalignment-indicators Public
Source code for the paper: Probing the Misaligned Thinking Process of Language Models
Uh oh!
There was an error while loading. Please reload this page.
safety-research/misalignment-indicators's past year of commit activity Uh oh!
There was an error while loading. Please reload this page.
safety-research/safety-tooling's past year of commit activity - SCONE-bench Public
Uh oh!
There was an error while loading. Please reload this page.
safety-research/SCONE-bench's past year of commit activity Uh oh!
There was an error while loading. Please reload this page.
safety-research/sleight-bench's past year of commit activity - aligning-ai-teams Public
Uh oh!
There was an error while loading. Please reload this page.
safety-research/aligning-ai-teams's past year of commit activity - faithful-cot Public
Code for "Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning"
Uh oh!
There was an error while loading. Please reload this page.
safety-research/faithful-cot's past year of commit activity Uh oh!
There was an error while loading. Please reload this page.
safety-research/bloom's past year of commit activity - legibility Public
Which models are illegible under what conditions, and why? How does that impact monitorability?
Uh oh!
There was an error while loading. Please reload this page.
safety-research/legibility's past year of commit activity Uh oh!
There was an error while loading. Please reload this page.
safety-research/agent-escape-bench's past year of commit activity
Top languages
Loading…
Uh oh!
There was an error while loading. Please reload this page.
Most used topics
Loading…
Uh oh!
There was an error while loading. Please reload this page.