Note that this will soon get outdated. To know more about me and my research, a much better place is my website: 7vik.io.
Independent Researcher | AGI Safety • Interpretability • Reinforcement Learning
📍 UC Berkeley (CHAI), MATS Program
📧 zsatvik@gmail.com | 🌐 7vik.io | Scholar | LinkedIn | GitHub
On a quest to understand intelligence and ensure that advanced AGI is safe and beneficial.
I’m an independent AI safety researcher currently working with:
- CHAI, UC Berkeley — Optimal exploration and long-horizon planning in RL.
- Adrià Garriga-Alonso (FAR AI) — Studying deceptive behavior in frontier AI systems at the MATS Program.
- Nandi Schoots (Oxford) — Hierarchical representations and modular training for interpretability.
Previously:
- Microsoft Research — Worked with Neeraj Kayal on representation learning theory, and Amit Sharma and Amit Deshpande on ICL robustness in LLMs.
- Wadhwani AI — Formulated AI problems in public health and trained robust and interpretable ML large-scale deployments in India.
- Mentored a SPAR 2025 project on zero-knowledge auditing for undesired behaviors.
- AI Alignment & Safety
- Interpretability & Feature Geometry
- Long-horizon RL & Planning
- Representation Learning & Theory
=Equal contribution; full list at Google Scholar
Intricacies of Feature Geometry in Large Language Models
ICLR 2025 (poster); Runner-up, ICLR Blog Awards
Code | BlogAmong Us: A Sandbox for Measuring and Detecting Agentic Deception
Under Review
Poster | BlogAuditing Language Models for Hidden Objectives
Anthropic (external collaboration)
Anthropic Blog | BlogProgress Measures for Grokking on Real-world Tasks
ICML 2024 Workshop on High-dimensional Learning Dynamics
CodeChallenges in Mechanistically Interpreting Model Representations
ICML 2024 Workshop on Mechanistic Interpretability
CodeA is for Absorption: Studying Feature Splitting and Absorption in SAEs
Under ReviewCataractBot: An LLM-Powered Expert-in-the-Loop Chat System
IMWUT / UbiComp 2025
CodePredicting Treatment Adherence of Tuberculosis Patients at Scale
PMLR 2022; Outstanding Paper, NeurIPS 2022
Media Coverage
AmongUs– Agentic deception sandboxnice-icl– ICL optimization toolsgrokking– Measuring grokking dynamicsbyoeb– Healthcare LLM deployment platform
- Email: zsatvik@gmail.com
- Website: 7vik.io
- LinkedIn: @7vik
- Open to collaborations in interpretability, alignment, deception audits, and theoretical ML.


