Psychology x AI governance. I bring the human-factors angle most governance people lack: I come from psychology (how people actually think, decide, and miscalibrate), and I work in AI risk, responsible AI, and AI evaluation.
Honours Psychology (Research Intensive) co-op student at the University of Waterloo. Previously AI Risk Governance co-op at Rogers Communications, where I benchmarked corporate AI policy against NIST AI RMF, ISO/IEC 42001, OSFI E-23, the EU AI Act, and Canada's AIA, and built incident frameworks and a Responsible AI evaluation plan with PASS/FLAG release gates.
I design the substance: the research question, the study design, the scoring logic, the KPIs, and the report structure. The implementation — Python, Streamlit, analysis code, front-end — is AI-assisted, and every repo says exactly which part is which. For Responsible AI work, being precise about that line is the point.
Live work (what each one shows)
- Do Large Language Models Know When They Are Wrong? : original research. 89 expert-level questions × 3 prompts × 2 models — 534 responses, run with a research partner. Claude Opus 4.8 and GPT-5.6 scored about the same, then missed their own accuracy in opposite directions: one underconfident, one overconfident. Telling them to be humble shifted stated confidence significantly (p = .004) and made calibration worse for both. The report runs the study's instrument on the reader before it tells them anything, and puts the ceiling effect that limits every claim on the page rather than in a footnote.
- Is It Smart, or Does It Just Sound Like Us? : original research. An interactive essay built on my own Turing-test study, where an AI told to "sound human" was judged more human than me (150 vs 101). Evaluation thinking, made tangible.
- AI Governance Audit Framework : governance design. 30 questions across 6 domains, Canada AIA impact estimation, OSFI E-23 domain, every finding mapped to its regulatory clause.
- Universal Analytics Engine : data plus governance. A dataset-agnostic BI platform with an embedded model-risk layer: PII detection, AIA impact scoring, proxy-bias flags, integrity scorecard.
- ESG Corporate Risk Scorecard : scoring logic. SASB-weighted, sector-adjusted ESG scoring with TCFD and GRI alignment.
Frameworks I work in: NIST AI RMF 1.0 · ISO/IEC 42001 · OSFI E-23 / B-13 · Canada AIA · EU AI Act · PIPEDA
What I want to work on next: the same question pointed the other way. We are getting better at measuring whether a model's stated confidence matches its accuracy, and we almost never measure that in the people reading the output — even though a confident person is making the same request a confident model is: stop checking. That gap is the human-factors core of AI evaluation, and psychology has decades of evidence on it that the governance conversation rarely uses.
Open to: Summer 2027 co-ops in responsible AI, AI evaluation, trust and safety, AI policy, technology and model risk, data and product analytics, and AI product.