AI Engineering student in Madrid, turning from shipping production systems toward research.
Over the last year and a half I've built production systems: LLM agents over Salesforce taking real traffic, institutional department tooling in real use, a university booking platform built to spec and pending rollout, and a local-first desktop stats app.
The bugs were the teacher. The bottleneck isn't the engineering, it's the science underneath.
So I'm working my way in from the applied side, with merged fixes to the eval infrastructure behind frontier-model safety and the lessons written down in public.
Builder in the concrete, learner in the deep.
Small fixes, real repos — the eval infrastructure behind frontier-model safety research.
- METR/hawk #627 — METR (Model Evaluation & Threat Research). Auth fix in
hawk, their cloud Inspect-AI eval runner: a passthrough credential header containing only whitespace is now treated as missing (401+anonymous) instead of surfacing an unactionable"invalid api key". Merged. - princeton-pli/hal-harness #182 — HAL, the Holistic Agent Leaderboard (Princeton PLI). Resolved a
wandb/weave/gqlversion conflict that broke every agent-env container at startup; aligned the three runners, the Docker image, packaging, and contributor docs. Merged. - affaan-m/ECC #920 — everything-claude-code, the most-starred agent toolkit on GitHub. Contributed the Token Budget Advisor skill. Merged.
Real systems built for clients and institutions, some live and taking traffic, some built to spec and pending rollout. Code is private; writeups on the architecture and the lessons are landing on xabier.me.
| reeljet AI short-form video ad generator, as a Claude Code skill | TBA — Token Budget Advisor Depth/cost choice before Claude answers · merged into ECC |
| xabier-blog Personal research notebook · Astro · xabier.me | save-session Claude Code skill — compressed session summaries to a vault |
Also early / idea-stage: harness-ablation (an agent-harness ablation study I sketched out) · xabier-arena-solutions (placeholder for the ARENA curriculum, starting soon).
- Writing up the production lessons as case studies on xabier.me: SAM's agent architecture and the LinceReservations build first.
- Contributing to eval-framework open source — more PRs in the pipeline after the three above.
- Next — the ARENA alignment curriculum, to build the research foundations under the applied experience.
| Highlight | Detail |
|---|---|
| Merged eval-infra fixes | 3 PRs into METR/hawk, Princeton HAL, and ECC — the eval stack behind frontier-model safety research |
| Production work | LLM agents over Salesforce taking real traffic, department tooling for ~30–40 professors in real use, a university booking platform built to spec (pending adoption), a local-first desktop stats app |
| Open source | reeljet, TBA (merged into ECC), save-session, and this blog |
| Writing in public | xabier.me — a research notebook, bilingual (ES/EN) |
| Studying | 2nd-year AI Systems Engineering at UFV, Madrid |
Focus — agent evals · Inspect-AI · production LLM systems · Stack — Python · TypeScript · FastAPI · React · Astro




