Skip to content
View mooner92's full-sized avatar
:shipit:
Bee's Knees
:shipit:
Bee's Knees

Organizations

@All-solve@LGHackerton@3-2-capstone@HaniumCapstone@smart-marin-logistics@Prometheus-AI-Hackathon@KY-MITE@lg-aimers-yes@4-1-Capstone@mocap4-1@programmers-project-july@Gemini-API-Developer-Competition@DevCourse-project-2@DSC-Hackathon@hack-seoul@place4coder

Block or report mooner92

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mooner92/README.md

Sean Choi · 최명헌

ML Platform & LLM Infrastructure Engineer
I build the infrastructure AI services run on — and measure whether it actually works.

Portfolio · LinkedIn · Email


Reliability is a number, not an adjective. Everything below was measured on a system I actually run.

Measured — not claimed

RAG answers that cite a source, or refuse80.6% → 100%
Strict retrieval accuracy · Hit@160.0% → 82.9%
Silent-failure detection time5.7 days → 40 min
Accelerator task execution time~70% faster than the default K8s scheduler
LLM serving under loadp95 18s single-request · throughput saturates at 4.3 req/min — bottleneck isolated to non-batched generation

Currently

  • Operating a multi-server GPU compute fleet (dual NVIDIA A40) serving LLM inference at the Korea Environment Institute
  • Building KEIwi — an on-prem GPU fleet observability & incident platform
  • Learning my way around Kubeflow, MLflow, Triton Inference Server, and feature stores
  • Open to ML Platform / LLM Infrastructure roles

Stack

Serving & inferencevLLM · Ollama · FastAPI · SSE streaming · GGUF quantization
RetrievalChroma · KURE-v1 embeddings · cross-encoder reranking
OrchestrationKubernetes · Docker · containerd · systemd
ObservabilityPrometheus · Grafana · Loki · DCGM · OpenSearch
Cloud & edgeGCP · AWS · Oracle Cloud · Cloudflare Zero Trust
LanguagesPython · TypeScript · C++ · C · Bash · SQL

Projects

KEIAdminSupervOn-prem RAG over 363 internal regulation documents — every answer cites its source or refusesFastAPIOllamaChroma
KEIwiGPU fleet observability & incident platform — fails safe: unknown state reports no-data, never a false downPrometheusGrafanaDCGM
tpuservAccelerator-aware pod scheduling on bare-metal Kubernetes — published at KIISE 2024KubernetesCoral TPU
dev-boothThree LLM agents coordinating through a Kanban board only — a human keeps final merge authorityvLLMSQLiteFastAPI
MineSweeperConflict-of-interest detection over mixed-format documents — it drafts, it never decidesQwen2.5-VLNext.js

Architecture diagrams and measured results for each → seanchoi.excusa.uk


Publications

Scheduling Techniques for Improving the Efficiency of TPU Work in a Cluster Environment — KIISE, 2024 · first authorResearch on Improving Pothole Object Detection Accuracy Through Similar-Object Data — KIISE, 2024 Proposal of a Korean History Education Application Based on Kubernetes Clustering — 2023 A Study on Chatbot for a Safe Harbor — 2023


Every project above was built alongside Claude — this is the receipt, not just the claim.

Tokscale Stats

Pinned Loading

  1. tpuservtpuservPublic

    클러스터 환경에서의 TPU 작업 효율 향상 스케줄링 기법 연구

    JavaScript 1

  2. dev-boothdev-boothPublic

    dev-booth — Autonomous Multi-Agent Development System

    Python 1

  3. KEIAdminSupervKEIAdminSupervPublic

    Horong — On-Prem RAG Chatbot for Internal Regulations

    Python 1

  4. KEIwiKEIwiPublic

    KEIwi — GPU Fleet Observability & Incident Platform

    Python

  5. MineSweeperMineSweeperPublic

    MineSweeper — Conflict-of-Interest Detection for Recruitment

    TypeScript

  6. RaypokeRaypokePublic

    Advanced Rayban glasses by using poke

    TypeScript