Senior Software Engineer @fullbeaker · Full-stack web · Self-hosting LLMs on the side
- Pacific North West
Highlights
- Pro
Pinned Loading
- vllm-qwen36
vllm-qwen36 PublicServe Qwen3.6 NVFP4 on Blackwell GPUs with vLLM - full 262K context, fp8 KV cache, and MTP speculative decoding via Docker Compose
Shell
- vllm-gemma4
vllm-gemma4 PublicServe Gemma 4 NVFP4 on Blackwell GPUs with vLLM - OpenAI-compatible API with full 262K context, vision, thinking, and tool calling via Docker Compose
Shell
- nvfp4-vllm
nvfp4-vllm PublicQuantize HuggingFace models to NVFP4 and serve them with vLLM on NVIDIA Blackwell GPUs.
Python 2
- llama-qwen36
llama-qwen36 PublicDockerized llama.cpp Vulkan server setup for running Qwen3.6 27B GGUF on AMD GPUs.
- slugvision
slugvision PublicTiny VLMs that turn an image + optional article title into a 3-5 word permalink slug — distillation pipeline, training, eval, and GGUF export
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.
Uh oh!
There was an error while loading. Please reload this page.






