Personal search for technical screenshots and notes. Hybrid vector and full-text retrieval on Amazon EKS, with Firn and S3 as the storage and index layer, and GPU capacity that scales to zero.
-
Updated
Jul 29, 2026 - Python
Personal search for technical screenshots and notes. Hybrid vector and full-text retrieval on Amazon EKS, with Firn and S3 as the storage and index layer, and GPU capacity that scales to zero.
Scale-to-zero with wake-on-request for Kubernetes. Sleep idle services on a schedule, wake them instantly on HTTP access. Single pod, no CRDs, no sidecars.
Serverless-GPU LLM serving: scale-to-zero with fast GPU snapshot/restore (cuda-checkpoint), multi-tenant packing, and an OpenAI-compatible API — built on vLLM.
Queue-driven, scale-from-zero GPU inference for any Kubernetes — bursts to cross-region VMs when GPUs run dry
Turn any Kubernetes cluster into a private serverless platform — self-hosted Cloud Run / Functions / Tasks. An Altikva product.
Scale-to-zero distributed AI news oracle. FastAPI orchestrator auto-provisions ARM EC2 instances on demand, runs a quantized Llama 3.2 3B model in 2 GB RAM, and shuts down to keep cloud spend near $0. Audio briefings via AWS Polly.
Add a description, image, and links to the scale-to-zero topic page so that developers can more easily learn about it.
To associate your repository with the scale-to-zero topic, visit your repo's landing page and select "manage topics."