llm theoretical performance analysis tools and support params, flops, memory and latency analysis.
-
Updated
Jul 11, 2025 - Python
llm theoretical performance analysis tools and support params, flops, memory and latency analysis.
Hands-on Machine Learning Infrastructure on Kubernetes. Using Microk8s/Ubuntu on Paperspace Cloud.
Evidence-driven PresentMon diagnostics and policy modeling paired with bounded owned-lab D3D11 runtime actuation.
Interactive theoretical Kimi-K3 inference roofline calculator for H200, B300, and GB300
Comprehensive performance analysis of DeepSeek V3 quantization levels (FP16, Q8_0, Q4_0) on 16GB GPU environments.
Reproducible long-context inference benchmark comparing vLLM, SGLang, and TensorRT-LLM on NVIDIA GB10.
Python lab for exploring memory bandwidth, cache effects, and locality in accelerator workloads
Hybrid XGBoost + PyTorch ML system for semiconductor scaling analysis and real-time GPU FPS prediction with SHAP, LIME, and Scaling Integrated Gradients explainability.
Comprehensive machine learning benchmarking framework for AMD MI300X GPUs on Dell PowerEdge XE9680 hardware. Supports both inference (vLLM) and training workloads with containerized test suites, hardware monitoring, and analysis tools for performance, power efficiency, and scalability research across the complete ML pipeline.
Inference optimization experiments on Qwen2.5-3B with vLLM (throughput, latency, prefix cache exploration).
Adjust GPU core clock and memory speeds to improve frame rates and hardware performance in gaming applications.
Correctness-first verification, profiling, and RL rewards for AI-written Triton kernels.
Add a description, image, and links to the gpu-performance topic page so that developers can more easily learn about it.
To associate your repository with the gpu-performance topic, visit your repo's landing page and select "manage topics."