A simplified LLAMA implementation for training and inference tasks.
-
Updated
Jul 9, 2025 - Python
A simplified LLAMA implementation for training and inference tasks.
A powerful Retrieval Augmented Generation (RAG) application built with NVIDIA AI endpoints and Streamlit. This solution enables intelligent document analysis and question-answering using state-of-the-art language models, featuring multi-PDF processing, FAISS vector store integration, and advanced prompt engineering.
An AI-powered interview practice agent that simulates real interviews, asks contextual follow-up questions, adapts to user behavior, and provides detailed feedback using Groq Llama 3.1. Built with FastAPI and JavaScript.
LLM quantization project built around `llama.cpp` + `Ollama` + `GGUF`
To associate your repository with the llama-models topic, visit your repo's landing page and select "manage topics."