View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
View atandra2000's full-sized avatar
💭
Learning has no ending
💭
Learning has no ending

Block or report atandra2000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
atandra2000/README.md
Atandra Bharati — Deep Learning Research Engineer

Building frontier AI architectures from scratch in raw PyTorch.
LLMs · Latent Diffusion · Multimodal · Video Understanding · Agentic ML · State-Space Models · Long-Context Attention


🎯 Open To:Deep Learning Research EngineerLLM EngineerGenAI / Diffusion EngineerAgentic ML Engineer
🌍 Remote-friendly · Available worldwide


18 Projects78% Memory Cut878 TestsBest Loss 0.09471.13B Params2× KV Cut at 128K


📌 Current Focus

Building and shipping production-grade from-scratch AI — from LLM pre-training infrastructure to autonomous multi-agent research orchestration. Every architecture is implemented layer-by-layer in raw PyTorch, verified with rigorous tests, and tracked on W&B / Comet.


🏆 Highlights

🧠 HyMo
Flagship Hybrid LLM
434M active · 1.13B stored
3:1 GDN/MLA · Asymmetric MoE · MTP
Custom Triton GDN · FSDP-2
📉 LLaMA-3-Lite
Memory-Engineered LLM
78% peak memory reduction
92 GB → 20 GB on A100
batch 96 · 2× headroom
🎨 Stable Diffusion
From-Scratch UNet
860M UNet · 7 training phases
0.0947 loss @ epoch 16
2× RTX 5090 · 1.3M+ images
🔭 GPT-OSS-Lite
Long-Context MoE
KV-cache cut at 128K
Sliding/Full alt + learned sinks
502M params · 130 tests
Mamba-3-Lite
Complex-Valued SSD
50% smaller state (N=64)
Parity loss vs Mamba-2 N=128
Pure PyTorch · no custom CUDA
🤖 AutoML Researcher
Multi-Agent Platform
15 phases · 23 agents · 878 tests
61 tools · 186 models
Paper → experiment → report


📂 Projects

🧬 LLM (7)

ProjectScaleKey InnovationHardwareRepo
HyMo434M act / 1.13B stored3:1 GDN/MLA hybrid · Asymmetric MoE · MTP · custom Triton GDN kernel · FSDP-24× A100 80GB
GPT-OSS-Lite502M / 247M activeSliding(128)/Full attn alt · learned sink bias · YaRN 128K · top-2-of-8 MoE · 2× KV-cache cut · 130 testsA100 80GB
Mamba-3-Lite~434MComplex64 SSD (N=64) · MIMO head mixing · zero causal conv · pure PyTorch · parity loss at half state sizeA100 80GB
DeepSeek-v3-Lite~422MMLA + AuxLossFree MoE + MTP · absorption-trick inference · 643-line MLA technical deep-diveA100 80GB
LLaMA-3-Lite~515MGQA · RoPE · SwiGLU · chunked CE · 78% memory cut (92 GB → 20 GB) · batch 96 on single A100A100 80GB
TranslationLM~44MEncoder-decoder Transformer · EN→IT · loss 6.17→2.28 · BLEU/CER/WER · attention vizP100
GPT-2~6MFrom-scratch GPT-style char-level decoder · learned positional embeddings · causal attention · top-k sampling · frozen educational foundationP100

👁️ Vision (8)

ProjectScaleKey InnovationHardwareRepo
Stable Diffusion 1.x860M UNetCustom UNet from random init · 7-phase curriculum · 1.3M+ images · best loss 0.0947 · epoch-42 checkpoint2× RTX 5090
Detect-Objects~50MRT-DETR/DINO deformable detector · no anchors/NMS · COCO 2017 · Gradio + ONNX2× RTX 5090
Upscale-SR~75M4× real-world SR · latent-diffusion UNet + SSM refiner · Real-ESRGAN degradation2× RTX 5090
ActionRecognition120 clsHRNet pose + Two-Stream CTR-GCN · ~30 FPS inference · ONNX + TensorRT + FastAPIRTX 3090
FaceAgingCycleGAN256²AdaIN per-layer conditioning · 3-scale PatchGAN · LSGAN + R1 GP · 31/50 epochsRTX 6000 Ada
FaceGenerationVAEβ-VAE50 epochs · recon MSE 0.0152 · linear KL annealing · bilinear-upsample decoder (no checkerboard)P100
DCGAN-Face-Generation6.4MN(0,0.02) init · 202K CelebA · D loss → ln 2 ≈ 0.693 GAN equilibrium2× T4
VisionLanguageModelPaliGemma-styleSigLIP ViT + Gemma decoder · linear projector · zero pretrained weights · COCO 2014P100

🤖 Agentic (3)

ProjectScaleKey InnovationHardwareRepo
LearnAgenticAI10 projects · monorepoLangChain / LangGraph / LangSmith / MCP reference portfolio · ReAct, RAG, Memory, Multi-Agent Supervisor, HITL, MCP Server, Deep Research, Structured Output, Eval Harness, Next.js 15 chat UILocal + Docker
Autonomous ML Research Engineer15 phases · 23 agents · 61 toolsFull paper → conclusions loop · self-repair · provider-agnostic LLM routing · 878 tests · 186-model registryLocal + Ollama
newsagent291 testsAutonomous AI research-intel agent · daily 15+ source sweep · LLM-reasoned reports with provenanceLocal

✍️ Writing

DocumentLengthCovers
Multi-Head Latent Attention — Technical Deep-Dive643 linesKV-cache math · low-rank compression algebra · absorption-trick derivation · decoupled RoPE · SDPA vs manual attention
Attention Sinks — StreamingLLM for GPT-OSS600 linesPer-head learned sink bias · BF16 stability (clamp [-10,15]) · sliding/full alt interaction
State-Space Duality — The Mamba-3 SSD AlgorithmFull derivationChunkwise SSD · complex64 packing · naive O(T) recurrence equivalence

🛠️ Tech Stack

Languages & ML Core
Python 3.12PyTorch 2.xCUDA 12.xTriton 2.x

Architectures
TransformersGQAMLAMoEGDNMTPSSD (real & complex64)MIMODiffusion UNetVAEGANCycleGANAdaINST-GCNHRNetSigLIPDeformable DETR

Optimization & Numerics
BF16FP16FP8Flash Attention 2SDPAtorch.compilechannels_lastGradient checkpointingμP scalingWSD LRNorMuonCautiousAdamWChunked CEFused optimizersChinchilla-optimal scaling

Hardware Validated
A100 80GBRTX 5090RTX 6000 AdaRTX 3090P1002× T4

Tooling & Frameworks
LangChainLangGraphLangSmithNext.jsHuggingFacediffusersW&BCometsafetensorsONNXFastAPIGradiopydantic

---

🔬 Engineering Philosophy

  • From-scratch PyTorch — no Trainer, no Lightning, no accelerate; every layer is written by hand
  • Single-GPU feasibility — every large project fits one consumer GPU via BF16, gradient checkpointing, FA2, channels_last, fused optimizers
  • Faithful reproductions — DeepSeek-V3, LLaMA-3, GPT-OSS, Mamba-3, PaliGemma, DCGAN — implemented to the paper
  • Novel hybrids — HyMo (GDN + MLA + MoE + MTP), FaceAgingCycleGAN (AdaIN-conditioned), GPT-OSS-Lite (sink bias + sliding/full alt)
  • Production hygiene — atomic checkpoints (.tmp.ptos.rename), full RNG-state reproducibility, W&B / Comet tracking, CI lint + tests
  • Hardware breadth — MPS/CPU → Kaggle T4/P100 → A100 80GB → 2× RTX 5090 → RTX 6000 Ada

🎓 Background

B.Tech, 2024 · Heritage Institute of Technology, Kolkata. Self-taught in deep learning through two years of from-scratch implementation — engineering discipline from infrastructure and constraint work translates directly to memory budgets, distributed training, and reproducible ML systems.


📫 Connect

PortfolioLinkedInGitHubW&BKaggleCometEmail


18 from-scratch projects (7 LLM · 8 Vision · 3 Agentic) · Updated 2026-08-29 · Open to remote and on-site DL/LLM/GenAI roles worldwide

GitHub starsGitHub followers

Pinned Loading

  1. StableDiffusionStableDiffusionPublic

    A Stable Diffusion 1.x-class latent diffusion model trained from scratch on 2× RTX 5090 (Blackwell) GPUs. Full UNet (~860M params), DDPM/DDIM, LAION pipeline, DDP+BF16.

    Python

  2. DeepSeek-v3-LiteDeepSeek-v3-LitePublic

    Faithful from-scratch reimplementation of DeepSeek-V3 (MLA + aux-loss-free MoE + MTP + speculative decoding), ~412M params, Chinchilla-optimal 8.4B-token training on a single A100 80GB.

    Python 3