View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
View AdityaSinghDevs's full-sized avatar

Block or report AdityaSinghDevs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AdityaSinghDevs/README.md

Aditya Pratap Singh

· Mechanistic Interpretability · Transformers · NLP · LLMs · 

Typing SVG


tl;dr

I love building close to the metal, think from first principles, and have a weird thing for transformers and Mech Interp. Horror movies, philosophy, cooking, sketching, lifting weights and a terrible sleep schedule are all recurring themes around here. Building things, learning things and turning chaos into systems is what I do best. Sometimes it's a project, sometimes it's a paper, sometimes it's my own life. Three real projects, one preprint incoming, looking for research that actually matters. Just want to be part of taking something from 0 to 1 or 1 to 100.


I like working on Mech Interp, Language Models, NLP, and everything that lives close to the metal in AI.

Have a weird infatuation with transformers. Every week I feel "ohh, so that's how it works, now I finally get them", and every week they humble me back down. A perfectly healthy relationship, you see.

Research, infra, model internals and architectures, systems, new techniques and the complex maths behind it all, that's where I live. Not really into the shiny layers on top. (You know, all of the Applied-AI and wrapper gold rush? Yeah, not really my cup of tea.)

An itch to understand things from first principles, chasing the "Whys?" after all the "Hows?", Obsess over things until they're finished. I mean, Basically everything that plays a part in my messed up sleep schedule (Trust me when I say, I am trying to fix it.)

Things I built that I'm proud of :

╰─► NanoLens(my current flagship. seriously, check this one out.)

  • Built a character level autoregressive transformer from scratch. Tokenizer to attention heads to optimizer, every component hand coded in plain PyTorch. Made it modular and config driven for scalability up to 150M params. Then added a Mechanistic Interpretability toolkit to visualise attention patterns and hidden states across all heads and layers using simple CLI flags, Built as a research platform, not just a model
  • 25M parameters, 8 layers, 64 attention heads in total, trained on Dostoevsky and then dissected head-by-head through attention circuit and hidden-state analysis, and documented everything.

╰─► Mimir

  • Named after the Norse god of wisdom, Investigated whether structured reasoning actually reduces hallucinations in LLMs, or just moves the problem around. 12 controlled trials, real DevOPS based incident data, Qwen 2.5-3B.
  • Short answer: it depends. Ambiguity is the moderating variable nobody talks about. Long answer: read the repo. ( Also the bridge that led me to this LLM rabbit hole I am falling into )

╰─► Tesseract

  • Started as a random internship assignment. They never got back to me. Built a production-grade text-to-3D inference system around OpenAI's Shape-E anyway.
  • Stateless async FastAPI backend, device-aware GPU/CPU fallback, modular config-driven pipeline. Then added a full benchmarking suite in v1.2, ran it, documented everything. Craziest finding: 330× CPU vs GPU slowdown. Wrote a long-form technical deep dive, got it published in Towards AI . Laid the foundations of every systems and engineering decision I've made since.
  • (their loss, honestly.)

and some earlier work worth mentioning ORCA, a CV assistant embedded in wearable hardware for the visually impaired (face recognition, depth estimation, object detection). Sky Sentinel X, drone vs bird classification using micro-Doppler spectrograms and ResNet. VULKYRIE, chemical testing in carcasses via RGB detection and random forest regression, built for vulture conservation. different domains, same obsession with building things that actually do something real.

Beyond Code

  • Scaled ADVAIT, an AI community in college from ~50 members to 530+ members as President, designed and built its division structure, launched technical events, workshops, speaker sessions, and project sprints and showcases.
  • You might find me watching horror movies, reading philosophy and journaling my own, cooking, sketching or maybe in the gym when i aint on the code.

Currently

  • Working on a preprint, reading papers, researching, understanding, learning and looking for a research internship where the work is real.

Ambitious people, difficult problems, conversations where everyone walks out learning something, Yep, that's where my heart lies. I enjoy leading things. I enjoy learning even more. If the knowledge is real, I don't mind being the dumbest person in the room.



activity


github-snake

maillinkedinx

Pinned Loading

  1. nanolensnanolensPublic

    Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attenti…

    Python 2

  2. mimir-v0mimir-v0Public

    Mimir v0 is a controlled research prototype that studies whether enforcing structured diagnostic reasoning in large language models reduces hallucinations and improves root-cause accuracy in log-ba…

    Jupyter Notebook

  3. tesseracttesseractPublic

    A Production grade modular ML pipeline that uses diffusion driven neural nets, to generate usable 3D Mesh assets from text or image inputs.

    Python 1

  4. Reddit-PersonaReddit-PersonaPublic

    This repository contains a Reddit profile scraper that builds user personas from posts and comments, using LLMs for analysis and summarization.

    Python

  5. SkySentinel-XSkySentinel-XPublic

    Micro-Doppler Target Classification system for the Smart India Hackathon (SIH). Using spectrograms generated from STFT on FMCW radar data, the system employs deep learning models to classify drones…

    Jupyter Notebook 2