I recently completed my PhD in Computer Vision at the University of Bologna, where I worked on Multimodal AI for 3D-Language Understanding under Prof. Luigi Di Stefano. During my PhD, I also worked as Research Intern at Niantic Spatial in London, working on spatial reasoning with multimodal LLMs.
My research focuses on teaching machines to understand and reason about the 3D world through language — bridging vision, language, and spatial understanding with large-scale multimodal models.
🔗 Website · Google Scholar · LinkedIn
| Paper | Venue | Code |
|---|---|---|
| Spatially-aware Weights Tokenization for NeRF-Language Models | NeurIPS 2025 | CVLAB-Unibo/Spatial-LLaNA |
| Scaling LLaNA | Under review, IEEE TPAMI | Coming soon |
| LLaNA: Large Language and NeRF Assistant | NeurIPS 2024 | CVLAB-Unibo/LLaNA |
| Looking at words and points with attention | ICCV 2023 Workshop | AndreAmaduzzi/CrossCoherence |
Python · PyTorch · Hugging Face Transformers · DeepSpeed · FSDP · CUDA · C++ · SLURM · Linux
Experienced in large-scale distributed training on GPU clusters (trained 10B+ parameter models on LEONARDO, a top-10 global supercomputer).
I'm currently exploring full-time Research Scientist and ML Engineer roles. Feel free to reach out at amaduzziandrea@gmail.com.
