A 434M-active / 1.13B-stored hybrid language model — Gated Delta Networks (linear attention) × Multi-Head Latent Attention (full attention) with Asymmetric Mixture-of-Experts. Pre-trained from scratch on 30B tokens, targeting held-out FineWeb-Edu perplexity ≤ 2.10.
The flagship model of the CoreProjects portfolio.
Hosted docs (GitHub Pages) — the full portal with interactive architecture visualizations, KaTeX math, and syntax-highlighted code.
The documentation is organized by how readers use the repository: concepts explain
why the architecture works, references describe stable APIs and config fields, and
guides show operational workflows. The full corpus lives under docs/. Quick links:
docs/guides/quickstart.md— install, first forward pass, tests and gatesdocs/concepts/model-architecture.md— the code walkthroughdocs/concepts/design.md— the v1.0 architecture & design documentdocs/references/config.md— the typed-config referencedocs/training.md— the training pipeline
Transformer attention scales quadratically with sequence length — the dominant cost of pretraining at scale. HyMo is a hybrid: it processes the bulk of the sequence through linear-complexity recurrence (Gated Delta Net) and reserves sparse full-attention anchors for genuine long-range reasoning. The result is a model that trains and infers far cheaper than an all-attention transformer of equal quality, while keeping the expressivity where it matters.
The headline design choices:
- 3:1 GDN-to-MLA ratio — 24 linear-attention layers interleaved with 8 full-attention layers, so 75% of the stack is sub-quadratic.
- Asymmetric feed-forward — MoE (sparse, expensive) lives only on the 8 full-attention MLA blocks; the 24 linear GDN blocks are recurrence-only (no FFN). Compute is spent where it buys the most.
- Custom Triton GDN kernel — a fused 1D selective scan with chunked recurrence (
chunk_size=64), written by hand insrc/hymo/models/gdn_triton.pyfor throughput and numerical parity with the eager reference. There is nofla-library dependency — the only sanctioned kernel path is this hand-written Triton kernel.
HyMo is a 32-layer stack with a 3:1 GDN-to-MLA ratio.
| Component | Layers | Type | Description |
|---|---|---|---|
| GDN | 24 | Linear attention | Gated Delta Net with 1D selective scan, partial RoPE, recurrence-only (no FFN) |
| MLA | 8 | Full attention | Multi-Head Latent Attention (DeepSeek-style low-rank KV compression, 4 KV heads) |
| MoE | On MLA layers | Sparse FFN | DeepSeekMoE (16 routed + 1 shared expert, top-2 routing, inter_dim = 2304) |
| FFN | On MLA layers only | SwiGLU | Inside the MoE experts, inter_dim = 2304; GDN blocks have no FFN |
| MTP | 2 heads | Multi-token prediction | Auxiliary heads predicting next 2 tokens, weighted [0.3, 0.1] |
Model footprint (v1.0 config):dim = 896, n_heads = 16, max_seq_len = 4096, vocab_size = 64,256 (BPE-64k + 256-byte tokenizer). ~434M active / ~1.13B stored parameters.
Key architectural invariants:
- Asymmetric feed-forward — MoE exclusively on MLA blocks; GDN blocks are recurrence-only (no FFN).
- Partial RoPE — applied to the first 25% of
head_dimat every position across all 32 layers. - MQA-4 — MLA compresses to 4 KV groups for efficient inference.
- FP32 master weights — full numerical stability; optimizer state held in float32.
- NorMuon / AdamW dual optimizer — NorMuon drives attention + GDN 2D matrices; AdamW handles embeddings, norms, gates, and MoE experts. Cautious weight decay enabled.
- Initialization — PyTorch module defaults plus the inline MoE-gate init (
bias=0,std=0.006) and the GDN recurrence init (A_log,dt_bias,D). The designed μP init was never wired intobuild_hymoand was removed in the 2026-08-04 cleanup (seedocs/concepts/optimization.md). - Logit softcap (15.0) — bounds logits for training stability.
- Custom Triton GDN kernel — a hand-written fused 1D selective scan in
src/hymo/models/gdn_triton.py(serial time loop, FP32 accumulation; the parallel-chunk algorithm from the GDN paper remains the design intent). Linux only — Triton does not ship on macOS/Windows; on those platforms the eager path insrc/hymo/models/gdn.pyis the reference and is what unit tests exercise. - FSDP-2 full parameter sharding — BF16 mixed precision, gradient clipping by global norm, NaN-step skipping with configurable tolerance.
- 10-source data pipeline — BPE-64k + 256-byte tokenizer and the held-out FineWeb-Edu validation-set builder remain in-repo; the 10 streaming loaders and shard writer moved to the workspace
LLM/shared_data/package in the 2026-08-04 cleanup (the trainer consumes a rawdata_iter). - Ablation framework — 4 families of config derivation (GDN variants, MLA variants, MoE variants, optimizer variants) via
dataclasses.replaceon the frozen configs; the in-repoablations/package was removed in the 2026-08-04 cleanup — the derivation helperderive_configlives inhymo.core.config. - Cool-by-design test suite — the full 1.13B model is never built in default tests; a ~760K-param surrogate is used instead. Heavy tests (full model construction) are opt-in via
--run-heavy. Defaultpytestfinishes in ~1 minute on an M1 Air. - DCP checkpointing — distributed checkpoint save/load with resume-from-arbitrary-step support.
Install, run the first forward pass, and run the test suite / gates in
docs/guides/quickstart.md. The 30-second
version:
uv sync --all-extrasimporttorchfromhymoimportload_config, build_hymoconfig=load_config("configs/hymo_750m.yaml")
model=build_hymo(config)
x=torch.randint(0, config.model.vocab_size, (2, 128))
# The main model interface returns one vocabulary distribution per input position.logits=model(x)
print(logits.shape) # (2, 128, 64256)pytest tests/ -v # ~1 min on CPU; heavy tests skipped (203 passed / 35 skipped (GPU-gated) as of 2026-08-20)
pytest tests/ --run-heavy # includes full 1.13B model construction
mypy src/hymo # type gate
ruff check src/hymo # lint gateAll hyperparameters live in YAML configs under configs/. The primary config is configs/hymo_750m.yaml, organized into 5 frozen dataclass groups:
| Group | Class | Key knobs |
|---|---|---|
model | ModelConfig | 32 layers, dim 896, 16 heads, 16 MoE experts, MTP depth 2, seq 4096 |
optimizer | OptimizerConfig | NorMuon LR 0.02, AdamW LR 3e-4, FP32 master weights, cautious WD |
scheduler | SchedulerConfig | WSD schedule, ~57.2k total steps, 2% warmup, linear decay |
training | TrainingConfig | Micro-batch 4, grad accum 8, FSDP BF16, eval every 2k steps |
run | RunConfig | Name + output directory |
Every field, validation rule, and the derivation helper are documented in
docs/references/config.md. Derive config
variants via hymo.core.config.derive_config() (e.g. dataclasses.replace on sub-configs).
hymo/
├── configs/ # YAML configurations
│ ├── hymo_750m.yaml # Primary v1.0 config
│ └── hymo_mixture.yaml # Data mixture config
├── src/hymo/
│ ├── core/ # Config dataclasses, types, exceptions, validation (PyTorch-free)
│ ├── models/ # GDN, MLA, MoE, MTP, RoPE, Triton kernel
│ ├── training/ # Trainer, dual optimizer, WSD scheduler, FSDP-2, checkpoint
│ └── data/ # Tokenizer + held-out validation-set builder
├── tests/
│ ├── unit/ # Module-level unit tests
│ ├── integration/ # Cross-module integration tests
│ └── conftest.py # Tiny model fixtures (760K params)
The model and training infrastructure are implemented; the remaining milestone is running the planned production pre-training job. Current phase status:
| Phase | Status | Description |
|---|---|---|
| 1 — Repository foundation | ✅ Done | Clean architecture, public API, config system, CI gates |
| 2 — Algorithmic model | ✅ Done | GDN, MLA, MoE, MTP, RoPE, μP init — all forward/backward finite |
| 3 — Training infrastructure | ✅ Done | Trainer, dual optimizer, WSD scheduler, FSDP-2, DCP checkpointing |
| 4 — Data & eval pipelines | ✅ Done | 10-source loader, tokenizer, sharding, eval harness, ablation framework |
| 5 — Deployment & 30B run | ⏳ Pending | RunPod scripts, 30B-token pre-training on 4× A100 80GB SXM |
The architecture, training, data, evaluation, and ablation pipelines are fully implemented. The 30B-token pre-training run on 4× A100 80GB is the remaining milestone.
- Raw PyTorch first — no HuggingFace
Trainer, no Lightning. The loop, kernels, and distributed training are hand-written and deeply optimized (torch.compile, FSDP-2). - Strong typing — every public function is fully annotated;
mypy --strictis a gate. - No magic numbers — all hyperparameters live in
configs/hymo_750m.yaml; code references them viahymo.core.config. - No circular dependencies —
core ← {models, training, data};modelsandtrainingshare state only through config. - Fully implemented — no
NotImplementedErrorplaceholders for core model logic.
Apache 2.0.