Skip to content

Repository files navigation


ablaut logo

ablaut

Fast, correct verb conjugation for 20+ languages.


crates.io PyPI npm CI license

Try it in the browser


The name is the German term for the vowel gradation in singen, sang, gesungen: Jacob Grimm formalized its seven classes in 1819, and they still describe every strong verb in the language.

Highlights

  • Measured correctness. Every language is validated against two independent machine-readable lexicons. Where the two sources agree, ablaut has zero known errors in nearly every language; the remaining disagreements are ruled on in published adjudication logs. CI re-checks 3.5 million forms on every change.
  • Generalizes to unseen verbs. Rule engines with curated exception tables, not lookup dumps: novel verbs (googeln, tweeter) conjugate correctly.
  • Fast and small. No I/O, no runtime data files, no dependencies in the core crate. A full table takes microseconds; every language fits in one small WebAssembly binary.
  • Permissively licensed. MIT OR Apache-2.0. Reference lexicons are used at test time only and never shipped.

Languages

German, French, Spanish, Catalan, Portuguese, Italian, Romanian, Swedish, English, Danish, Czech, Slovenian, Estonian, Finnish, Irish, Ukrainian, Icelandic, Japanese, Korean, Dutch, Russian, Eastern Armenian, Turkish, Hindi, Swahili, Tamil, Telugu, Tagalog, Persian, Kannada, Gujarati, Urdu, Bengali, and Marathi.

Each language is scored against the slots where its two reference lexicons agree; the second column is that score.

Language Verified forms Accuracy
🇩🇪 German 194,254 99.2%
🇫🇷 French 284,034 100.00%
🇪🇸 Spanish 339,200 100.00%
🇦🇩 Catalan 187,186 100.00%
🇵🇹 Portuguese 373,163 100.00%
🇮🇹 Italian 280,143 100.00%
🇷🇴 Romanian 225,494 100.00%
🇸🇪 Swedish 39,453 100.00%
🇬🇧 English 74,600 100.00%
🇩🇰 Danish 20,205 100.00%
🇨🇿 Czech 107,812 100.00%
🇸🇮 Slovenian 10,757 100.00%
🇪🇪 Estonian 24,184 100.00%
🇫🇮 Finnish 406,209 100.00%
🇮🇪 Irish 27,404 100.00%
🇺🇦 Ukrainian 71,990 100.00%
🇮🇸 Icelandic 5,749 100.00%
🇯🇵 Japanese 9,421 100.00%
🇰🇷 Korean 3,576 100.00%
🇳🇱 Dutch 27,836 100.00%
🇷🇺 Russian 151,942 99.98%
🇦🇲 Armenian 50,521 100.00%
🇹🇷 Turkish 44,156 100.00%
🇮🇳 Hindi 48,421 100.00%
🇹🇿 Swahili 1,134 100.00%
🇮🇳 Tamil 31,829 100.00%
🇮🇷 Persian 2,683 100.00%
🇮🇳 Kannada 411 100.00%
🇮🇳 Telugu † 1,159 100.00%

† Telugu is scored against UniMorph alone. Its intended second oracle, kaikki.org, fills conjugation tables for only ~11 verbs (one shared with UniMorph), too few for a two-oracle agreement gate; see docs/tel/oracles.md. | 🇵🇭 | Tagalog | 469 | 100.00% |

Tagalog is a work in progress: the aspect × voice paradigm is built with a new infixation and reduplication mechanism (see below), and it scores 100% on the slots where its two lexicons agree — but that agreement set is small (138 roots, 469 slots), because the oracles frequently differ on which voice a root lexicalizes. See docs/tgl/oracles.md for the honest scope. | 🇮🇳 | Gujarati † | 2,880 | 100.00% |

Gujarati is scored against UniMorph alone: the available kaikki.org Gujarati tables share the same English-Wiktionary (gu-conj) lineage, so they are a spot check rather than an independent agreement partner. The engine covers the full paradigm at 100% on the 2,880 UniMorph forms across 90 lemmas; see docs/guj/oracles.md. | 🇮🇳 | Marathi † | 56,551 | 100.00% |

† Marathi is scored per cell against apertium-mar, a hand-built, fully person/gender/number-tagged FST (there is no UniMorph Marathi). kaikki.org is the independent second oracle, but Wiktextract lost the person/number on its finite cells, so it forms a per-cell agreement loop over the non-finite forms only (493/493) and corroborates the finite paradigm at the set level (98.9%). The engine covers the full paradigm at 100% on 56,551 forms across 1,304 lemmas; see docs/mar/oracles.md. | 🇵🇰 | Urdu | 777 | 100.00% |

Urdu reuses the Hindi engine (src/hin.rs) and the shared Perso-Arabic normalizer, and is scored against two independent oracles — kaikki.org and apertium-urd. kaikki's ur-conj template emits no person tags, so the per-cell agreement gate covers the person-independent core (infinitive, oblique, participles) at 777/777; the finite paradigm is corroborated by apertium and UniMorph separately. See docs/urd/oracles.md. | 🇧🇩 | Bengali † | 3,864 | 100.00% |

† Bengali is scored against UniMorph alone: kaikki.org Bengali shares the same Wiktionary lineage as UniMorph ben (whose source is Wikipedia), so it is a spot check rather than an independent agreement partner. The engine covers the full paradigm at 100% on the 3,864 UniMorph forms; see docs/ben/oracles.md.

Details per language, including which lexicons are used and every adjudicated disagreement, live in docs/{lang}/.

Usage

Rust

cargo add ablaut
use ablaut::{conjugate, Conjugation, Lang};

match conjugate("vorbi", Lang::Ron)? {
    Conjugation::Ron(t) => assert_eq!(t.present[0], "vorbesc"),
    _ => unreachable!(),
}

Per-language modules expose richer APIs; the shared contract is Verb::from_infinitive and Table::build:

use ablaut::{Mood, Number, Person, Tense, Verb};

let v = Verb::from_infinitive("aufstehen")?;
v.conjugate(Tense::Present, Mood::Indicative, Person::First, Number::Singular);
// "stehe auf"

use ablaut::fra;
let v = fra::Verb::from_infinitive("appeler")?;
v.conjugate(fra::SimpleTense::Future, fra::Person::First, fra::Number::Singular);
// "appellerai"

Reverse lookup maps a conjugated form back to its infinitive(s) and the slots it fills — fully productive for German, English, French and Spanish, irregular-index-backed for every language:

use ablaut::{reverse, Lang};

let m = reverse("suis", Lang::Fra);
// être (present 1sg) and suivre (present 1sg, present 2sg)
assert_eq!(m.len(), 2);
assert_eq!(reverse("war", Lang::Deu)[0].infinitive, "sein");
assert_eq!(reverse("hablé", Lang::Spa)[0].infinitive, "hablar");

Python

pip install ablaut
import ablaut

c = ablaut.conjugate("aufstehen")   # German is the default
c.present[0]                        # "stehe auf"
c.auxiliary                         # "sein"

ablaut.conjugate("appeler", lang="fr").present[0]  # "appelle"
ablaut.conjugate("delati", lang="sl").present[3]   # "delava" (the dual)

Wheels are abi3 (Python 3.11+); no Rust toolchain needed.

WebAssembly

wasm-pack build --target web -- --features wasm
import init, { conjugate } from "./pkg/ablaut.js";
await init();
conjugate("aufstehen").present[0];   // "stehe auf"
conjugate("andare", "it").auxiliary; // "essere"

Language codes are ISO 639-1/-3 or English names, case-insensitive.

Coverage

Each language covers its full synthetic paradigm plus the analytic tenses of its written standard: German's separable prefixes and both Konjunktiv rows, French spelling doublets and reflexive clitics, Spanish enclitic imperatives with written stress, Portuguese's personal infinitive, Italian's perfect auxiliary, Romanian's synthetic pluperfect, the Scandinavian s-passives, Czech gendered participles, Slovenian's dual, Estonian and Finnish impersonals and potentials, Irish's initial mutations, Armenian's converbs with the fronted negative copula (գրում եմ / չեմ գրում), and Tamil's tense–PNG suffix stacking across its gendered third person (செய்தான், செய்கிறது, செய்வார்கள்). docs/design.md documents the German engine in depth; each other language documents its scope in docs/{lang}/oracles.md.

How correctness is verified

For each language, two independently derived lexicons (a Wiktionary extraction and a national or academic resource) are converted to a common format. The engine is scored on every slot where the two agree, and CI fails if accuracy regresses. Disagreements between the sources are ruled on case by case in docs/{lang}/adjudications.tsv, with references to standard grammars. The process regularly finds errors in the reference data itself; several fixes have been filed upstream.

Design

Each engine is a productive default rule plus a finite table of stored exceptions, the dual-mechanism model of inflection (Marcus, Pinker et al., 1995). Features follow the UniMorph schema.

License

MIT OR Apache-2.0.

Data provenance: the exception tables are factual data about each language (principal parts, class membership, auxiliaries), established by the agreement of independent sources and verified against reference grammars. The reference lexicons themselves are fetched at development time and never shipped, whatever their license.

About

Fast, correct verb conjugation for 50+ languages

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages