← ablaut

Accuracy

Every language is scored at test time against the slots where two independent oracles agree — a Wiktionary extraction crossed with a national or academic lexicon. The engine hits 100% there; the residual is the two oracles disagreeing with each other, which is ruled on case by case.

What “accurate” means

A form is correct when it is the form a competent native speaker writes for that cell of the paradigm. We can’t poll speakers form by form, so we approximate truth with agreement between two independent authorities: for each language we take a Wiktionary-derived lexicon and a second, independently built source — a national or academic lexicon, a UniMorph dataset, or an Apertium dictionary — and score the engine only on the forms where those two already agree.

Accuracy is the share of that agreed set the engine reproduces exactly:

accuracy = engine-matched forms ÷ forms the two oracles agree on

Two things are deliberately kept out of that number. Where the oracles disagree with each other there is no ground truth to score against, so those slots are excluded and instead adjudicated by hand against reference grammars and published (the Disagreements column). And a slot the engine doesn’t yet produce is a coverage gap, tracked separately (the Slots column) rather than counted as a wrong answer. A form the engine invents that neither oracle attests is always a failure — that is the error class this whole method exists to catch.

Tiers

Everything shipped is at least Beta: 100% correct on one oracle.

Across versions

Correctness is measured by the golden harness, which landed in 0.8.0; every release since adds a point here, so a regression would show up as a number going down.

ReleaseLanguagesVerified formsOverall agreement
0.9.0283,041,19699.998%
0.8.0222,914,02899.998%

By language

LanguageOraclesAgreed formsAgreedSlotsDisagreementsTier
Armenianuniparser ∩ unimorph50,521 / 50,521100.00%63/63179/179Verified
Catalanfreeling ∩ kaikki187,186 / 187,186100.00%50/5013/110Verified·gaps
Czechmorfflex ∩ kaikki107,812 / 107,812100.00%29/320/1397Verified·gaps
Danishcor-verbs ∩ kaikki20,205 / 20,205100.00%9/90/165Verified·gaps
Dutchapertium ∩ kaikki27,836 / 27,836100.00%12/120/66Verified·gaps
Englishagid ∩ kaikki74,600 / 74,600100.00%5/50/163Verified·gaps
Estonianvabamorf ∩ kaikki24,184 / 24,184100.00%41/410/77Verified·gaps
Finnishomorfi ∩ kaikki406,209 / 406,209100.00%35/350/8201Verified·gaps
Frenchlefff ∩ kaikki284,034 / 284,034100.00%49/490/342Verified·gaps
Germanunimorph ∩ kaikki193,038 / 193,06799.98%30/300/2141Verified·gaps
Hindikaikki ∩ unimorph48,421 / 48,421100.00%205/2053708/3708Verified
Icelandicbin ∩ kaikki5,749 / 5,749100.00%26/260/42Verified·gaps
Irishbunamo ∩ kaikki27,404 / 27,404100.00%25/250/488Verified·gaps
Italianmorphit ∩ kaikki280,143 / 280,143100.00%49/490/1185Verified·gaps
Japaneseipadic ∩ kaikki9,421 / 9,421100.00%3/30/22Verified·gaps
Koreankrdict ∩ kaikki3,576 / 3,576100.00%4/40/17Verified·gaps
Portuguesemorphobr ∩ kaikki373,163 / 373,163100.00%70/700/191Verified·gaps
Romaniandexonline ∩ kaikki225,494 / 225,494100.00%35/350/1977Verified·gaps
Russianopencorpora ∩ kaikki151,912 / 151,94299.98%19/190/273Verified·gaps
Sloveniansloleks ∩ kaikki10,757 / 10,757100.00%25/250/226Verified·gaps
Spanishfreeling ∩ kaikki339,200 / 339,200100.00%59/590/723Verified·gaps
Swahiliswc ∩ kaikki1,134 / 1,134100.00%26/2636/36Verified
Swedishsaldo ∩ kaikki39,594 / 39,594100.00%11/130/296Verified·gaps
Tagalogkaikki ∩ unimorph469 / 469100.00%7/70/191Verified·gaps
Tamilthamizhi ∩ kaikki31,829 / 31,829100.00%37/370/980Verified·gaps
Teluguunimorph1,159 / 1,159100.00%24/24Beta
Turkishkaikki ∩ unimorph44,156 / 44,156100.00%93/930/10763Verified·gaps
Ukrainianvesum ∩ kaikki71,990 / 71,990100.00%14/140/329Verified·gaps

Agreed forms — engine matches / slots the two oracles agree on. Slots — paradigm slot types the engine covers / present in the gold. Disagreements — oracle-vs-oracle splits resolved / total.

Measured for ablaut 0.9.0 by the golden harness (regenerated each release). Per-language detail lives in the crate under docs/<lang>/.