I would like to propose an updated Swedish dictionary based on a conservative conversion of yeager/hunspell-sv, combined with the existing aspell-sv-0.51-0 word forms.
This is a request for technical review of a working prototype, not a request to publish it as a release yet.
Reproducible prototype and validation
Source scripts, pinned inputs, checksums, build instructions and local test results.
The converter expands supported suffix rules, validates candidates using libhunspell, combines them with the old Aspell word list, and builds an Aspell dictionary. Reviewed project test vocabulary is also exported explicitly after validation. The Hunspell input is pinned to 167ef93ff69812279701979e01bdc70675a859a1; related source/build corrections are in hunspell-sv PR #1.
Local results with Aspell 0.60.8.2 and libhunspell 1.7 on Linux:
- 943,330 compiled word forms; 823,749 are absent as explicit entries in the old Aspell list. These are surface-form counts, not lemma counts or a quality measure.
- The entire compiled dictionary matches the exported word set exactly, without build warnings.
- Two clean builds produced byte-identical word lists, RWS files and language data.
- Both Hunspell and Aspell accept 38 positive examples and reject eight negative examples in a targeted regression test.
- The Hunspell source/build corrections pass nine local regression tests.
Open design questions
- Swedish compounds: Hunspell's compound grammar is not ported. With Aspell's original
run-together true, the merged dictionary accepted errors such as översätning and säkerhett, as well as bilbil, datordator and falllucka. The prototype disables run-together, which avoids these cases but can reject correct unseen compounds. What approach would you recommend for an upstream Swedish update?
- Character data: the old ISO-8859-1 Swedish data is retained. 4,379 candidate forms are excluded for character/format reasons and recorded in
excluded.txt during the build. Would extending the Swedish character data be preferable?
- Special flags: compound-only, NOSUGGEST and FORCEUCASE entries are conservatively omitted, and NEEDAFFIX stems are not exported alone. Explicit forbidden entries are removed from the union. This does not provide full Hunspell semantic parity.
- Provenance and packaging: the original Aspell package states LGPL 2.1, while hunspell-sv states LGPL 3.0 and combines several lexical sources. Their notices are preserved separately; compatibility and all underlying source rights still need review. The prototype is not yet in aspell-lang distribution format.
Would this be a useful starting point for modernizing the Swedish dictionary, and who should coordinate review of the language data? I have seen the aspell-lang guidance to send release-ready dictionary submissions to aspell-dict@gnu.org; I am opening this here for review of the unresolved technical questions before preparing such a release.
I would like to propose an updated Swedish dictionary based on a conservative conversion of yeager/hunspell-sv, combined with the existing
aspell-sv-0.51-0word forms.This is a request for technical review of a working prototype, not a request to publish it as a release yet.
Reproducible prototype and validation
Source scripts, pinned inputs, checksums, build instructions and local test results.
The converter expands supported suffix rules, validates candidates using libhunspell, combines them with the old Aspell word list, and builds an Aspell dictionary. Reviewed project test vocabulary is also exported explicitly after validation. The Hunspell input is pinned to
167ef93ff69812279701979e01bdc70675a859a1; related source/build corrections are in hunspell-sv PR #1.Local results with Aspell 0.60.8.2 and libhunspell 1.7 on Linux:
Open design questions
run-together true, the merged dictionary accepted errors such asöversätningandsäkerhett, as well asbilbil,datordatorandfalllucka. The prototype disables run-together, which avoids these cases but can reject correct unseen compounds. What approach would you recommend for an upstream Swedish update?excluded.txtduring the build. Would extending the Swedish character data be preferable?Would this be a useful starting point for modernizing the Swedish dictionary, and who should coordinate review of the language data? I have seen the aspell-lang guidance to send release-ready dictionary submissions to aspell-dict@gnu.org; I am opening this here for review of the unresolved technical questions before preparing such a release.