Spanish word list data, Free Tool Explorer ========================================== These files are the Spanish word lists used by the Free Tool Explorer word tools (language code es). They are open data, built offline by "npm run words:build -- --lang=es" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a b c d e f g h i j k l m n ñ o p q r s t u v w x y z. The n with tilde is kept as its own letter; every other accent and the diaeresis are removed. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. RLA-ES (Recursos Libres para el Aprendizaje del Español), Spanish Hunspell dictionary es.dic and es.aff, v2.9 (2025-01-02), es.oxt Used for: validity list (main tier) Source: https://github.com/sbosio/rla-es/releases/download/v2.9/es.oxt SHA-256: b08a1a0e3e044697f63a67184f591f7e2c37bbb53bbfbb4780bcbd84929d6e8c Licence: MPL-1.1 Credit: Spanish word list derived from RLA-ES (Recursos Libres para el Aprendizaje del Español), Copyright Santiago Bosio and contributors (https://github.com/sbosio/rla-es). The authors offer RLA-ES under a choice of three licences (GPL-3.0-or-later, LGPL-3.0-or-later or MPL-1.1-or-later); we use it under the Mozilla Public License, version 1.1 or later, and have modified it as described below. The MPL-1.1 text and the authors' own LICENSE.md, which states the choice they offer, are reproduced verbatim as MPL-1.1.txt and RLA-ES-LICENSE.md. Changes: - Dropped the dictionary entries that are not made only of lowercase letters (capitalised proper nouns and acronyms, mixed-case brand names, and the two entries with a space or a soft hyphen), and, after the expansion, the lowercase abbreviations (cm, apdo, aprox) that the noRAE Abreviaturas lists name, except those that also stand in the RAE lemma lists (del). - Expanded the remaining stems into every inflected and derived form with unmunch from the hunspell tools: the affix file was first copied with every flag renamed to one byte (unmunch reads only the first byte of a flag, and the original file uses 99 UTF-8 flags), the forms of each stem were kept together, and a second pass applied the continuation flags of two-step suffixes. Every generated form was accepted by hunspell -l with the original dictionary, and every frequency token that hunspell accepts was among the generated forms. - Folded the forms to the stored form (the n with tilde is kept; every other accent and the diaeresis are removed), and kept only words of 2 to 15 letters. Note: The release bundles a thesaurus and hyphenation patterns as well; they are not used and not shipped. 2. RLA-ES lemma lists (ortografia/palabras), source repository snapshot, tag v2.9, commit ea82c1214ead57740798acf66a1e18e5ac874c41 Used for: accept-only list (extended tier) Source: https://github.com/sbosio/rla-es/archive/ea82c1214ead57740798acf66a1e18e5ac874c41.tar.gz SHA-256: bf76b4daa9958bc43b9e602c51ae3a1ee29859831f9137107a0aa1940d580cf8 Licence: GPL-3.0-or-later Credit: RLA-ES lemma lists, Copyright Santiago Bosio and contributors (https://github.com/sbosio/rla-es), modified as described below. The header of every list file names the GNU General Public License, version 3 or later; the repository's LICENSE.md (shipped as RLA-ES-LICENSE.md) offers the project and its dictionaries under a choice of GPL, LGPL or MPL. Because the extended tier is read from these list files, it is shared under GPL-3.0-or-later; the GPL text is shipped as GPL-3.0.txt. Changes: - Read the lemma lists of palabras/RAE and the Abreviaturas lists of palabras/noRAE only to decide which lowercase abbreviation entries of the dictionary are ordinary words (an abbreviation that is also an RAE lemma, such as del, is kept). - Read the regional lemma lists of palabras/noRAE/l10n (words used in particular Spanish-speaking countries; abbreviations and proper nouns left out), expanded them with the same affix file as the main list, and kept the forms that are not in the main list as the extended tier (accepted by the word checker only). The generic dictionary already holds every regional lemma, so no form is left and the extended tier is empty. Note: The archive is a GitHub snapshot of the commit that the v2.9 tag points to, so its bytes are pinned by SHA-256; only palabras/RAE, palabras/noRAE and the GPL text are unpacked. 3. FrequencyWords, Spanish (content/2018/es/es_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/es/es_50k.txt SHA-256: dcff3ad4316192f4dc4ff7d26e637c6ff314ef1ca0f3f720c5649018a71056c0 Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form and dropped those with letters outside the Spanish alphabet or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the 20,000 most frequent lines to restrict the main list to common words for the answers tier; because that selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. 4. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), Spanish, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/es SHA-256: 073334261cc2e7c08339faefd2f19f9623e8240a0ac18fe2a26c368b897496ce Licence: CC-BY-4.0 Credit: LDNOOBW, the Spanish list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries and folded them to the stored form, removed the ordinary, topic and identity words the project does not block (murder, drugs, hell, a donkey, a shell, a trio, urine, sex, a prostitute, a cross-dresser and similar), and expanded every remaining entry to all the forms of the dictionary stem it names (the dictionary's own affix classes) and to its own gender and number forms that are real words, without touching a form that an unrelated, unblocked stem shares unless a collision was accepted on purpose. Phrases were dropped. 5. cuss, Spanish (es.js), commit 6bab3fef250481e34ba55bc400fac5c6d25f1429 (cuss 2.2.0) Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/words/cuss/6bab3fef250481e34ba55bc400fac5c6d25f1429/es.js SHA-256: c924edc8f94561699a97f767d511804fe2ce448a8562f4a72eb0296f7d089679 Licence: MIT Credit: cuss by Titus Wormer (https://github.com/words/cuss), MIT licence, Copyright (c) 2016 Titus Wormer. The licence text is shipped as MIT-cuss.txt. The Spanish list was compiled by its author from several public sources, including LDNOOBW. Changes: - Read the file as data (it is never run), kept only the entries with the top sureness rating of 2, reviewed every one of them by hand, blocked the words that are vulgar, obscene or slurs in everyday use, and declined the ordinary words and the mild insults it also lists (a bun, a leech, a fool, a simpleton). Every blocked word was expanded to all the forms of the dictionary stem it names and to its own gender and number forms that are real words. Entries with a lower rating and phrases were not used. 6. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 7. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. 8. MIT licence of cuss (the license file of the repository), commit 6bab3fef250481e34ba55bc400fac5c6d25f1429 Used for: licence text Source: https://raw.githubusercontent.com/words/cuss/6bab3fef250481e34ba55bc400fac5c6d25f1429/license SHA-256: ca4662cb5d1b738fbe5350c0d5485ba11773b4b7208974082ae6e129a52d631d Licence: MIT Credit: Licence text and copyright notice of cuss. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the RLA-ES v2.9 dictionary (es.dic with 71198 entries, es.aff with 99 affix classes), used under its MPL option and expanded with unmunch into every form: 606684 words after folding. Entries that are not made only of lowercase letters were dropped before the expansion (11527: capitalised proper nouns and acronyms, mixed-case brand names, one entry with a space and one with a soft hyphen), and so were 129 lowercase abbreviations (cm, apdo, aprox) that the noRAE Abreviaturas lists name and the RAE lemma lists do not (del, which both name, is kept). Because unmunch reads only the first byte of a flag, the affix flags were renamed to single bytes in a private copy of the dictionary; the build runs unmunch twice (the second pass applies the continuation flags of two-step suffixes). Every one of the 707649 generated forms was then accepted by hunspell -l with the original dictionary, and each of the 37887 frequency tokens that hunspell accepts was among the generated forms. Accents and the diaeresis were removed (the n with tilde is kept) and only forms of 2 to 15 letters were kept. - The extended tier is meant for the words of the regional lemma lists (palabras/noRAE/l10n, 116 files, 845 entries) that the main list lacks. They were expanded with the same affix file: 0 forms are not in the main list, because the generic RLA-ES dictionary already holds every regional lemma, so the extended tier is empty. - The answers tier is the 20000 most frequent FrequencyWords lines that are also main words (16246 of them), minus the offensive words and minus 81 words that are valid but not offered unasked: the policy list names 71 words (childish bodily words, ordinary words with a common crude sense, mild insults and adult vocabulary), which with the other forms of their stems are 358 forms. A form that a dictionary stem outside the list also generates (a ring, a piece, an instrument: 26 forms) is an ordinary word and goes with the stem only when the list names it, and 1 form that only an excluded stem generates is kept on purpose. These stay valid and unblocked. - The offensive list is the single words of LDNOOBW es (minus 27 ordinary, topic and identity words it flags) and the 53 entries of words/cuss es with rating 2 that survived a review of all 274 of them, plus 63 in-house words (derived words of the blocked roots that the audit found, and words the frequency list has but the dictionary lacks). Each is expanded to every form of the dictionary stem it names, which is the source's own morphology, and to the gender and number forms (bolleras, malparidas, pendeja) of every blocked entry that is a real word, because a blocked entry can be a form of another stem whose paradigm has none: 87 stems, 1134 blocked forms. 14 ordinary words that a blocked stem also generates (the bellows, a straw loft, youth, a cheese, a verb form) are never blocked; 17 forms are blocked on purpose although an unrelated stem shares them. Entries with a lower rating, phrases and HurtLex are not used, and no semantic tag is used as a blocklist. - Five audits run inside the build and stop it on an undecided word: every main word of a blocked stem is blocked or protected; every blocked form that an unblocked stem shares is protected or accepted on purpose; every stem that derives from a blocked root, or that starts or ends with a blocked form of 5 or more letters (64 roots; 92 stems derive from them without being blocked, and 91 reviewed stems record why they are ordinary words), is blocked, made of blocked or protected forms, or reviewed; every frequency token that no dictionary stem generates and that matches a root or the compound rule (9 tokens; 9 of them reviewed as ordinary) is blocked, protected or reviewed, so the published frequency list holds no vulgar token the paradigms miss; every word of the two source lists is blocked or declined with a reason. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 606684 words extended 0 words answers 16106 words frequency 46948 tokens offensive 1134 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-4.0.txt GPL-3.0.txt MIT-cuss.txt MPL-1.1.txt RLA-ES-LICENSE.md