Arabic word list data, Free Tool Explorer ========================================= These files are the Arabic word lists used by the Free Tool Explorer word tools (language code ar). They are open data, built offline by "npm run words:build -- --lang=ar" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters ا ب ت ث ج ح خ د ذ ر ز س ش ص ض ط ظ ع غ ف ق ك ل م ن ه و ي ء ؤ ئ. Vowel marks, shadda, Quranic marks and tatweel are removed; presentation forms and letter variants become plain letters; the loose folds are applied (alef with hamza or madda becomes alef, teh marbuta becomes heh, alef maksura becomes yeh). Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. Ayaspell hunspell-ar, Arabic Hunspell dictionary ar.dic (as packaged in the LibreOffice dictionaries repository), LibreOffice/dictionaries commit 951bb429a4d846699b02e7adb4fc6fca128c6bae, ar/ar.dic (last changed 2018-12-29) Used for: validity list (main tier) Source: https://raw.githubusercontent.com/LibreOffice/dictionaries/951bb429a4d846699b02e7adb4fc6fca128c6bae/ar/ar.dic SHA-256: 2a3e5367f61c1583734db9d66734f5603e6be5c2d227cf5c5cd7e4ca586e34fe Licence: MPL-2.0 Credit: Arabic word list derived from Ayaspell hunspell-ar (Arabic Spellchecker for Hunspell), by Taha Zerrouki (the Ayaspell project, https://github.com/linuxscout/ayaspell) with the data collected and packaged by Mohamed Kebdani (2006-2008). Ayaspell is offered under a choice of the GNU GPL 2 or later, the GNU LGPL 2.1 or later, or the Mozilla Public License 1.1 or later; we use it under the Mozilla Public License, version 2.0 (a later version of the licence, which MPL 1.1 allows), and have modified it as described below. The MPL-2.0 text and the upstream COPYING and AUTHORS files are reproduced verbatim as MPL-2.0.txt, Ayaspell-COPYING.txt and Ayaspell-AUTHORS.txt. The modified list is the file set data/words/ar/main/*.txt (plain text, one word per line), and the program that produces it from the original dictionary is scripts/words/langs/ar.ts in this project's repository. Changes: - Read ar.dic (465,928 lines in five labelled sections: stopwords, tools, names, nouns and adjectives with a Candidate3.4 tail of proper nouns, and verb forms) as data. Dropped the whole names section (lines 13553 to 14083: letter names, continents, countries, capitals and personal names; the range is checked against the section markers of the pinned file), every entry that the dictionary files only under its own proper-noun affix classes (the aliases of the classes np and mp: the Candidate3.4 place and person names and four credit lines), the 220 place, person and company names that the nouns section carries among its ordinary nouns (cities, countries, regions and seas, rivers, persons; the entries are listed by stem in PROPER_NOUN_STEMS of scripts/words/langs/ar.ts, and a spelling whose usual reading is an ordinary word, such as the one for Egypt, is kept), every entry that does not fold to one word of 2 to 15 Arabic letters (single letters, and entries with a full stop or a colon in them) and so keeps no entry shorter than 2 or longer than 15 letters. - Dropped the particle-joined and definite-article forms that the dictionary lists as separate entries: the mass-generated combinations of the stopword section (a particle such as wa, fa, bi, li, ka or the question hamza, alone or stacked, in front of another function word, for example the entry that reads wa + kam), except the function words that carry a particle inside the word (la-qad, kadhalika, li-dhalika, li-dha, fa-hasb, la-tala-ma) and the entries of the curated function-word section, and the article forms of the function-word sections and of the fixed-form nouns (al + a word that is itself an entry: the dictionary writes the article form of the numerals, the dual and the plurals, and the names of God, as separate entries). The relative pronouns that carry the article inside the word (alladhi, allati and their dual and plural forms) and the adverb al-an are kept. - Expanded the remaining entries with unmunch from the hunspell tools, with a reduced affix file that has only the dictionary's own plural, dual and feminine suffix classes (the feminine ta marbuta of BA and BB, the dual of CA to GB, the sound feminine plural of HA to HD, the sound masculine plural of IA and IB and the same plural of a stem that is already a nisba adjective, NA and NB, through the rules that replace its final ya only: mabni gives mabniyyun, while taqlid does not give taqlidiyyun; in each class only the rules that add the feminine ending, -an or -ayn, -un or -in, or -at; the continuation flags are removed). No prefix class (the definite article, the conjunction and preposition particles, the future particle), no pronoun-suffix class, no nisba derivation (no nisba singular and no nisba plural of a stem that is not a nisba) and no tanwin-alef class is expanded. Every generated form was accepted by hunspell -l with the original, unmodified dictionary. - Folded every form to the stored form (Tier 1: harakat, shadda, tatweel and bidirectional controls removed; presentation forms composed; alef with hamza above or below or with madda folded to alef, ta marbuta to ha, alef maksura to ya; hamza, waw with hamza and ya with hamza kept) and kept the distinct words of 2 to 15 letters. display.json maps each folded word back to the spellings of the dictionary where they differ. Note: The dictionary is the copy that LibreOffice ships (Ahmad Farghal packaged Ayaspell for OpenOffice.org 3.0 in 2009; the LibreOffice README_ar.txt states no licence and points to Ayaspell), so the licence comes from the upstream Ayaspell COPYING file (repository linuxscout/ayaspell, commit 36a90abb22c5033be58e709c8980160e0ce6c88c). The ar.dic of this commit is byte-identical to the file of the research fact-check (7,217,161 bytes, 465,928 lines). 2. Ayaspell hunspell-ar, affix file ar.aff (as packaged in the LibreOffice dictionaries repository), LibreOffice/dictionaries commit 951bb429a4d846699b02e7adb4fc6fca128c6bae, ar/ar.aff (last changed 2018-02-01) Used for: validity list (main tier) Source: https://raw.githubusercontent.com/LibreOffice/dictionaries/951bb429a4d846699b02e7adb4fc6fca128c6bae/ar/ar.aff SHA-256: cec30b8621001e49618feb05aec1984c5fcfbf7d2ec309901d5cbf66585217a3 Licence: MPL-2.0 Credit: Affix file of Ayaspell hunspell-ar (Taha Zerrouki, Mohamed Kebdani), used under MPL-2.0 like the dictionary; see the ayaspell-dic entry. Changes: - Read the 243 affix classes (41 prefix classes, 202 suffix classes, FLAG long, 333 alias lines) as data and kept only the audited plural, dual and feminine suffix classes, with their continuation flags removed, in a private copy used for the expansion (see the ayaspell-dic entry). Nothing of the affix file is shipped. 3. FrequencyWords, Arabic (content/2018/ar/ar_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/ar/ar_50k.txt SHA-256: bbe98b4b92902b392bdefa2e555a108fdb42a5dd79d261674be5ab666229e19f Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form (Tier 1) and dropped those with letters outside the Arabic alphabet (punctuation, digits, Latin letters) or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list; its tokens include forms with joined particles and the definite article, which the word list does not have. - Used the 20,000 most frequent lines to restrict the main list to common words for the answers tier, leaving out the tokens that read as the definite article or a joined particle (wa, fa, bi, li, ka, the question hamza, sa) in front of another word, and the spellings of the names of places and persons that the word list drops (they reach it only through a rare verb or noun that is spelled the same), except in both cases the ordinary words that only look like one; because that selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. wordfreq is not used: its author says extracting its data into another form is not compliant. The Ayaspell list (MPL-2.0) and this list (CC BY-SA 4.0) are kept in separate files: main/*.txt holds no frequency information. 4. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), Arabic, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/ar SHA-256: 23dbd9c55257bf2a4965f46664dc7eac04b61ff4eb3571e41f7e964c7c3ed96b Licence: CC-BY-4.0 Credit: LDNOOBW, the Arabic list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries, folded them to the Tier 1 stored form on both sides (the list writes ta marbuta where the dictionary's folded form has ha), reviewed each one, and removed the ordinary, clinical, identity and everyday words the project does not block (a lioness, a draper, a heat exchanger, lick, suck, sex-organ terms of anatomy, a swinger term, intersex and lesbian as identity terms). Expanded every remaining entry to all the forms of the dictionary stems it names (the dictionary's own plural, dual and feminine classes). Phrases were dropped. 5. safetext (DeepSafe), Arabic profanity list (safetext/languages/ar/words.txt), commit afe96c8053e577e7e1e3a2eb5c9503ac157efbb5 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/viddexa/safetext/afe96c8053e577e7e1e3a2eb5c9503ac157efbb5/safetext/languages/ar/words.txt SHA-256: 99fcc05d386d89f9bf41930710fa231c2103584847d642766f627ea1ed24644d Licence: MIT Credit: safetext by DeepSafe (https://github.com/viddexa/safetext), MIT licence, Copyright (c) 2023 DeepSafe. The licence text is shipped as MIT-safetext.txt. The list is an aggregate whose upstream sources are not stated; LDNOOBW and the uxbertlabs list are credited as well. Changes: - Kept the single-word entries written in Arabic letters (most of the file is Latin transliteration or multi-word phrases, which are not used), folded them to the Tier 1 stored form and reviewed every one that is a word of the dictionary or of the frequency list: the vulgar, obscene and abusive words are blocked and expanded to the forms of the dictionary stems they name; the ordinary words the list also holds (the name of God, pronouns and prepositions with a suffix, an animal, an adjective, a verb, a mild insult, vocabulary of anatomy and identity) are not blocked. Entries that are not words of the dictionary or of the frequency list (colloquial and transliterated spellings, forms with the article) are blocked as written, as whole words only. 6. uxbertlabs/arabic_bad_dirty_word_filter_list (arabic-profanity-bad-words-dictionary.txt), commit 15f64b425d0f34c0be3e3809f5eb0002eaf6e723 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/uxbertlabs/arabic_bad_dirty_word_filter_list/15f64b425d0f34c0be3e3809f5eb0002eaf6e723/arabic-profanity-bad-words-dictionary.txt SHA-256: ae6c16617ea35da51cae2eab7250b7a8a7f7e7f6e66f5332a616aa06d6a60ac1 Licence: MIT Credit: Arabic bad-word list by UXBERT (https://github.com/uxbertlabs/arabic_bad_dirty_word_filter_list), MIT licence, Copyright (c) 2017 UXBERT. The licence text is shipped as MIT-uxbertlabs.txt. Changes: - Kept the single-word entries (14 of its 16 lines; the other two are phrases), folded them to the Tier 1 stored form and reviewed them like the safetext entries: the vulgar words are blocked and expanded to the forms of the dictionary stems they name; the everyday words the list also holds (your mother, your sister, a wineskin, the commanders) are not blocked. 7. Mozilla Public License, version 2.0 (the text published by Mozilla), 2.0 Used for: licence text Source: https://www.mozilla.org/media/MPL/2.0/index.txt SHA-256: 3f3d9e0024b1921b067d6f7f88deb4a60cbe7a78e76c64e3f1d7fc3b779b9d04 Licence: MPL-2.0 Credit: Licence text of the Ayaspell word list, as published by the Mozilla Foundation. 8. Ayaspell COPYING (the tri-licence notice: GPL 2 or later, LGPL 2.1 or later, MPL 1.1 or later), linuxscout/ayaspell commit 36a90abb22c5033be58e709c8980160e0ce6c88c Used for: licence text Source: https://raw.githubusercontent.com/linuxscout/ayaspell/36a90abb22c5033be58e709c8980160e0ce6c88c/COPYING SHA-256: 73480d918cd78d655508e80aa88c022cbd53634524a0f714b5f2519f2cf19124 Licence: MPL-2.0 Credit: Licence notice of Ayaspell, from the upstream repository (the LibreOffice copy of the dictionary states no licence). 9. Ayaspell AUTHORS (author of Hunspell-ar), linuxscout/ayaspell commit 36a90abb22c5033be58e709c8980160e0ce6c88c Used for: licence text Source: https://raw.githubusercontent.com/linuxscout/ayaspell/36a90abb22c5033be58e709c8980160e0ce6c88c/AUTHORS SHA-256: a70637e80421ea80d29c20388c211cff715b4b1896c0e99194418996b511d3d6 Licence: MPL-2.0 Credit: Author notice of Hunspell-ar (Mohamed Kebdani, 2006-2008), from the upstream Ayaspell repository. 10. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 11. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. 12. MIT licence of safetext (the LICENSE file of the repository), commit afe96c8053e577e7e1e3a2eb5c9503ac157efbb5 Used for: licence text Source: https://raw.githubusercontent.com/viddexa/safetext/afe96c8053e577e7e1e3a2eb5c9503ac157efbb5/LICENSE SHA-256: 4541b3066d5a63fcd260245287077d8a1db2bfe3e99fd09c2947525919956d95 Licence: MIT Credit: Licence text and copyright notice of safetext. 13. MIT licence of uxbertlabs/arabic_bad_dirty_word_filter_list (the LICENSE file of the repository), commit 15f64b425d0f34c0be3e3809f5eb0002eaf6e723 Used for: licence text Source: https://raw.githubusercontent.com/uxbertlabs/arabic_bad_dirty_word_filter_list/15f64b425d0f34c0be3e3809f5eb0002eaf6e723/LICENSE SHA-256: 8ec6461737ec830c8d84e099cec2cd973997b9a6ffa90fcdf1ef3a5de31732e7 Licence: MIT Credit: Licence text and copyright notice of the uxbertlabs Arabic bad-word list. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the Ayaspell hunspell-ar dictionary (ar.dic with 465885 entries in five sections, ar.aff with 243 affix classes), used under its MPL option (the authors offer "MPL 1.1 or later"; we take MPL-2.0) and modified. Dropped before the expansion: the names section (lines 13553 to 14083, 523 entries, checked against the section markers), 1940 entries filed only under the dictionary's proper-noun classes (np and mp: the Candidate3.4 place and person names and four credit lines), 223 entries (220 stems) of the nouns section that are the names of places, persons and one company (the dictionary files them among the ordinary nouns, mostly in its section of miscellaneous vocabulary; the usual reading of each spelling is the name, and a spelling whose usual reading is an ordinary word, such as the one for Egypt or Faisal, is kept), 62 entries that are not made only of Arabic letters, 11656 mass-generated particle combinations of the stopword section (a particle such as wa, fa, bi, li, ka or the question hamza in front of another function word), and 63 article forms (the article in front of a numeral, a plural, a name or a function word; a particle merged with the article; the relative pronouns and the adverb al-an are kept). 451418 entries remain. - The remaining entries were expanded with unmunch and a reduced affix file made from 20 audited suffix classes of ar.aff (170 rules, keeping only the rules that add a feminine, dual, sound feminine plural or sound masculine plural ending, with the continuation flags removed; the two nisba plural classes through the rules that replace the final ya of a stem that is already a nisba adjective only, so that a stem that is not a nisba never gets a nisba plural): no prefix class (the article, the conjunction and preposition particles, the future particle), no pronoun suffix, no nisba derivation and no tanwin form was expanded. Every one of the 99807 generated forms was accepted by hunspell -l with the original dictionary. Tier 1 folding (alef with hamza or madda to alef, ta marbuta to ha, alef maksura to ya) gives 359088 distinct words of 2 to 15 letters; display.json keeps the spellings of the dictionary. - The answers tier is the 20000 most frequent FrequencyWords tokens (of 39047 valid ones after folding) that are also main words (9355 of them), minus 20 words that carry the article (the relative pronouns and the like), minus 227 tokens that read as the article or a joined particle (wa, fa, bi, li, ka, the question hamza, sa) in front of another word although folding merges them with a rare dictionary spelling, so that they pass the test of being a main word (the 245 candidates that are ordinary words in their usual reading, such as one, duty, to him, thousands, I leave or the former, were reviewed one by one and stay), minus 24 spellings of the names that the dictionary drops (the names section, the proper-noun classes and the listed proper nouns: Rome, Sophia, Korea, Chad, Boston) that reach the main list only through a rare verb or noun spelled the same, so that they are valid words but the subtitle corpus reads each one as the name (the 43 that are ordinary words in their usual reading, such as the letter names, was, with her, he lives or he agreed, were reviewed one by one and stay), minus 45 words that are valid but not offered unasked (the policy list names 61 words: strong insults, their feminine and accusative forms and their plurals, the feminine of the adjective fallen that the corpus reads as a slur, anatomy and sex vocabulary, bodily words, identity terms, an ethnic term felt as a slur, and a pronoun form of a blocked word that the corpus reads as a name; 183 forms with the other forms of their stems), and minus the offensive words. These stay valid and unblocked. freq.txt is the frequency list's own tokens and is not joined to the Ayaspell list. - The offensive list is the single words of safetext ar, LDNOOBW ar and the uxbertlabs list, folded to Tier 1 on both sides. 133 of the 219 entries are words of the dictionary or the frequency list and were reviewed one by one: 16 are blocked (vulgar words for the sexual organs and acts, slurs for the prostitute, the pimp and the effeminate man, the vulgar word for shit) and 117 are declined (function words and phrases cut out of a curse, ordinary verbs, nouns and adjectives with a crude second sense, mild insults, the name of God, anatomy and identity vocabulary). The 86 entries that are not words of the vocabulary (colloquial and transliterated spellings, forms with the article) are blocked as written, whole words only. 20 in-house words were added (a broken plural, forms with the article or a pronoun that the frequency list has, the formal word for a prostitute, the feminine and plural of the adjective sexy, the verb behind the blocked words for the act). The feminine, dual and plural guesses of the spellings blocked as written are blocked too when they are main words (4), each one decided like any other blocked form. Every blocked word is expanded to all the forms of the dictionary stem it names (the source's own morphology): 25 stems, 150 blocked forms. 4 words that a blocked stem or a guess from a blocked spelling generates (the country Iran, a name, the forms of an ordinary stem for pruning) are never blocked; 0 forms are blocked on purpose although another stem shares them. Phrases, HurtLex and wordfreq are not used, and no semantic tag is a blocklist. - Five audits run inside the build and stop it on an undecided word: every main word of a blocked stem is blocked or protected; every blocked form that another stem or a top-5000 frequency token shares is protected or accepted on purpose; every stem that derives from a blocked root (12 roots; 34 stems derive from them without being blocked, and 34 of them are reviewed as ordinary) is blocked, made of blocked or protected forms, or reviewed; every frequency token that no stem generates and that matches a root, also behind the article or a joined particle (0 tokens; 0 reviewed), is blocked, protected or reviewed; every source entry that is a word of the vocabulary is blocked or declined with a reason. Arabic derives words from roots and patterns, which a list of letters cannot follow, so the root checks find the forms that keep the root's letters in a row (with the usual prefixes and suffixes) and the broken plurals are listed by hand. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 359088 words extended 0 words answers 9036 words frequency 39028 tokens offensive 150 words LICENCE TEXTS IN THIS FOLDER ---------------------------- Ayaspell-AUTHORS.txt Ayaspell-COPYING.txt CC-BY-4.0.txt CC-BY-SA-4.0.txt MIT-safetext.txt MIT-uxbertlabs.txt MPL-2.0.txt