Italian word list data, Free Tool Explorer ========================================== These files are the Italian word lists used by the Free Tool Explorer word tools (language code it). They are open data, built offline by "npm run words:build -- --lang=it" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a b c d e f g h i j k l m n o p q r s t u v w x y z. Accents are removed. All 26 letters are accepted in the stored form; the 21-letter native alphabet is applied by the tools that need it. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. Morph-it! (a free morphological lexicon for the Italian language), 0.4.8, 23 February 2009 Used for: validity list (main tier) Source: https://docs.sslmit.unibo.it/lib/exe/fetch.php?media=resources:morph-it.tgz SHA-256: b137343dc095e038ebfaef7c743185e85ec35f5322a1b8629b600e6e898c7024 Licence: LGPL-2.1-only Credit: Morph-it! Copyright (c) 2004-2009 Marco Baroni and Eros Zanchetta (SSLMIT, University of Bologna, http://sslmit.unibo.it/morphit; info page https://docs.sslmit.unibo.it/doku.php?id=resources:morph-it). The authors dual-license it under the GNU Lesser General Public License and the Creative Commons Attribution-ShareAlike 2.0 licence; this project uses and redistributes it under the LGPL option. The LGPL 2.1 text (lgpl.txt of the archive) and the authors' README with both licence notices are reproduced verbatim as LGPL-2.1.txt and Morph-it-README.txt. Changes: - Read the lexicon (ISO-8859-1, one form, lemma and tag per line), dropped the tags NPR (proper nouns), ABL (abbreviations), PON (punctuation), SYM (symbols), SENT (sentence markers) and SMI (emoticons), dropped every form that is not written entirely in lowercase letters, removed accents, and kept only entries of 2 to 15 letters a-z, so forms with an apostrophe, hyphen, space, digit or full stop are dropped too. - Used the lemma column to expand the blocked words and the words kept out of the answers tier to their other inflected forms (the lemma column is not shipped). - Left out of the spelling list that feeds display.json the rows that only repeat their lemma under another accent (perchè and perche for perché, piú and piu for più, affinchè, nonchè, menù for menu) and a reviewed list of corpus misspellings that Morph-it! files as words of their own (citta, caffe, ventitre and similar), so that the accented spelling is the one shown. This changes which spellings are shown, not which words are in the list. - Replaced, in that spelling list, five corpus misspellings that are the only spelling Morph-it! records for their word by the correct one (ahimé, ohimé, potè and koiné are shown as ahimè, ohimè, poté and koinè; pò, the form of the lemma po', is shown as po because po' cannot be stored). Each replacement folds to the same stored form, so the list of words does not change. - Kept the loanword spellings (j, k, w, x, y) in the main list; the answers tier uses the 21-letter native alphabet only. Note: The README that ships with the data says the LGPL grant is 'version 2 or (at your option) any later version' in its notice text while the bundled licence text is LGPL 2.1, so the exact grant is not fully clear; both readings allow redistribution under the LGPL. The wiki footer of the authors' info page says CC BY-NC-SA 4.0 'except where otherwise noted'; the data carries its own dual licence notice (above), which is what applies to the lexicon, and the footer is recorded here for transparency. The authors warn that the data comes from a newspaper corpus (la Repubblica) of 2009 and has many gaps in everyday vocabulary. answers/*.txt is a subset of these words, picked with the help of the FrequencyWords ranking (see the next entries). It is distributed under GPL-3.0-only, one licence for the whole tier: LGPL-2.1 section 3 lets a recipient apply the GNU GPL to a copy of the lexicon instead of the LGPL, and CC BY-SA 4.0 material may be adapted and shared under GPL 3, so both inputs are covered; the GPL 3 text is shipped as it_IT-README-GPL-3.0.txt. main/*.txt stays under the LGPL option (LGPL-2.1-only). This choice is among the copyleft questions for counsel (plan risk table, owner decision C2). 2. Italian spelling dictionary it_IT of LibreItalia (LibreOffice dictionaries), dictionary file it_IT.dic, 5.1.1 (7 November 2022), repository commit b871dff9e73eefc187b2c979803c619ceccaed3a Used for: accept-only list (extended tier) Source: https://raw.githubusercontent.com/LibreOffice/dictionaries/b871dff9e73eefc187b2c979803c619ceccaed3a/it_IT/it_IT.dic SHA-256: bae1e3501dcd2a923669592493b3fde6c02aae7c7aab83bf5e5b49077e73dd64 Licence: GPL-3.0-only Credit: Italian spelling dictionary it_IT (Estensione linguistica italiana, Italian Writing Aids extension), version 5.1.1 of 7 November 2022, Copyright (C) 2020-2022 LibreItalia - Marina Latini, with portions Copyright (C) 2001-2015 Gianluca Turconi, Davide Prina, Andrea Pescetti and other authors (https://libreitalia.org; https://github.com/LibreOffice/dictionaries/tree/master/it_IT), licensed GNU GPL version 3. extended/*.txt is derived from it and is shared under the same licence. The README of the extension, which carries these notices and the GPL 3 text, is reproduced verbatim as it_IT-README-GPL-3.0.txt. Changes: - Used only as a spelling oracle, with the hunspell program and the original it_IT.dic and it_IT.aff: the frequency tokens that Morph-it! lacks (written in lowercase, accents kept) were checked with hunspell -l, and the tokens it accepts are the candidates of the extended tier. The dictionary is not expanded and none of its own entries is copied. - Removed from the candidates, one by one after reading every one of them: Roman numerals, unit symbols and abbreviations, English words, given names, surnames, place names and brands, prefixes, apocopes and fragments of Latin and other foreign phrases (the list is in scripts/words/langs/it.ts). What remains is extended/*.txt: accept-only words that main lacks, never in answers. Note: The licence is GPL version 3 only (the wording says 'version 3' without 'or later'). The extended files are kept apart from main and answers, with this licence text shipped next to them. The extended selection also uses the frequency tokens (CC BY-SA 4.0), which that licence allows to be adapted under GPL 3. Hunspell's own output was used, not the dictionary's entries, so the file holds only words (facts about the language) that the dictionary accepts. 3. Italian spelling dictionary it_IT of LibreItalia (LibreOffice dictionaries), affix file it_IT.aff, 5.1.1 (7 November 2022), repository commit b871dff9e73eefc187b2c979803c619ceccaed3a Used for: accept-only list (extended tier) Source: https://raw.githubusercontent.com/LibreOffice/dictionaries/b871dff9e73eefc187b2c979803c619ceccaed3a/it_IT/it_IT.aff SHA-256: 951afaa19272f13555b8823e8bcf9ccf78f8fe1a07835bdfb912ab3e4d537c2b Licence: GPL-3.0-only Credit: Italian spelling dictionary it_IT (Estensione linguistica italiana, Italian Writing Aids extension), version 5.1.1 of 7 November 2022, Copyright (C) 2020-2022 LibreItalia - Marina Latini, with portions Copyright (C) 2001-2015 Gianluca Turconi, Davide Prina, Andrea Pescetti and other authors (https://libreitalia.org; https://github.com/LibreOffice/dictionaries/tree/master/it_IT), licensed GNU GPL version 3. extended/*.txt is derived from it and is shared under the same licence. The README of the extension, which carries these notices and the GPL 3 text, is reproduced verbatim as it_IT-README-GPL-3.0.txt. Changes: - Read by hunspell together with it_IT.dic (see the dictionary file above); not copied or changed. 4. FrequencyWords, Italian (content/2018/it/it_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/it/it_50k.txt SHA-256: bb96cdcb56d28342c1e909db6b2525448b7767136b7e86a2ccc649a1be66fc19 Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form and dropped those with letters outside a-z or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the 20,000 most frequent lines to pick, among the Morph-it! words, the common words written with the 21 native letters for the answers tier. answers/*.txt holds only Morph-it! words, so it is not a CC BY-SA 4.0 adaptation shared under that licence: it is distributed under GPL-3.0-only (see the Morph-it! note), a licence that accepts CC BY-SA 4.0 adaptations. Only freq.txt, which lists this source's own tokens, is shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. 5. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), Italian, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/it SHA-256: d413f41018d58aa79d2c4809d0161f29ffd5652c4dc34c5b7b9ade680ee5599a Licence: CC-BY-4.0 Credit: LDNOOBW, the Italian list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries, folded to the stored form, expanded them to their inflected forms through the Morph-it! lemma column (a form that is also a form of an ordinary word is not blocked), and removed the ordinary words and the words the project only keeps out of the answers tier (see the notes below). Phrases were dropped. 6. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 7. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. 8. README_it_IT.txt of the it_IT extension: the project's notices and the GNU General Public License version 3, 5.1.1, repository commit b871dff9e73eefc187b2c979803c619ceccaed3a Used for: licence text Source: https://raw.githubusercontent.com/LibreOffice/dictionaries/b871dff9e73eefc187b2c979803c619ceccaed3a/it_IT/README_it_IT.txt SHA-256: 34c3e93595cf3cf5f3afc9bc0d98eea12593750383f1e82ed1d3ca29f9681283 Licence: GPL-3.0-only Credit: Licence text and notices of the it_IT dictionary (LibreItalia and the earlier authors of Dizionario italiano), as shipped with the extension. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The extended tier holds 2853 words that the Italian spelling dictionary it_IT (LibreItalia, GPL 3) accepts and Morph-it! lacks. The candidates are the FrequencyWords tokens that main does not have (14017); each was checked with hunspell -l against the original dictionary (3285 words accepted, written in lowercase, accents kept), and every accepted word was read: 432 that are not Italian words (Roman numerals, unit symbols and abbreviations, English words, given names, surnames, place names and brands, prefixes, apocopes and pieces of foreign phrases) were removed. The dictionary itself is not expanded or copied. These files are shared under GPL 3, separate from main and answers, and are never used by the answers tier. - The main list is Morph-it! 0.4.8 (505074 lines read as ISO-8859-1: 505072 with a form, a lemma and a tag; malformed lines skipped: 2). The tags NPR, ABL, PON, SYM, SENT and SMI were dropped, as was every form that is not written entirely in lowercase letters; accents were then removed and only forms of 2 to 15 letters a-z were kept, so forms with an apostrophe, hyphen, space, digit or full stop are dropped. Loanword spellings with j, k, w, x or y stay in the main list. - The display spellings in display.json come from the Morph-it! spellings, minus 8 rows that only repeat their lemma under another accent (perchè and perche for perché; cosi is kept, the plural of coso) and 24 corpus misspellings that Morph-it! files as words of their own (citta, caffe, ventitre), each checked against the it_IT dictionary. 5 spellings that are the only one Morph-it! records for their word are shown corrected (ahimé as ahimè, potè as poté, and pò as po, because its lemma po' cannot be stored). The main word list is unchanged by this: every spelling left out folds to a word that another spelling keeps, and every corrected spelling folds to the same stored form. - The answers tier is the 20000 most frequent FrequencyWords lines that are also main words written with the 21 native letters (no j, k, w, x, y), minus the offensive words and minus 49 entries (with the other forms of their lemma) that are valid words not offered unasked: ordinary words with a vulgar or insulting sense, ethnic and disability terms many people find offensive, and adult vocabulary. A form that an ordinary word also has is not taken through the lemma: it is either an entry of its own (peni, the plural of pene and a form of penare) or a reviewed form that stays in the answers tier (ammucchiate, porci). - The offensive list is LDNOOBW plus 134 in-house words, expanded to their inflected forms through the Morph-it! lemma column; a form that is also a form of an ordinary word (scopo, the noun for "purpose", is also a form of a vulgar verb) is not blocked through the lemma. A protected list of 40 words is never blocked: everyday words the topic-based list flags (battere, tirare, regina, pesce) and identity words. The words LDNOOBW lists that are ordinary in their main sense (a bean, a pea, a cow, a saw) are not blocked either; they are only kept out of the answers tier. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 371289 words extended 2853 words answers 15325 words frequency 47682 tokens offensive 1079 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-4.0.txt LGPL-2.1.txt Morph-it-README.txt it_IT-README-GPL-3.0.txt