Attribution and licences of distributed model data

The keyboard's language models are trained on, or contain data from, the sources below. Models built from them are distributed under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) with this attribution. Personal models (trained on the owner's own writing) are never distributed.

File Source Licence
lexicon.kbng (general word list) wordfreq word frequencies, by Robyn Speer CC BY-SA 4.0
generic.kbng, generic-neural.mlpackage (British English base) English Wikipedia, plain text (wikimedia/wikipedia, 2023-11-01), by Wikipedia contributors CC BY-SA 3.0 / 4.0 (and GFDL)
Stack Exchange data dump (English Language, English Language Learners, DIY, Cooking, Travel, Money, Parenting, Pets, Gardening, Bicycles, The Great Outdoors, The Workplace, Interpersonal Skills, Academia, Movies & TV, Photography, Motor Vehicle Maintenance, Lifehacks, Super User, Ask Ubuntu, Ask Different, Unix & Linux), by Stack Exchange contributors CC BY-SA 2.5 / 3.0 / 4.0

| generic-terminal.kbng (shell commands for terminal mode) | tldr-pages example commands (common, Linux, macOS pages), by the tldr-pages contributors | CC BY 4.0 |

Tools used to build them (not distributed): breame American → British spelling list (Apache-2.0).

Changes: text was sampled, cleaned (markup, code, tables removed), redacted (URLs, email addresses, numbers replaced by placeholders), converted to British spelling, tokenised and used to estimate word frequencies and train a neural next-word model. No text is reproduced verbatim beyond single words.

Download the models

The language models in the app, under the licences above. The file format is described in spec/model-format.md of the source.