Sources
Attribution & Licences
Everything on this site is built from open data. The vocabulary, the reading texts and the questions derived from them are published under CC BY-SA 4.0, inherited from Wiktionary and Klexikon. If you reuse any of it, keep this attribution reachable.
Where the data comes from
| Source | Provides | Licence |
|---|---|---|
| Wiktionary via wiktextract | Gender, article, plural, genitive, verb principal parts, auxiliary, comparative and superlative, English glosses, IPA, audio links, topic labels, synonyms, antonyms | CC BY-SA 4.0 |
| Tatoeba | Example sentences, and the corpus counts used to pick each word's primary reading | CC BY 2.0 FR |
| FrequencyWords (OpenSubtitles 2018) | The frequency ranking every list on this site is sorted by | CC BY-SA 4.0 |
| OpenSubtitles v2024 via OPUS | Most example sentences, mined from 65.7 million German–English pairs | See note below |
| Klexikon | Every reading text, and the passages the comprehension questions are built on | CC BY-SA 4.0 |
| minddory.com | The word list and its CEFR level assignments | None stated |
OpenSubtitles asks that work using the corpus link to opensubtitles.org and cite Lison & Tiedemann (2016), OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles (LREC). OPUS states that it does not own the underlying text and redistributes only what it believes it may.
Pronunciation audio is linked to Wikimedia Commons, not copied. Individual recordings carry their own licences — check the Commons file page before redistributing a file itself.
Machine-generated content
Three things on this site were written by a language model rather than a person:
- The 2,303 practice questions and their explanations.
- The sentence-by-sentence English translations of the reading texts.
- The glossaries and grammar notes attached to each text.
They are derivative of the Klexikon passages and the vocabulary examples they were built from, so they inherit CC BY-SA 4.0.
They are machine-written and machine-checked, not reviewed by a teacher. The checks were real ones — comprehension items must quote the passage verbatim; gap-fill items must restore a word that genuinely occurs in the source; every word in a stem or option is checked against the vocabulary list up to that level; and a second model pass answered each item blind with the options already shuffled, with disagreements dropped rather than kept. But automated checks are not the same as review. Sample before putting any of it in front of learners, and treat anything that looks wrong as possibly wrong.
Known limitations
- The source word list was fully lowercased; capitalisation is restored only for words matched in Wiktionary.
- Homographs differing only by case (essen / Essen) are collapsed in the source; the alternative reading is kept alongside the entry.
- Verb government — the case and preposition a verb takes — is not available as structured data in Wiktionary and is therefore absent here.
- The level distribution is skewed: C1 holds about 42% of the list, so treat C1 as a soft boundary rather than a precise one.
- Reading coverage tops out near 90% even at C2, because roughly a quarter of Klexikon's running words fall outside a 7,164-word list entirely.
Reusing this material
CC BY-SA 4.0 lets you copy, adapt and redistribute all of it, including commercially, on two conditions: credit the sources above, and license what you build under the same terms. A link back to this page satisfies the attribution requirement.