Research findings
Is wordfreq useful for our Arabic decks?
Short answer: as a coverage sanity-check against written Arabic, yes.
As a priority list for a spoken Palestinian course, no — it would teach the wrong word first.
Measured against wordfreq 3.1.1 (Robyn Speer), Arabic “large” list of
619,906 words. Its ten most frequent Arabic tokens: في · من · على · أن · لا · إلى · و · ما · عن · هذا.
Applicable — as a gap-check
Coverage sanity-check ✓
wordfreq recognises 98.9% of our Palestinian single words and 100.0% of the MSA ones.
It’s a fair way to spot common written words we haven’t taught.
Not applicable — as a priority list
Frequency ranking ✗
Its data is Modern Standard Arabic. Where a Levantine word differs from the textbook word,
the textbook word ranks higher 78.7% of the time.
01The coverage numbers
We compared two of our decks against wordfreq: the Palestinian deck (our real, native-reviewed
dialect deck) and the MSA deck (the fair like-for-like, since wordfreq’s Arabic is mostly written MSA).
Percentages are over single-word entries; multi-word idioms are counted separately.
265
Palestinian single words analysed (+23 idioms)
43.0%
of them sit in wordfreq’s top 5,000 written words
51.8%
of the MSA deck sits in the top 5,000 — the register gap
78.7%
of differing word-pairs: the MSA form is commoner in writing
wordfreq strips diacritics and recognises almost everything we teach — so “does it know our words?”
is the wrong question. The signal is where words sit. Our Palestinian words cluster lower in the
frequency distribution than the MSA words, because everyday spoken forms are simply rarer in the
Wikipedia-and-news text wordfreq is built from.
02Same meaning, different word — who wins in writing?
Our decks are concept-aligned, so for every idea taught in both we can line the Levantine word up against
the MSA word. Across 47 concepts where the two differ, the written (MSA) form is more
common 37 times; the spoken (dialect) form wins only 2.
Longer bar = more common in written Arabic (Zipf scale, higher = commoner).
Palestinian (spoken)
MSA (written)
بقديش
b'addeish · How much? (alternative)
زغير
zghiir · Small, little
زهقان
zah'aan · Bored, fed up
مهضوم
mahdoom · Charming, likeable
مزنوق
maznoo' · Pressed for time; in a tight spot
فرحان
farhaan · Glad, joyful
This is the crux: sort a Levantine syllabus by wordfreq and you’d front-load
صغير, جيد,
غدا over the words people actually say —
زغير, منيح,
بكرا.
03How common are our words, really?
The same 265 Palestinian and 164 MSA single words, bucketed by Zipf value. Zipf 6 ≈ a word you meet
about once per thousand; Zipf 3 ≈ once per million. The MSA deck leans right (commoner in writing);
the Palestinian deck spreads left, with a genuinely dialectal tail wordfreq barely sees.
Palestinian deck
by Zipf band · median 4.28
30 (unknown)
40.01–2
252–3
663–4
1134–5
505–6
46+
MSA deck
by Zipf band · median 4.47
00 (unknown)
00.01–2
32–3
403–4
824–5
395–6
06+
04What that tail looks like
Words we teach that are common in writing too
- هُوّ huwwe He #16
- هِيّ hiyye She #30
- تِمّ timm Mouth #43
- يَوْم yoom Day #51
- أَنَا ana I #96
- كِيف kiif How #99
- وَاحَد waahad One #121
- هُمّ humme They #146
- سَنَة sane Year #157
- كْبِير kbiir Big, large #163
Core Levantine words wordfreq barely knows
- تُقْبُرْنِي tu'burni I’d rather die first than lose you Zipf 0.0
- بْقَدِّيش b'addeish How much? (alternative) Zipf 0.0
- إِيمْتَى eimta When Zipf 0.0
- مَزْنُوق maznoo' Pressed for time; in a tight spot Zipf 1.6
- زَهْقَان zah'aan Bored, fed up Zipf 1.8
- سِرْفِيس serviis Shared taxi Zipf 1.9
- تْمَانْيَة tmaanye Eight Zipf 2.0
- عَطَى 'ata To give Zipf 2.0
- إِجَا ija To come Zipf 2.1
- زْغِير zghiir Small, little Zipf 2.2
Look at the right-hand column: تمانية (eight),
عطى (to give), إجا (to come),
زغير (small). These are beginner essentials that wordfreq ranks
as obscure — the clearest sign its ranking can’t drive a dialect syllabus.
05Common written words neither deck teaches
wordfreq’s most frequent Arabic words that appear in neither deck. Read the caveat first:
most are MSA grammar words (الذي “which”,
لم negation, قد a verb particle)
that Levantine either doesn’t use or forms differently — so “missing” here often means “correctly absent
from a spoken course”, not a real gap.
- أن#4 · Zipf 7.0
- إلى#6 · Zipf 6.9
- و#7 · Zipf 6.9
- ما#8 · Zipf 6.8
- عن#9 · Zipf 6.8
- هذا#10 · Zipf 6.7
- التي#12 · Zipf 6.6
- كل#13 · Zipf 6.6
- هذه#14 · Zipf 6.6
- أو#15 · Zipf 6.6
- كان#17 · Zipf 6.5
- الذي#18 · Zipf 6.5
- ذلك#19 · Zipf 6.5
- بعد#20 · Zipf 6.4
- الله#21 · Zipf 6.4
- لم#22 · Zipf 6.4
- بين#23 · Zipf 6.3
- ان#24 · Zipf 6.3
- كانت#25 · Zipf 6.3
- حتى#27 · Zipf 6.3
- قبل#28 · Zipf 6.3
- قد#29 · Zipf 6.3
- إن#31 · Zipf 6.3
- كما#32 · Zipf 6.3
06What wordfreq is (and its expiry date)
wordfreq is a Python library by Robyn Speer giving word-frequency data for 40+ languages. Its Arabic
data is pooled from Wikipedia, news crawls, film subtitles, the OSCAR web corpus and Twitter —
overwhelmingly Modern Standard Arabic. It’s the standard open frequency source, but it stopped in 2024:
“I don’t think anyone has reliable information about post-2021 language usage by humans… large language
models generate text that masquerades as real language.”
— Robyn Speer, wordfreq sunset note, Sept 2024
- Language
- Arabic (
ar) — one shared code for all dialects and MSA
- Register
- Mostly written / Modern Standard Arabic
- Frozen at
- ~2021 language usage; no further data updates
- Diacritics
- Stripped — it matches bare forms, so we normalise our vocalised text before comparing
- Clitics
- Not split — the definite article ال and prefixes stay attached to the token
- Licence
- Apache-2.0 (code); data CC-BY-SA-4.0
- API used
zipf_frequency, top_n_list
07Recommendation
- Use it as a one-off gap-check. Run the script now and then to catch high-frequency
written words we’ve missed, and reviewed together against whether they belong in a spoken course.
- Don’t rank or order the decks by it. It would systematically bury the spoken forms
(زغير, منيح,
بكرا) our learners actually need.
- If we want a real frequency signal later, it has to come from a spoken/dialect
corpus — Levantine subtitle collections or a spoken-Arabic corpus — not from written MSA.