The short answer is five new vocabulary items a day for two weeks—then adjust. Five is a conservative starting experiment, not a scientifically optimal quota.
The research reviewed here does not establish one daily number that works across learners, languages, goals, item types, schedules, and available study time. This is a limited synthesis of studies concerning learning activities, spacing, practice, and item variables—not proof that no relevant daily-quota study exists anywhere. The reviewed studies instead document substantial variation in activity outcomes, practice conditions, and predictors of vocabulary knowledge. (Wiley, Cambridge University Press)
A practical quota can be set as the lower of two limits:
- The number of useful items your real reading supplies.
- The number whose future reviews fit your study budget.
That means zero can be the right number on a busy day. It also means ten may be reasonable for a learner with ample time, a manageable review queue, and plenty of relevant input. The aim is not to win a daily counting contest. It is to build vocabulary without letting collection and review displace actual language use.
First, decide what counts as one word
A numerical target is meaningless until you define the unit being counted. In vocabulary research, a lemma generally groups a headword with inflected forms belonging to the same part of speech, while a word family may also include derived forms. The appropriate unit depends on the language and the purpose of the count. (Cambridge University Press)
For personal study, use a simpler operational definition:
One learning item = one target word or phrase, one relevant meaning, and one saved context.
Under this rule, the English forms runs, ran, and running would not automatically become four items if you are learning ordinary forms of run. But a distinct meaning or fixed phrase might deserve its own item. For example, run a company could be separate from run five kilometres if both meanings matter to you.
Consistency matters more than choosing the theoretically perfect unit. Avoid switching between counting surface forms, dictionary headwords, translations, phrases, and meanings while treating them as equivalent.
A saved item is also not the same as a mastered word. Vocabulary knowledge develops incrementally and can include written form, sound, meaning, grammatical behaviour, derivation, associations, collocations, and constraints on use. Many studies test only one or two of those dimensions. (Cambridge University Press)
Why a universal daily quota is misleading
A quota such as “learn 20 words every day” hides differences in both the words and the learners. Item difficulty can be associated with properties including frequency, concreteness, length, part of speech, and similarity to a learner’s other languages, although findings for some properties are mixed or dependent on the task. Learner proficiency and educational background can also predict receptive vocabulary knowledge. These associations do not determine how difficult a given item will be for an individual. (Cambridge University Press)
Goals change the workload too. Recognising a word’s meaning while reading is a narrower assessment target than recalling its exact form, pronouncing it, choosing an appropriate collocation, and producing it independently. Vocabulary knowledge is multidimensional, and studies use tests with different recognition, recall, and production demands. Two learners who each add five items may therefore be preparing for substantially different assessments and uses. (Cambridge University Press)
Immediate performance can also create false confidence. A meta-analysis of intentional vocabulary activities reported average immediate gains of 60.1% for meaning recall and 58.5% for form recall. On delayed tests, the corresponding averages were 39.4% and 25.1%. The included activities and studies varied considerably, so these figures are not personal forecasts. They show why an immediate result should not be assumed to predict performance on a later assessment. (Wiley)
Repeated practice can improve measured learning outcomes, but more repetitions are not free. In one controlled study of 98 Japanese learners studying 16 English–Japanese word pairs, five and seven within-session retrievals produced higher raw test scores than one and three retrievals, while one retrieval produced the greatest gain after controlling for time on task. That result does not prescribe one retrieval per word; it demonstrates a trade-off in that study between item-level test performance and time efficiency. (Cambridge University Press)
The hidden denominator in every daily quota is therefore future review time. Adding an item today may create retrievals tomorrow and later. The exact burden depends on your memory, target, scheduler, accuracy, prior knowledge, and the item itself, so there is no defensible universal “reviews per new word” multiplier.
Use the Two-Gate Quota
The Two-Gate Quota is an editorial workload model, not a research-validated formula. Its starting value, gates, calibration period, and adjustment thresholds have not been tested in a controlled trial. They are practical heuristics inferred from evidence about variable outcomes, incomplete vocabulary gains, item differences, spacing, and time-on-task trade-offs.
The model asks every potential new item to pass through two gates.
Gate 1: useful reading supply
Add only items that appear in material you genuinely want or need to read, watch, or listen to. A candidate might help you understand the current material, recur in your subject area, support a real communicative goal, or simply be memorable enough to deserve attention. These are selection heuristics, not experimentally validated criteria.
Meaning-focused input produces real but incomplete vocabulary gains. A meta-analysis found that the average proportion of target words learned ranged from 9% to 18% on first posttests and from 6% to 17% on follow-up tests. The estimates are constrained by study designs, test formats, target-word selection, and a relatively small delayed-test evidence base. (Cambridge University Press)
The practical interpretation is not that reading makes deliberate study unnecessary. Meaning-focused input and intentional study are complementary, and neither research tradition suggests that one approach alone will quickly create comprehensive word knowledge. This is a synthesis across research traditions, not a direct randomized comparison of every possible combined program. (Cambridge University Press, Wiley) Reading supplies context and relevant candidates; deliberate retrieval gives selected items focused attention.
Do not fill an unused quota with arbitrary list words merely to preserve a streak. If today’s reading produces only two worthwhile candidates, learn two.
Gate 2: review headroom
Set a fixed daily vocabulary-study budget and protect a separate minimum for real language use. For example, you might allow 15 minutes for vocabulary review while preserving at least 20 minutes for reading. These numbers are personal choices, not research findings.
New items pass Gate 2 only when due reviews fit comfortably inside the vocabulary budget without consuming the protected reading floor. If reviews repeatedly carry over or crowd out reading, stop or reduce additions.
Spacing research requires careful attribution. A controlled primary experiment found a limited advantage for expanding over equal spacing under its particular conditions. A later meta-analysis across second-language studies found equal and expanding schedules statistically equivalent overall. The synthesis also found an overall advantage for spaced over massed practice, particularly on delayed tests, while cautioning that results depend on the retention interval, task, feedback, and other conditions. Neither source supports one universal schedule. (Cambridge University Press, Wiley)
For the quota decision, the narrower point is that practice may continue after the day an item is added, so current intake should leave time for later reviews.
A 14-day calibration protocol
Use the following feedback loop to set a manageable pace. The starting value and adjustment thresholds are transparent editorial heuristics, not validated learning laws.
- Define one learning item. Use one form or phrase, one relevant meaning, and one context.
- Choose a vocabulary budget. Decide how many minutes reviews may consume each day.
- Protect a reading floor. Reserve a minimum amount of time for actual reading or other meaningful input.
- Start with five new items per day. Continue for 14 days, but add fewer when reading supplies fewer useful candidates.
- Record three signals: review minutes, unfinished due reviews, and useful candidates encountered while reading.
- Adjust after each week:
- Add one item per day if reviews stayed within budget on at least six of seven days, no reviews carried over, and reading supplied enough candidates.
- Subtract two items per day if reviews carried over on two or more days or displaced your reading floor.
- Otherwise, hold the quota steady.
This makes the quota a ceiling rather than an obligation. If your current setting is six but only three useful items appear, add three. If reviews already exceed the budget, add none.
Weekly checklist
- [ ] Am I counting items consistently?
- [ ] Did most candidates come from material I genuinely use?
- [ ] Did reviews fit the vocabulary budget on at least six days?
- [ ] Did any due reviews carry over?
- [ ] Did vocabulary work reduce my protected reading time?
- [ ] Am I checking performance later rather than relying on immediate familiarity?
- [ ] Does my target require recognition only, or accurate production too?
- [ ] Should next week’s ceiling rise by one, stay fixed, or fall by two?
Three worked examples
These examples illustrate the editorial model; their numbers are not research findings.
A busy learner with stable reviews: Maya protects 20 minutes of reading and allows 10 minutes for vocabulary review. She begins with five new items a day. After one week, her reviews stayed under ten minutes on six days, none carried over, and her reading produced more than five useful candidates most days. Under the proposed rule, she raises her ceiling to six.
A learner with a growing backlog: Tomas starts at five, but reviews exceed his 15-minute budget on three days and replace part of his reading twice. He lowers the next week’s quota to three. If the backlog remains, he can temporarily set the quota to zero. That is workload control, not failure.
A learner with limited useful supply: Aisha’s review queue is light, but her specialist reading produces only one or two unfamiliar terms worth keeping each day. She does not supplement them with random vocabulary to reach five. Gate 1 sets her actual intake below the nominal ceiling.
These cases also show why a monthly promise such as “learn 600 words” does not fit this adaptive model. It cannot guarantee a fixed total because it responds to observed review burden, useful input, available time, and differences among items.
Common mistakes to avoid
- Counting every inflection as a new word. Decide whether related forms belong to one item before tracking totals.
- Adding words to satisfy a streak. A quota is a capacity limit, not a command to collect filler.
- Treating recognition as mastery. One correct answer immediately after study does not establish later recall or independent use.
- Ignoring the target skill. Exact-form recall, pronunciation, collocation, and production involve different or broader assessment demands than recognising a meaning in context. (Cambridge University Press)
- Optimising repetitions without considering time. More practice produced higher raw scores in one controlled study, while fewer retrievals were more efficient after controlling for time on task. That result is a trade-off, not a universal repetition rule. (Cambridge University Press)
- Letting review replace language use. Research supports treating meaning-focused input and deliberate study as complementary, so preserve time for both.
Within this unvalidated workload model, the most useful daily setting is not necessarily the largest number you can add today. It is a ceiling that continues to fit both gates next week: enough worthwhile candidates from real input and enough headroom for later reviews without sacrificing reading.
Stickly callout: Stickly can support this approach by letting you deliberately translate selected text while browsing, save useful words with context, and review words that are due. Treat your daily additions as an adjustable ceiling rather than a required streak.
Sources
- How Effective Are Intentional Vocabulary-Learning Activities? A Meta-Analysis
- How Effective Is Second Language Incidental Vocabulary Learning? A Meta-Analysis
- The Effects of Spaced Practice on Second Language Learning: A Meta-Analysis
- Effects of Expanding and Equal Spacing on Second Language Vocabulary Learning
- Does Repeated Practice Make Perfect?
- Predicting Vocabulary Knowledge in Adult L2 Learners
- Understanding L2-Derived Words in Context
AI-assisted research and automated checks by Stickly Editorial
