All articles

Learning science

How comprehensible input builds vocabulary

Learn what comprehensible input can and cannot do for vocabulary, how lexical coverage affects reading, and how to use a gist–gap–gloss–re-entry loop with authentic material.

Stickly Editorial11 min read
A language learner reading an authentic webpage and connecting an unfamiliar word back to its surrounding sentence

Comprehensible input builds vocabulary when you can follow a message while still meeting language that is not fully known. Familiar words, syntax, topic knowledge, and discourse give you clues; selected unfamiliar words then become learnable rather than mere noise.

That does not mean every understandable page teaches every unknown word. Meaning-focused input usually produces positive but incomplete vocabulary gains, and immediate comprehension can coexist with weak later recall.A meta-analysis of incidental second-language vocabulary learning found gains across several input modes, but only a subset of target words was typically learned.

For intermediate learners browsing authentic material, the practical aim is therefore not to eliminate uncertainty. It is to preserve the gist, repair a few consequential gaps, and return each repaired word to the message that made it meaningful.

What “comprehensible input” actually means

Krashen’s influential input hypothesis proposes that acquisition occurs through comprehended language containing features somewhat beyond a learner’s current level, commonly represented as i+1.Krashen presents this formulation in Principles and Practice in Second Language Acquisition. It is best understood as a theoretical proposal, not a numerical formula that can calculate the perfect webpage for an individual learner.

The crucial word is comprehended. A podcast, article, or social post may be available to your eyes or ears without its message becoming understandable. Conversely, you may understand a passage despite several unknown words because the title, surrounding sentences, images, and your knowledge of the topic constrain their possible meanings.

Comprehensibility also cannot be reduced to vocabulary alone. Syntax, genre, discourse structure, background knowledge, reading purpose, and the standard used for “adequate” comprehension all affect whether a text works for a learner.Replication and processing research indicates that the relationship between lexical coverage and comprehension varies with tasks, texts, and measures rather than producing one universal cutoff.

This creates two necessary conditions for vocabulary growth:

  1. The message must remain sufficiently understandable. Known language supplies a framework in which an unfamiliar form can be interpreted.
  2. Some attention must reach the unfamiliar item. A learner can understand the overall message while overlooking a word, guessing it incorrectly, or forgetting it immediately.

Input creates an opportunity. It does not guarantee what the learner will notice, retain, or later use.

Lexical coverage is a diagnostic, not a law

Lexical coverage is the proportion of running words in a text that a reader knows. In an influential study, Hu and Nation found that comprehension generally improved as coverage increased. None of the readers in the 80% condition met the study’s criterion for adequate comprehension, while results varied substantially at 90% and 95%.The study involved 66 pre-university learners reading manipulated versions of one fiction text.

Hu and Nation inferred through regression that 98% coverage was necessary for most learners to achieve adequate comprehension; 98% was not directly tested as an experimental condition in their study.The conclusion came from one manipulated fiction text, a particular learner population, and the study’s chosen comprehension criterion.

A later registered replication directly included a 98% condition, used both narrative and informational material, and tested 104 Sri Lankan adult learners. It partially supported a mostly linear coverage–comprehension relationship but did not reproduce one threshold that generalized across genres and test formats.The replication authors therefore called for a more nuanced treatment rather than treating a single percentage as universally sufficient.

The arithmetic still helps illustrate reading load. At 98% coverage, a 1,000-token text contains roughly 20 unknown tokens—about one in every 50. At 95%, it contains about 50.These are running-word calculations, not personal vocabulary estimates or instructions to count every word. Repetitions, names, transparent word families, partially known words, and strategically important terms make real pages more complicated.

Use coverage as a description of friction, not as a pass–fail score. A familiar news topic may remain easy with several unknown words. A legal explanation may collapse because one unknown connector reverses the argument. Your reading purpose matters too: browsing for the main point tolerates more uncertainty than following medical instructions.

A useful decision rule is:

Reading state What to do
You can state the paragraph’s gist and unknown words do not change it Continue without lookup
The gist survives, but a few words control the argument, tone, or action Repair those words selectively
Unknown language repeatedly prevents paragraph-level understanding Switch to an easier or parallel source

This is an evidence-informed heuristic rather than a validated assessment scale. It reflects the continuous, task-sensitive relationship between coverage and comprehension found in coverage research. Hu and Nation concluded that 98% was necessary for most learners to meet their adequate-comprehension criterion, whereas the later replication did not reproduce a generally applicable 98% threshold and cautioned against treating one percentage as sufficient across genres and test formats.The replication discusses both the original conclusion and the limits of a universal threshold.

What reading can—and cannot—teach

Meaning-focused exposure can build vocabulary incidentally. A 2023 meta-analysis covering 24 primary studies and 2,771 participants found positive learning through reading, listening, reading while listening, and viewing. Across those modes, mean proportions of target words learned ranged from 9–18% on immediate tests and 6–17% on delayed tests.These pooled estimates varied with learner, material, mode, test, spacing, and study design; they are not promised returns for one browsing session.

Authentic material is not automatically better. The same meta-analysis reported larger effects for material designed for second-language learners than for native-user material, although this was a study-level association rather than a randomized comparison of otherwise identical texts.Native-user and learner-oriented materials can differ in many ways besides lexical difficulty. Authentic pages are useful when interest and relevance keep you engaged and enough of the message remains available for learning.

One encounter can also produce only partial knowledge. You might recognize a word’s form without recalling its meaning, recognize the meaning in context without producing the word, or know one sense without knowing its pronunciation, register, grammar, or common collocations.

In a controlled reading experiment, learners encountered target items eight times and later recognized more forms and meanings than they could recall without support.The result illustrates that recognition and unaided recall are distinct outcomes; eight encounters should not be treated as a universal prescription.

Repeated encounters can strengthen knowledge incrementally, but exposure count is not the whole story. In a more naturalistic eye-tracking study based on chapters of a novel, the number of exposures was the strongest predictor of form and meaning learning, while total reading time independently predicted meaning learning.These are predictive relationships, not proof that accumulating a fixed number of sightings or staring longer will cause mastery.

A separate eye-tracking experiment with 87 advanced readers found very low average vocabulary uptake—2.26%—despite controlled conditions. Total and second-pass reading time predicted meaning recall, but the authors cautioned against strong conclusions because so few words were learned.Longer processing may reflect inference, difficulty, rereading, or several processes at once; it should not be interpreted as a simple causal technique.

The lesson is modest: comprehension gives vocabulary somewhere to attach. Attention, informative context, later encounters, and assessments of recognition, recall, or production indicate different aspects of what the learner knows; the cited studies do not establish retrieval testing as a cause of durable vocabulary retention.

The gist–gap–gloss–re-entry loop

The following loop is an editorial synthesis for authentic browsing, not a protocol directly tested as one intervention. It combines meaning-first reading, selective support, contextual rereading, and a later check of the kind of knowledge you actually need.

1. Gist: read before interrupting yourself

Read one paragraph or short screenful without opening a dictionary. Then state its point in one sentence—mentally or in writing.

For example:

The author argues that the city’s new transport plan may reduce traffic, but funding is uncertain.

If you can produce a stable summary, the input is doing its main job: carrying a message. If you cannot, reread once. When the paragraph still has no coherent point, you may need selective repair or an easier parallel source.

2. Gap: identify the smallest consequential problem

Do not treat every unfamiliar item as equally important. Classify it:

  • Blocker: It prevents the main claim, action, referent, or relationship from making sense.
  • Recurring topic word: It appears repeatedly and is likely to matter across this page or related pages.
  • Optional detail: The gist survives without it.

Intervene for blockers and useful recurring terms. Let optional details pass unless your purpose requires precision.

Suppose you read, “The proposal was shelved after the committee failed to secure funding.” If shelved is unknown, it controls what happened to the proposal. It is a blocker. An unknown adjective describing the committee room probably is not.

3. Gloss: infer first, then verify briefly

Before looking up the word, make a small hypothesis about its role and meaning:

  • Is it an action, entity, quality, or connector?
  • Does the sentence suggest continuation, rejection, contrast, or cause?
  • What meaning would preserve the paragraph’s logic?

Then check a concise definition, translation, or gloss. Glossed reading has produced greater vocabulary learning than nonglossed reading in a meta-analysis of 42 studies, although outcomes varied with gloss format, proficiency, tests, and study design.The evidence supports word information as a useful aid, not translating every unfamiliar item or assuming all lookup tools have identical effects.

Treat the lookup as correction of a hypothesis, not as the end of the task. For shelved, you might predict “delayed or stopped,” then verify the relevant sense.

4. Re-entry: put the word back into the message

Immediately reread the original sentence and its surrounding paragraph. Ask:

  • Does the verified meaning change my gist?
  • Does it alter the tone or strength of the claim?
  • Does it resolve a pronoun, cause, contrast, or timeline?
  • Can I now read the sentence without mentally replacing the word with a dictionary entry?

This re-entry step prevents the lookup from becoming an isolated vocabulary event. The word returns to the discourse that constrained its meaning. In the example, you now understand not merely that shelved can mean “postponed,” but that the funding failure caused the proposal to stop moving forward.

5. Transfer: assess the knowledge you need later

On another page or in a later review, choose a check that matches your goal:

  • Receptive goal: Notice shelved and check whether you can recognize its relevant meaning before revealing help.
  • Productive goal: Check whether you can paraphrase the sentence or write one short example using the relevant sense.

These checks assess different dimensions of word knowledge rather than proving that the checking process causes durable retention. Production may expose gaps that comprehension allowed you to bypass, which is why a brief paraphrase can be informative when active use is the goal.Swain and Lapkin argue that producing language can prompt learners to notice gaps under some conditions; they do not establish production or retrieval testing as a guarantee of lasting vocabulary learning.

A practical browsing checklist

Before committing to a page, ask:

  • Can I explain what the last paragraph was about?
  • Are unknown words occasional gaps, or is the entire message unstable?
  • Which one or two words materially control the meaning?
  • Can I infer each selected word’s role before checking it?
  • After checking, did I reread the original sentence and paragraph?
  • Am I assessing recognition, recall, or productive use?
  • Is this word recurring enough to deserve later review?
  • Would an easier article on the same topic give me a better foundation?

The best input is not necessarily the hardest material you can endure. It is material that sustains meaning while leaving manageable uncertainty for attention and repair. Comprehensible input creates vocabulary opportunities. Later recognition, recall, and production checks can show which aspects of selected words are available to you, while the cited evidence does not support promising that those checks will cause durable retention.

A restrained Stickly option: If you want tool support for this process, Stickly can translate deliberately selected text, save useful words with their context, show remembered words on later pages, and provide due-word review. These features can support selective repair and later checks of word knowledge, but they do not determine page difficulty or guarantee learning.

Sources

AI-assisted research and automated checks by Stickly Editorial

Stickly Editorial uses AI-assisted research and writing tools. Every published article passes automated source, product-accuracy, and quality checks.