Root CauseWhat broke, why, and the fix.

71.2% of top YouTube comments contain a word that is not in the textbook

· english learning, measurement, vocabulary
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

What we hit

A learner can pass a vocabulary test and still struggle to read the comment section under a popular video. Familiar words sit beside abbreviations, informal spellings, and expressions that require context. We wanted a number for that gap instead of an impression.

We narrowed the question to something we could measure: how often does a comment contain a word outside the vocabulary lists we use for learning? That does not tell us whether a learner understands the comment. It tells us how often reading takes them beyond those lists.

The distinction matters. We were measuring a boundary in our learning material, not testing readers or grading their English.

What we assumed

We assumed that standard vocabulary lists covered everyday online English well enough. Our baseline combined NGSL with the TOEIC, TOEFL, and Eiken exam vocabulary bundled with Words Admin.

That assumption was convenient because list membership is easy to check. A word either matches an entry or does not. We could organize study around that boundary and treat progress through the lists as progress toward reading ordinary English.

But a useful study list is not automatically a description of language in use. We needed to compare our baseline with comments people actually encounter.

What we measured

We used the YouTube Data API to retrieve the most-popular video charts for US, GB, CA, and AU. We removed duplicate videos, excluded News & Politics and Nonprofits categories, and applied filters for predominantly ASCII titles and a minimum available comment count. That left 82 videos.

We fetched top-level comments in relevance order, collecting 2,631 comments. We did not sample replies or put comments in chronological order. The resulting dataset therefore represented the comments surfaced near the top of those videos.

We then applied an English filter. It excluded comments containing CJK, Cyrillic, Arabic, or Thai characters, rejected URLs, and imposed minimum English-letter and English-word requirements. A vocabulary-ratio gate also required a sufficient share of tokens to appear in the combined dictionary, NGSL, and exam vocabulary. That gate removed 191 non-English comments, including Spanish-language material.

After filtering, 2,408 comments remained. We used the first 1,000. This was an ordered sample from the filtered collection, not a random sample of everything posted on YouTube.

We compared tokens against these resources:

  • NGSL: 2,800 words, with inflected forms providing 8,481 forms.
  • Exam vocabulary: 3,191 entries covering TOEIC, TOEFL, and Eiken.
  • Dictionary: an older system word list containing 234,456 entries.

Tokenisation retained English letters and apostrophes, then converted tokens to lowercase. Before matching, a proper-noun heuristic excluded capitalized words outside sentence beginnings and words appearing in video titles or channel names.

We also normalized contractions and stripped common endings to look for matching stems. These included plural, past-tense, progressive, adverbial, comparative, superlative, and noun-forming endings. Elongated spellings were excluded from evaluation.

We defined metric A as the share of comments containing at least a word outside the union of NGSL and exam vocabulary. Metric B applied the additional dictionary check: a comment qualified when it contained a word outside those learning lists and outside the dictionary.

What the counts said

Metric A was 712/1,000 = 71.2%.

Metric B was 312/1,000 = 31.2%.

Both metrics count comments containing an unmatched word. They do not measure the proportion of all tokens that were unmatched, and they do not mean that the entire comment was unfamiliar. A mostly ordinary sentence can cross the boundary because of an abbreviation.

Among the leading out-of-dictionary tokens, excluding apparent names from this short list, were:

  • bro: 21 comments
  • lol: 13 comments
  • kinda: 9 comments
  • hype: 8 comments
  • gameplay: 8 comments

These are isolated token counts, not reproduced comments. We used comment text for the measurement; we do not reproduce or paraphrase it here.

The tokens also show why “outside the dictionary” needs care. A missing entry can reflect the dictionary’s coverage rather than an unusual expression.

Limits

Metric B is overstated. The dictionary is old and omits ordinary words such as gameplay, internet, and wheelchair. Some game-character and streamer names also slipped past the proper-noun heuristic. We therefore treat B as a ceiling, not a clean estimate of slang or unfamiliar vocabulary.

Our “textbook” definition is operational: NGSL plus the exam vocabulary. It does not represent every textbook. Some proper nouns and derivations that appear in textbooks can still be counted outside our lists, pushing A upward.

The opposite error also exists. Familiar textbook words can acquire slang meanings in combination, while token-level matching still counts their component words as covered. We measured membership, not meaning, so that kind of gap can remain invisible.

Relevance ordering introduces another limit. Highly liked top comments skew toward short, casual language. We cannot extend the result to all comments, long-form writing, or everyday English generally.

A, at 71.2%, is the number we stand behind for this sample and definition. B is a ceiling. Neither metric directly measures comprehension.

What we changed

We stopped treating the word lists as evidence of coverage. We built a feature that takes real comments and explains the words outside those lists, connecting vocabulary study to the language that prompted the question.

The lesson

We had treated a learning resource as a coverage model without checking the boundary against actual use. Measuring comments made that boundary visible, even with imperfect matching. The useful result is a scoped claim about this sample, supported by explicit rules and limits. We now use the lists as a starting point and real comments to reveal what still needs explanation.

Root Cause is written by TechAthletes. This measurement is why we built Native Comments inside Words Admin: real comments, with the words outside the textbook explained.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy