Stopword

Frequently occurring words excluded from search and text analysis, such as "a," "the," "is," and "in."

Stop words are high-frequency words excluded from text analysis and search engine indexing. In Japanese, particles like の, は, が, を, に, で, and と qualify; in English, articles, prepositions, and be-verbs like "a," "the," "is," "in," "and," and "of." While these words appear extremely frequently, they carry little semantic information on their own, making them noise in text analysis.

The main purposes of removing stop words are improving search accuracy and reducing index size. When a full-text search engine drops high-frequency words from its index, the number of postings the inverted index has to hold falls considerably. In measurements on the Reuters-RCV1 corpus reported in Introduction to Information Retrieval, a leading textbook on information retrieval, removing the 150 highest-frequency words cut the number of nonpositional postings by roughly 30%. In text mining, calculating TF-IDF (Term Frequency-Inverse Document Frequency) after removing stop words enables more accurate extraction of keywords that characterize documents.

Stop word lists vary by language and use case. Python's natural language processing libraries ship with a standard English stop word list, but the number of entries varies by library and version, so in an environment without pinned dependencies the list that is actually loaded should be checked before use. Japanese stop word lists are built based on morphological analysis results, centering on particles, auxiliary verbs, and conjunctions. Adding domain-specific stop words (e.g., "patient" in medical contexts, "clause" in legal contexts) can further improve analysis accuracy.

However, blanket stop word removal requires caution. In cases like "to be or not to be" where stop words carry core meaning, or "The Who" (band name) where proper nouns contain stop words, removal causes information loss. Phrase searches ("New York," etc.) also depend on stop word positional information, making complete removal inappropriate.

As the same textbook notes, web search engines have largely done away with stop lists: compression lowers the cost of keeping postings for frequent words, and idf weighting leaves extremely frequent words with almost no effect on ranking, so there is little to gain from discarding them. Search engines and large language models judge the intent of a query from the whole sentence, function words included. Transformer models like BERT learn context from entire sentences including stop words, so preprocessing stop word removal can actually be counterproductive.

In relation to character counting, stop words characteristically account for a large proportion of total text character count. In English, the same corpus measurements confirm that the 30 highest-frequency words alone account for roughly 30% of all tokens, and Japanese particles likewise take up a substantial share of the character count. For character-limited content (tweets, meta descriptions, etc.), consciously reducing stop words allows packing more information into limited character counts.

Share this article