We opened our keyword density analyzer to check where its numbers come from, and the formula hides two decisions that change every result you see. The denominator counts every whitespace separated token in your text. The numerator counts matches after lowercasing, after stripping punctuation, and after dropping single character words. Those two counts cover different sets of words, so the percentage is an estimate by construction. This article explains the formula, where it stops meaning anything, and how to read the three views the tool gives you.
The formula behind the number
The analyzer computes density as the phrase count divided by the total word count, multiplied by 100, displayed with two decimals. A phrase appearing 12 times in a 1,000 word text shows 1.20 percent.
That sentence sounds precise and is not. Both inputs pass through filters first, and the filters are not symmetric.
The total word count splits your raw text on whitespace, so it includes single letters, it includes numbers, and it includes anything glued together with punctuation.
The keyword side is stricter. Text is lowercased, punctuation is replaced with spaces, and only tokens longer than one character survive. The word count and the keyword count therefore run on different views of your text, and the gap grows with messy input.
Why the stop word toggle changes the math
The stop word filter defaults to on. With it on, the English list holds 132 entries, and any phrase containing one of them is excluded from the table. The Turkish list holds 42 entries. The tool picks the list that matches the interface language, so analyze Turkish text with the Turkish interface or the filter applies the wrong list.
The filter removes phrases, and this matters for one specific reason. Two word and three word phrases in natural writing are full of stop words. With the filter on, the 2-gram view shows only phrases built entirely from content words. With it off, you see phrases like "of the" climb to the top and push useful rows out of the top 20.
Use the toggle as a lens, in this order.
1. Read the 1-gram view with stop words on to see your content words.
2. Read the 2-gram view with stop words on to find repeated term pairs.
3. Turn the filter off only to check how much filler dominates, then turn it back on.
Limits where density stops meaning anything
Our own project guidelines set a 1 to 2 percent range for a primary keyword. That range is a writing target, and it makes a poor ranking target, for three reasons visible in the tool itself.
First, the number moves with the denominator. Add a 300 word FAQ section that never repeats the keyword and your density falls even though the keyword section is unchanged. Density rewards or punishes length changes that have nothing to do with the keyword.
Second, the top 20 cutoff hides the tail. In a long text, hundreds of distinct phrases appear twice. The table shows the top 20 by count, so the view saturates quickly and stops distinguishing between texts.
Third, the tokenizer only keeps Latin letters, digits, the Latin extended range, the Cyrillic range, and hyphens. Scripts outside those ranges, including Chinese, Japanese without Latin characters, Arabic, and Thai, are stripped before counting. If you analyze such text the table reports nothing, and the honest reading is that this analyzer is built for Latin and Cyrillic scripts.
One more mechanical note. Apostrophes are treated as separators, so a possessive form splits into two tokens. The phrase "reader's data" never matches a search for "reader data" in the table, because the tokens differ.
How to read the three views together
The tabs run 1-gram, 2-gram, and 3-gram. Each builds phrases from consecutive tokens after filtering, then ranks by count and keeps the top 20. Read them as one sequence.
1. The 1-gram view shows which single words dominate. Look for accidental repetition of a generic word, such as a product name in every sentence.
2. The 2-gram view shows fixed pairs. This is the view that catches unintended echo, where the same pair appears in four consecutive paragraphs.
3. The 3-gram view is the tightest filter. Anything that tops this view is a phrase you repeat verbatim, and verbatim repetition is what readers register as monotone.
The stats row above the tabs gives total words, unique words, sentences, paragraphs, and characters. The unique to total ratio is the fastest quality signal in the tool. A text where the top few words crowd out variety shows up there before you read a single row.
Steps to run a useful check
1. Paste the final draft, never an outline, into the analyzer.
2. Confirm the interface language matches the text language.
3. Keep stop words on and read the 1-gram view. Note the top three rows.
4. Switch to the 2-gram view and scan for a phrase above 5 total uses in a 1,000 word draft.
5. Switch to the 3-gram view and check that nothing exceeds 2 uses.
6. Vary one repetition per paragraph at most, then re run the check on the edited draft only.
Checklist before you trust the number
1. The text language and the interface language match.
2. The stop word filter is on for the reading pass.
3. No 3-gram phrase repeats more than twice.
4. The unique to total word ratio did not drop after your last edit.
5. You treat any density value as a repetition report, not as a ranking dial.
Run your next draft through the [keyword density analyzer](https://webrecast.com/en/keyword-density) and read the 3-gram view first. If your rankings moved after you changed repetition patterns in either direction, send us the before and after. Real cases beat the 1 to 2 percent folk rule.