Why does information density matter to AI models?
A model answering a user has limited room for a quotation. It picks the passage that packs the most checkable substance per sentence - a number, a date, a unit, a proper name. A paragraph built from “many companies”, “modern solutions” and “broad experience” cannot be quoted, because it contains no claim anyone could confirm or refute.
This reverses an old SEO instinct. Volume used to help - longer text ranked better, so stretching paragraphs paid off. In GEO the opposite ratio counts: diluting facts lowers the odds of being cited, even when the facts are there. The same material written half as long performs better.
What will you see in the report?
You get the share of fact-bearing sentences with a score, plus two lists pulled from your own text: sentences counted as facts and sentences counted as filler. Both lists are built from your own sentences, so you can see exactly which paragraphs drag the result down.
Recommendations come in a Before / After format with a concrete rewrite proposed.
Sample recommendations
Fragments of a report from an audit of an electronics store category page.
Problem: The paragraph opens with a rhetorical question — there is no claim in it that could be confirmed or refuted.
Want to know more? Use our recommendations and browse products in specific categories.
Replace the question with a fact: “Buying guides and related categories — 12 topics”.
Problem: The sentence is built entirely from generalities: “varied”, “perfectly”, “maximum precision” — no verifiable data.
Thanks to varied sensors, shapes and functions you will match the mouse perfectly to your needs and reach maximum precision in every game.
Give the criterion: “A 26,000 DPI optical sensor and a claw-grip profile cut cursor tracking errors in FPS games”.
How do we measure information density?
The basis is algorithmic: we split the text into sentences and use patterns to detect which ones carry a fact - a number, date, unit, amount, proper name or a concrete claim. The result is the share of such sentences, rescaled to a ten-point score.
The formula: sentences carrying a fact divided by all sentences, times 10. The algorithm alone computes this score, and the language model supplies the examples: actual sentences from your text on both sides of the line.
Signals that adjust the result
| Signal in a sentence | Effect |
|---|---|
| A concrete number, amount or date | up |
| A value attached to a property (“21.3% efficiency”) | up |
| A single, checkable claim | up |
| A proper name, brand, standard or institution | slightly up |
| Filler phrase (“it is worth knowing”, “in today’s world”) | strongly down |
| Hedging word (“may”, “usually”, “rather”) | down |
| Unsupported evaluative adjective (“best”, “effective”) | down |
| Rhetorical question | down |
What raises and what lowers the score?
Raises
- Numbers, amounts, dates and units instead of approximations
- One claim per sentence instead of a sentence built from three caveats
- Proper names: brands, standards, institutions, models
- Values tied to specific properties rather than general praise of the product
- Shortening the text without removing facts - changing the ratio alone raises the score
Lowers
- Introductory paragraphs without a single checkable statement
- Filler phrases: “it is worth knowing”, “in today’s world”, “as we all know”
- Hedging: “may”, “usually”, “rather”, “in a sense”
- Evaluative adjectives with no data behind them
- Rhetorical questions instead of answers
The fastest fix is trading generalities for data - same sentence, different density
| Instead of | Write |
|---|---|
| “many people” | “67% of respondents” |
| “expensive to run” | “from $300 a month” |
| “modern technology” | the name of the technology and the year it shipped |
| “significantly more efficient” | “18% more efficient than model X” |
When hard data is missing, ranges, proportions and comparisons still work - “3-5 days”, “1 in 4 tickets”, “twice the average”.
Frequently asked questions
Is shorter text always better?
No - the ratio counts, not the length. A long text with a high share of facts beats a short one built from generalities. Cutting helps only when you remove empty sentences rather than the facts themselves.
What if I do not have hard data?
Use ranges, proportions and comparisons. “3-5 business days”, “1 in 4 tickets”, “twice the average” all count as concrete, because they can be verified. Only statements nobody can check are worthless.
Does the score depend on what the model decides?
No. The share of fact-bearing sentences is computed by an algorithm over patterns - number, date, unit, amount, proper name. The language model only adds examples from your text so you can see which paragraphs are meant.