What is information density in content?

How many facts, data points and concrete statements a single paragraph contains. Low-density content is filler - AI skips it. This dimension rewards every sentence that adds something new and penalizes text that repeats the same idea in different words.

Context

Why does information density matter to AI models?

A model answering a user has limited room for a quotation. It picks the passage that packs the most checkable substance per sentence - a number, a date, a unit, a proper name. A paragraph built from “many companies”, “modern solutions” and “broad experience” cannot be quoted, because it contains no claim anyone could confirm or refute.

This reverses an old SEO instinct. Volume used to help - longer text ranked better, so stretching paragraphs paid off. In GEO the opposite ratio counts: diluting facts lowers the odds of being cited, even when the facts are there. The same material written half as long performs better.

In practice

What will you see in the report?

You get the share of fact-bearing sentences with a score, plus two lists pulled from your own text: sentences counted as facts and sentences counted as filler. Both lists are built from your own sentences, so you can see exactly which paragraphs drag the result down.

Recommendations come in a Before / After format with a concrete rewrite proposed.

Sample recommendations

Fragments of a report from an audit of an electronics store category page.

Problem: The paragraph opens with a rhetorical question — there is no claim in it that could be confirmed or refuted.

Before

Want to know more? Use our recommendations and browse products in specific categories.

After

Replace the question with a fact: “Buying guides and related categories — 12 topics”.

Problem: The sentence is built entirely from generalities: “varied”, “perfectly”, “maximum precision” — no verifiable data.

Before

Thanks to varied sensors, shapes and functions you will match the mouse perfectly to your needs and reach maximum precision in every game.

After

Give the criterion: “A 26,000 DPI optical sensor and a claw-grip profile cut cursor tracking errors in FPS games”.

Method

How do we measure information density?

The basis is algorithmic: we split the text into sentences and use patterns to detect which ones carry a fact - a number, date, unit, amount, proper name or a concrete claim. The result is the share of such sentences, rescaled to a ten-point score.

The formula: sentences carrying a fact divided by all sentences, times 10. The algorithm alone computes this score, and the language model supplies the examples: actual sentences from your text on both sides of the line.

Signals that adjust the result

Signal in a sentenceEffect
A concrete number, amount or dateup
A value attached to a property (“21.3% efficiency”)up
A single, checkable claimup
A proper name, brand, standard or institutionslightly up
Filler phrase (“it is worth knowing”, “in today’s world”)strongly down
Hedging word (“may”, “usually”, “rather”)down
Unsupported evaluative adjective (“best”, “effective”)down
Rhetorical questiondown
Factors

What raises and what lowers the score?

Raises

  • Numbers, amounts, dates and units instead of approximations
  • One claim per sentence instead of a sentence built from three caveats
  • Proper names: brands, standards, institutions, models
  • Values tied to specific properties rather than general praise of the product
  • Shortening the text without removing facts - changing the ratio alone raises the score

Lowers

  • Introductory paragraphs without a single checkable statement
  • Filler phrases: “it is worth knowing”, “in today’s world”, “as we all know”
  • Hedging: “may”, “usually”, “rather”, “in a sense”
  • Evaluative adjectives with no data behind them
  • Rhetorical questions instead of answers

The fastest fix is trading generalities for data - same sentence, different density

Instead ofWrite
“many people”“67% of respondents”
“expensive to run”“from $300 a month”
“modern technology”the name of the technology and the year it shipped
“significantly more efficient”“18% more efficient than model X”

When hard data is missing, ranges, proportions and comparisons still work - “3-5 days”, “1 in 4 tickets”, “twice the average”.

Questions

Frequently asked questions

Is shorter text always better?

No - the ratio counts, not the length. A long text with a high share of facts beats a short one built from generalities. Cutting helps only when you remove empty sentences rather than the facts themselves.

What if I do not have hard data?

Use ranges, proportions and comparisons. “3-5 business days”, “1 in 4 tickets”, “twice the average” all count as concrete, because they can be verified. Only statements nobody can check are worthless.

Does the score depend on what the model decides?

No. The share of fact-bearing sentences is computed by an algorithm over patterns - number, date, unit, amount, proper name. The language model only adds examples from your text so you can see which paragraphs are meant.

Related

Related dimensions