What is TF-IDF in a content audit?

TF-IDF compares your page terminology with the vocabulary of the Top 10. A missing specialist term usually means a missing angle, so CitationOne shows which concepts you skip and in what context competitors use them.

Context

Why does TF-IDF matter to AI models?

Terminology is the cheapest competence signal to verify. A text about mortgages that never mentions “loan-to-value ratio”, “creditworthiness” or “bank margin” covers the topic superficially, and the model compares it with ten pages that do use those concepts.

In GEO this carries extra weight, because a missing term usually means a missing angle. If nine competitors write about something absent from your page, you have a substantive gap the model will see when assembling its answer.

In practice

What will you see in the report?

You get a list of missing terms ordered by how many competitors use them and how strongly they bind to the page topic.

Where a term consistently appears alongside another concept among competitors, we surface that context - the recommendation then reads “add term X in the context of Y, 6 of 10 competitors connect them”.

Sample recommendations

Fragments of a report from an audit of a gaming mice category page in an electronics store.

Problem: The phrase “optical dpi” appears at 5 of 8 competitors when describing the sensor, and never once in your content.

Before

The optical sensor affects tracking accuracy and response speed.

After

Use the term in an explanation: “an optical dpi resolution of 26,000 translates into tracking accuracy without interpolation”.

Problem: The advisory phrase “best gaming mouse” appears at 4 of 8 competitors, usually next to selection criteria - you do not use it at all.

Before

Thanks to varied sensors and shapes you will match the mouse to your needs.

After

Add an advisory section: “Best gaming mouse for FPS - the parameters that matter”.

Method

How do we measure TF-IDF?

We build a reference set from competitor content and identify the terms with high informational value: specialist, industry-specific, often multi-word. Function words and generic nouns are discarded. The score is the ratio of terms present in your content to those expected, rescaled to ten.

The formula: specialist terms present in the content divided by expected terms, times 10.

Which terms count

Type of termWeight for the score
Specialist, industry-specific, multi-wordHigh - this is what proves subject familiarity
Generic noun, generic adjectiveIrrelevant - appears in every text
Function wordIgnored
Factors

What raises and what lowers the score?

Raises

  • Industry concepts used inside an explanation, not merely listed
  • Multi-word phrases the industry actually uses, rather than colloquial substitutes
  • Terms present at most of the Top 10 competitors
  • A term used in the context competitors pair it with - “cortisol” next to “elevated”, not in isolation

Lowers

  • Describing the topic in general language, skipping the names of the phenomena involved
  • Missing concepts that are standard among competitors
  • Terms dumped as a list at the end of the text, with no explanation
  • Replacing a term with a description (“that indicator which shows...”) instead of naming it
Questions

Frequently asked questions

Is this not keyword stuffing?

No - we measure concept coverage, not phrase frequency. A term counts when it appears inside an explanation; dumped at the end of the text it builds neither understanding nor score.

How many terms do I need to add?

As many real gaps as the report identifies, starting from the top. Terms are ordered by the number of competitors using them and the strength of their link to the page topic, so the first entries deliver the largest gain.

Where does the list of missing terms come from?

From the content of the Top 10 competitors for your phrase.

Related

Related dimensions