Skip to main content

SEO for Journalists UK: Keyword Research, Schema & Core Web Vitals

Get your journalism found on Google. Practical guidance on keyword research, entity-based SEO, NewsArticle structured data, and Core Web Vitals for UK news publishers.

Last reviewed: Next review due:

What is SEO for journalism and why does it matter?

Search engine optimisation (SEO) for journalism is the practice of structuring, framing, and publishing news content so that it surfaces when people search for it on Google and other search engines. It is not about manipulating search results — it is about writing clearly, using the right words, and providing the technical signals that tell search engines what your content is about.

Commercial SEO focuses on ranking for competitive search terms over months. News SEO is different: it prioritises timeliness, entity recognition, and appearing in Google's Top Stories carousel and Discover feed. The keyword research tools are similar — Google Trends and Answer The Public UK are the most accessible — but the approach is tailored to the news cycle rather than evergreen content.

Entity-based SEO is particularly important for news. Google increasingly understands content in terms of entities — real-world people, organisations, and places — rather than keywords alone. Using the exact, consistent name of a person or organisation (rather than varying it with epithets or abbreviations) helps Google correctly identify what your story is about and connect it to related coverage.

When SEO matters most

Breaking news

Speed and clarity matter most. Your headline must name the subject, action, and place in plain language so Google can match it to search queries instantly.

Evergreen guides

These benefit most from keyword research. Use Google Trends to confirm search demand before commissioning. Update them regularly to maintain freshness signals.

Local journalism

Local search SEO rewards specific place names, postcode-level content, and Google Business Profile optimisation for the publication itself.

National news

NewsArticle schema and Core Web Vitals matter most here — Top Stories placement drives enormous traffic to national titles on breaking stories.

Investigations

Long-form investigations benefit from entity consistency and internal linking. A multi-part series should link between instalments and to relevant evergreen explainers.

Live blogs

Live blogs need LiveBlogPosting schema to surface the update stream in Google Search. Each update should have a clear timestamp and a distinct headline.

Red flags to watch for

  • Keyword stuffing in headlines — repeating a search term unnaturally damages readability and is penalised by Google.
  • Misleading headlines written for clicks rather than accuracy — a breach of IPSO Clause 1 that also creates a high bounce rate signal.
  • Thin content: articles under 300 words with little context or analysis. Google's helpful content systems actively demote thin news pages.
  • Missing or malformed NewsArticle structured data — the validator at schema.org will show errors that prevent rich result eligibility.
  • Slow Core Web Vitals caused by unoptimised images or blocking third-party scripts — directly impacts ranking.
  • Duplicate title tags across multiple articles on the same story — each URL needs a unique, descriptive title.
  • No canonical tag on syndicated content — if your article appears on partner sites, a canonical tag prevents duplicate content issues.
  • Ignoring robots.txt and sitemap — search engines need permission and a map to crawl your content efficiently.

SEO checklist for UK journalists

  • Headline includes the primary keyword or subject name in the first three words.
  • Meta description is 150–160 characters and summarises the story accurately.
  • NewsArticle schema includes datePublished, dateModified, author name, and publisher.
  • All images have descriptive alt text naming the subject and context.
  • At least two relevant internal links to related articles or evergreen guides.
  • Canonical URL is set, especially if content is syndicated elsewhere.
  • Article is included in the XML sitemap (or sitemap auto-updates via CMS).
  • robots.txt does not accidentally block Googlebot from the article URL.
  • Core Web Vitals pass: LCP under 2.5s, CLS under 0.1, INP under 200ms.
  • An llms.txt file is present at the domain root indicating AI crawler permissions.

Tool recommendations

Google Search Console (free)

Track impressions, clicks, CTR, and average position for your articles. Essential for understanding what is working and what is not.

Google Trends (free)

Compare search interest for different keywords and find rising topics. Use the UK filter to see local demand.

Answer The Public

Visualises questions and prepositions people search around a topic. Useful for evergreen guide structures and FAQ sections.

Screaming Frog (free up to 500 URLs)

Crawl your own site to find broken links, missing meta descriptions, duplicate titles, and missing canonical tags.

Schema Markup Validator

Validate your NewsArticle or LiveBlogPosting structured data at validator.schema.org before publishing.

Common mistakes

  • Writing headlines for print or social and forgetting that the HTML title tag is what Google indexes — they should often be different.
  • Using the same meta description template for every article — duplicate meta descriptions confuse search engines and reduce CTR.
  • Not updating dateModified when an article is substantially revised — Google uses this to assess freshness.
  • Uploading full-resolution images without compression — a 5MB hero image will cause LCP failures.
  • Blocking Googlebot via robots.txt accidentally during CMS migration.
  • Adding structured data that does not match the visible page content — Google will ignore or penalise mismatched schema.
  • Never checking Google Search Console — you cannot improve what you do not measure.
  • Treating SEO as a separate team's job rather than a skill integrated into the editorial workflow.

Related guides

Primary sources

Frequently asked questions

What is the most important SEO signal for UK news articles?
Relevance and timeliness are the two dominant signals for news content in Google Search. Your headline and first paragraph must clearly signal what the story is about, including the who, what, and where. Entity recognition — Google understanding that your article is about a specific person, organisation, or location — matters more than keyword density. Use the exact name of a person or organisation consistently rather than varying it with synonyms.
Does NewsArticle structured data improve rankings?
Not directly — structured data does not boost rankings as a direct signal. However, it improves how Google understands and presents your content, which can increase click-through rates via rich results including Top Stories carousels. Implementing NewsArticle schema with correct datePublished, dateModified, author, and publisher fields is recommended for all UK news publishers.
What are Core Web Vitals and why do they matter for journalism sites?
Core Web Vitals are Google's user experience metrics: Largest Contentful Paint (LCP, loading speed), Interaction to Next Paint (INP, interactivity), and Cumulative Layout Shift (CLS, visual stability). Poor scores are a confirmed ranking signal. For journalism sites the most common problems are large hero images without proper sizing, third-party advertising scripts delaying interactivity, and layout shifts caused by ads or embeds loading after the page.
Should journalism sites have an llms.txt file?
llms.txt is an emerging convention (not yet a formal standard) that tells AI crawlers which content on your site is available for training or summarisation. Some UK publishers are experimenting with it to communicate licensing terms to large language model providers. Whether it has any practical effect on AI crawlers today is unclear, but it costs nothing to implement and signals intent about content use.