Content Marketing Intelligence: Which Content Earns Its Production Cost
Content Marketing Intelligence applies NLP, density clustering, path modelling and survival analysis to content libraries — mapping topical saturation and gaps, detecting editorial drift, reallocating attribution beyond last-touch, and reconstructing where readers actually disengage. It answers a production question: what should be written, refreshed, consolidated or retired.
Every Standard Content Metric Has a Measurement Flaw
Content analytics has run on the same small set of numbers for fifteen years: pageviews, sessions, time on page, bounce rate, scroll depth, conversion rate. Each one has a specific, well-known problem, and together they produce a function that publishes a great deal and cannot demonstrate which of it created value.
- Time on page cannot be measured for the last page in a session — and where it can, it does not distinguish reading from an abandoned browser tab
- Bounce rate treats two opposite outcomes identically — the reader who found exactly what they needed and left, and the one who found nothing and left
- Scroll depth cannot tell reading from scanning — someone hunting for one fact and someone reading every line produce similar numbers
- Conversion rate credits the landing page — ignoring the sequence of content that made the visit happen at all
- Assisted conversion reports use arbitrary credit rules — linear, time-decay and position-based are conventions, not measurements
- Nothing detects a library competing with itself — or a piece that has drifted away from how the audience now talks about the subject
The consequence is a content programme where production, refresh and retirement decisions are made by whoever has the most authority in the room, because there is no evidence in the room that could settle it differently.
Where This Ends and SEO Diagnostics Begin
Worth being explicit, because two of the solutions below touch subjects that also appear in Organic Growth Intelligence, and running both without understanding the boundary wastes money on overlapping work.
Why is this page losing position, and what is competing with it? Semantic distance from the current SERP, authority flowing through the internal link graph, which specific pairs of URLs are splitting intent.
The output is a fix list for pages that already exist and are already ranking.
What should we write next, what should we refresh, what should we retire, and which pieces justified their cost? The shape of the library as a whole, where demand exists with no coverage, and where budget is being spent on ground already held.
The output is an editorial roadmap and a budget argument.
They overlap deliberately at one point — a library that is over-saturated is both a ranking problem and a production problem — and the audit establishes which framing your situation actually needs. Most engagements need one, not both.
Diagnostics on the Library, Not the Page
Each solution targets one structural failure in a content programme, uses a named method, and carries a stated limit. They are run selectively — the audit determines which apply.
Topical Topology and Saturation Mapping
Transformer embeddings + UMAP + density clustering
The problem: Content libraries carry two opposite faults at once. Over-saturation, where several pieces occupy the same semantic territory and split authority between them — which is broader than keyword overlap, because pieces targeting different phrases can still address one underlying need. And under-coverage, where genuine audience questions have no piece at all, often because keyword tools do not surface demand that has not yet accumulated search volume.
The method: The library is embedded, projected into a navigable topology, and clustered by density to reveal where content is stacked and where space is empty. The topical territory audiences are actually searching is modelled separately and compared against that coverage, producing a gap list ranked by demand, competition and the site’s existing authority.
An honest limit: below roughly fifty published pieces there is not enough structure for topology modelling to reveal anything an editor could not see by reading the list. This is a technique for libraries, not for content programmes still finding their shape.
Topic and Editorial Voice Drift
Topic modelling + temporal shift detection against a voice profile
The problem: Libraries decay in two dimensions. Topical decay, where a piece that matched an audience need at publication drifts away from how people now think and search about the subject. And stylistic drift, where multiple contributors over years — or one whose voice has evolved — leave a library that no longer sounds like one brand. Both are invested budget quietly losing value, and neither shows up in any standard content audit.
The method: Topic distributions are extracted across the library and compared against two reference points: what currently ranks for the same queries, and a voice profile derived from the brand’s own strongest work. Pieces are ranked for update, consolidation or retirement by drift magnitude weighted against current traffic contribution.
An honest limit: stylistic drift is not automatically bad. A brand voice that evolved deliberately should not be measured against its own past. The model finds divergence; deciding which direction is correct is an editorial judgement and stays with a person.
Multi-Touch Content Attribution
Markov removal effect + Shapley value allocation
The problem: Last-touch attribution is not a neutral default. It systematically undervalues everything above the bottom of the funnel, which is why educational content appears to generate nothing and gets cut — removing the foundation the rest of the funnel was standing on.
The method: Content touchpoints are structured as a directed graph. Markov modelling estimates each piece’s removal effect — how much conversion probability the graph loses without it — and Shapley value allocation distributes credit by marginal contribution across possible orderings. Together they produce a far more defensible distribution than crediting the last click.
An honest limit, and it is important: this is not causal. Removal effect is a model-based reallocation of paths that were observed, not evidence that a piece caused a conversion. It is better than last-touch and it is not proof. Establishing genuine incrementality needs holdout experiments or marketing mix modelling. It also requires user-level journey data, which consent requirements and cross-device behaviour have made materially less complete than it was.
Scraping and Near-Duplicate Detection
MinHash shingling + locality-sensitive hashing + embedding similarity
The problem: High-performing content gets scraped and republished, sometimes within hours. Standard plagiarism checkers find exact and near-exact text matches. They do not find semantic paraphrasing, structural copying that reuses your architecture with different surface language, or systematic republication with randomised sentence-level variation.
The method: Compact fingerprints of each original piece are indexed for efficient similarity search, and a monitoring set — known aggregators, scraper networks, competitor domains, and index searches on distinctive phrasing — is checked against it on a schedule. Embedding similarity catches the rewritten cases that fingerprinting alone misses. Findings are ranked by the authority of the offending domain and estimated organic impact, so takedown effort concentrates where it matters.
Two honest corrections: nothing crawls the whole web — this monitors a defined and expanding set, and coverage is a scope decision. And scraped content is not a link problem: disavow files do nothing here, and you cannot place a canonical on someone else’s site. The realistic remedies are takedown requests, self-referencing canonicals and fast indexation of originals.
Micro-Engagement Dropout Modelling
Survival analysis on session-level scroll and reading behaviour
The problem: Knowing that average scroll depth is sixty-five percent tells you nothing actionable. It does not distinguish readers who finished and left satisfied from readers who lost interest at a specific structural point from readers who found their answer early and had no reason to continue. Those three outcomes need three different responses, and the aggregate number hides all of them.
The method: Session-level behavioural events are pulled from BigQuery and used to reconstruct approximate reading trajectories. Survival analysis models the hazard of dropout at each position in a piece — the probability of leaving at that point, given that the reader reached it — revealing where engagement terminates at anomalous rates rather than where it naturally completes.
A prerequisite worth knowing up front: a default analytics setup does not produce this data. Standard scroll tracking fires at a single threshold near the bottom of the page, which is far too coarse to reconstruct a trajectory. Granular scroll and reading-pause instrumentation has to be implemented first, and then data has to accumulate before any modelling is possible. That sequencing is usually the longest part of this particular engagement.
Content structure decisions — where the call to action sits, how long the introduction runs, whether the technical section belongs before or after the example — are almost universally made on editorial instinct.
A dropout hazard curve is the first evidence most content teams have ever had about where readers actually leave. It frequently contradicts the assumption everyone was operating on, and it applies across the whole library rather than to one page.
It also compounds with the attribution work: knowing which pieces carry the conversion path, and separately knowing where readers abandon those specific pieces, is a much sharper instruction than either finding alone.
How This Connects to the Rest of the Practice
Content Intelligence uses the same methods documented elsewhere on this site, pointed at editorial data. That matters for sequencing more than for cost — some of this work only makes sense once other things are in place.
Ranking diagnostics run alongside
Incrementality sits above attribution
Clustering carries the same warning
The sequencing that actually works: fix the ranking problems on pages that already earn traffic, then map the library to decide what to produce, then instrument engagement so the next round of production is informed by evidence rather than by the previous round’s assumptions.
The Questions Serious Clients Ask
We track content performance in Google Analytics. Why is this different?
Analytics gives you accurate descriptive aggregates — what happened, at the surface. This works on raw session-level event data, reconstructs individual reading behaviour, models the library as a structure rather than as a list of pages, and produces prescriptive output: what to write, refresh, consolidate or retire. The question is not whether you are tracking content; it is whether anything in your reporting could settle an argument about what to produce next.
Our content ranks well. Do we still need saturation mapping?
Possibly, and possibly not. Content ranking well in an over-saturated area is operating below its potential, because authority is split across pieces competing for one intent rather than concentrated. But the honest version is that consolidation carries risk — merging two ranking pages can lose you both — so this is worth doing when the library is large enough that the pattern is systematic, not when two pages happen to overlap.
Is Markov attribution actually causal?
No, and any page claiming otherwise is overstating it. Removal effect estimates how much conversion probability a graph loses without a given node, based on paths that were observed. That is a substantially better allocation than last-touch, and it is not evidence that the content caused anything. Genuine incrementality requires a holdout or an aggregate causal model. Attribution output belongs in a budget argument as strong evidence, not as proof.
What data does the attribution work need?
User-level journey data — every content touchpoint, in order, for converting and non-converting users. This is the prerequisite that most often ends the conversation, because consent requirements, cross-device behaviour and tracking prevention have made that data materially less complete than it was. The audit establishes how much of your journey data is actually reconstructable before any modelling is scoped.
Can you find everyone scraping our content?
No. Nothing crawls the entire web, and any vendor implying otherwise is describing something that does not exist. What is realistic is a defined monitoring set — known aggregators and scraper networks, competitor domains, and index searches on distinctive phrasing — expanded over time, with detection that catches paraphrased and restructured copies rather than only exact matches. Coverage is a scope decision with a cost attached, and it is discussed as one.
What can actually be done about scraped content?
Less than most people assume, and it is worth correcting two common suggestions. Disavow files do nothing here — they address links pointing at you, and scraped content is not a link problem. And you cannot place a canonical tag on someone else’s site. What does work: takedown requests prioritised by the offending domain’s authority, self-referencing canonicals on your originals, and getting new content indexed quickly so your version is established first.
How long before the engagement modelling produces anything?
Longer than the others, because a default analytics setup does not collect the necessary data. Granular scroll and reading-behaviour instrumentation has to be implemented, and then enough sessions have to accumulate for survival analysis to be meaningful per page. That sequencing is stated up front rather than discovered later, and it is the reason this solution is usually scheduled after the others rather than alongside them.
Who Content Marketing Intelligence Is Built For
This suits organisations with libraries large enough to have structure, content budgets large enough that allocation matters, and someone being asked to justify the spend. It is a poor fit for young content programmes, where the honest answer is to publish more and model later.
B2B SaaS and technology
Publishers and content-first businesses
Enterprise teams defending budget
E-commerce with editorial programmes
Agencies managing content at scale
Work is delivered remotely from Lahore, Pakistan, for brands across Pakistan, the United Kingdom, the United States and the UAE. Content topology is modelled per market and per language — semantic space differs enough between markets that a topology built on one country’s SERPs does not describe another’s, and multilingual libraries need separate embedding treatment rather than one pooled model.
Where AI Agents Fit Into Content Work
A SaaS tool is someone else’s generic model. An AI agent is your own model, run autonomously. Cognitive Intelligence decides what to build; agents are how it keeps running. Content is the area where the distinction matters most, because generating text is the one thing language models do effortlessly — and it is not the part that was ever the bottleneck.
Built on ML and data science. The agent decides its next step from live data — flagging a piece whose topic distribution has drifted past threshold, detecting a new gap opening in the topology as search behaviour shifts, spotting a near-duplicate appearing on a high-authority domain.
n8n, Make.com, Zapier. Scheduled fingerprint checks against the monitoring set, embeddings refreshed after publication, engagement data pulled into the modelling table, takedown drafts queued for review, drift reports circulated to the editorial team on a cadence.
Model Context Protocol lets an agent query the CMS, BigQuery, Search Console and the analytics warehouse directly rather than working from exports. For content work this also means an agent can read the actual library rather than a summary of it.
So every run meets the same standard: the same embedding model and drift threshold, the same clustering validation, the same similarity cut-off before a duplicate is flagged. Skills are what stop an automated content review from changing its own definition of a problem between runs.
What Stays With a Person
The part nobody else writes, and on this page it carries more weight than elsewhere — because content is where the temptation to automate the wrong thing is strongest.
- Writing the thing. A model can tell you a gap exists and that a piece has drifted. It cannot supply the argument, the example from your own work, or the point of view that makes the piece worth reading rather than worth skimming.
- Deciding which direction the voice should go. Drift analysis measures divergence from a past profile. Whether the brand has degraded or evolved is an editorial judgement, and treating divergence as error would freeze a voice that ought to change.
- Choosing what to consolidate or retire. Merging ranking pages carries real downside, and the commercial, brand and legal considerations behind retiring a piece are not in the data.
- Calling it off. If the diagnosis is that the content is competent and the problem is the offer or the product, someone has to say so instead of prescribing another content roadmap.
Content marketing agent channels are being documented separately. The AI agents hub is the current starting point, and the SEO agents section covers the closest adjacent work.
Frequently Asked Questions
What is topical saturation mapping?
It is embedding a whole content library into semantic space and clustering it by density, to see where several pieces occupy the same territory and where territory exists with no coverage at all. It is broader than keyword cannibalisation, because pieces targeting completely different phrases can still address one underlying informational need — which is how a library ends up competing with itself without any keyword report showing it.
How is this different from an SEO content audit?
An SEO audit asks why specific pages are losing position and what is competing with them. This asks what the library should contain — what to write, refresh, consolidate or retire, and which pieces justified their production cost. Ranking diagnostics fix pages that exist; this decides what should exist. They overlap where over-saturation is both a ranking and a production problem, and most engagements need one framing rather than both.
Why is last-touch attribution a problem for content specifically?
Because content works furthest from the conversion. Last-touch credits whatever happened immediately before the purchase, which is usually a branded search or a retargeting ad, and gives nothing to the piece that created the awareness months earlier. Top-of-funnel content therefore appears to generate no revenue, gets cut, and the funnel that was standing on it degrades some months later without an obvious cause.
How much content do we need for this to be worth doing?
Roughly fifty published pieces before topology modelling reveals anything an editor could not see by reading the list, and considerably more before drift analysis is meaningful. Below that, the honest recommendation is to keep publishing and revisit later. The exception is attribution work, which depends on journey data volume rather than on library size.
Does scraping actually damage our rankings?
Sometimes, and less often than people fear. Search engines are generally good at identifying the original, and the more common damage is fragmented authority — the engagement, links and social signals that would have accrued to your version accruing elsewhere instead. The cases worth acting on are copies on high-authority domains and systematic republication at scale, which is why detection output is ranked by estimated impact rather than presented as an undifferentiated list.
What is dropout hazard modelling telling us that scroll depth does not?
Where readers leave, and whether that point is a natural completion or a structural failure. An average scroll depth conflates the reader who finished, the reader who lost interest at a specific section, and the reader who found their answer in the first paragraph. Hazard modelling separates them by estimating the probability of leaving at each position given that the reader got there — which turns content structure decisions into evidence rather than instinct.
Do you work with businesses outside Pakistan?
Yes. Delivery is remote from Lahore, with clients across Pakistan, the United Kingdom, the United States and the UAE. Content topology is modelled per market and per language, because semantic space and SERP composition differ enough between markets that a single pooled model describes the largest one and misrepresents the rest. Working hours overlap comfortably with the Gulf and the UK, and partially with US mornings.

AI-Driven Digital Marketing Intelligence Consultant & Growth Engineer in Pakistan. Usman Saeed specializes in engineering resilient digital growth architectures — helping enterprise brands eliminate tracking data drops, secure conversion signals, and maximize profitability through E-commerce Engineering, server-side Signal Engineering, and Predictive Intelligence. With 12+ years of experience and advanced data science expertise, marketing guesswork is replaced with mathematical precision — automated systems that bridge execution with business intelligence, ensuring your investment delivers measurable scale.
Where to Go Next
Organic Growth Intelligence
Marketing Mix Modeling
Customer Segmentation
All Solutions
The Diagnosis Starts With Your Content Data
A data audit maps your library’s actual topology, checks how much journey data is reconstructable, and reports which of these five are worth running — including the case where the library is not yet large enough to model.
