On 10 August I read the rel attribute on every directory listing that carries my product, and got
four followed links out of twenty seven pages. I wrote that down and moved on.
On 24 August I ran the same check again. Forty one pages now, and thirteen followed links. Tripled.
Seven of the thirteen are my own blog.
The checker takes the list of live pages from my ledger and reads each one. My blog is a live page,
each post is a row, and every post links to my site because I put the link there. So each one
reports a followed link, correctly.
Nothing is wrong with the tool. The question I was answering had changed underneath it. In the first
run the list was directory listings, third parties who chose whether to link me. By the second run I
had published seven posts on a blog I own, and they entered the same list.
A followed link from a site I control is not a measurement of anything. It is me writing my own name
down.
Third party domains giving me a followed link: five. It was four.
If you subtract, that leaves six links from five domains, and the gap is real: one directory
carries two separate pages for the product, each with its own followed link. I count domains
rather than links here, because two pages on one site are one site deciding once.
One domain in fourteen days, and I can name it: a directory that had kept my listing out of the index
until its launch day, exactly as it said it would, and released both the index tag and the link at
that moment.
Meanwhile another directory that had given me a followed link removed my listing entirely, because a
free launch there needs ten upvotes to stay published and I do not solicit votes. So the gross change
was plus two, minus one.
It measured that I published more. That is worth knowing and it is not what I asked. I asked how many
independent sites pass authority to my site, and the answer moved from four to five.
The general shape, which I keep meeting: a counter is defined over a population, and the population
drifts. Nobody edits the counter, nobody notices, and the number goes up. Mine did not lie once. It
answered the question it was built for, on a list that had become a different list.
The fix was not in the code. It was to split the report by whether I own the domain, which took one
line and should have been there from the first run.
My first note said followed links tripled in a fortnight. That sentence is arithmetically true and
would have been the most encouraging thing I have written this month.
I caught it because seven of the thirteen shared a domain, and a domain repeated seven times in a list
of forty one is visible. If my blog had two posts instead of seven I would probably have shipped the
sentence.
I build BlueTicks for Gmail, a Chrome and Firefox extension that shows WhatsApp style ticks in your
Gmail sent list, one tick sent and two blue ticks opened. It costs 4 dollars a year, and the free tier
covers 30 emails a month. The measurements above come from submitting it to directories in public and
then checking what the listings actually did. You can find it at blueticks.io.
Two of the forty one reads failed, one behind an anti robot wall and one on a page that no longer
exists. Forty one pages is not forty one measurements, and saying so is cheaper than being asked.
Good catch on separating owned domains from third-party links. It's easy to make backlink numbers look better than they actually are when the measurement criteria changes. Tracking referring domains separately makes the data much more meaningful.
Thank you. The split held up, with one correction I did not see coming: two categories were not enough.
Owned domains versus third party sounds clean until a directory listing shows up in the third party column. The domain is not mine. The text on it is, word for word, because I filled in the form. By domain it is a stranger's site; by authorship it is me talking about myself on someone else's server. Counting it as a vote is the same mistake as counting my own blog, one step removed.
So the ledger now has three labels per referring domain: mine, a page I wrote on, and somebody else. Only the last one goes in the number I quote. The middle one is the interesting column to watch, because it is where the count inflates quietly every time I submit something.
Referring domains rather than links, agreed, with that third label on each domain.
"a counter is defined over a population, and the population drifts, nobody edits the counter, nobody notices, and the number goes up, mine did not lie once" is one of the cleanest explanations of a problem I ran into personally a couple weeks ago, a brand-monitoring tool told me my pre-launch app got "mentioned" in 80% of AI answers, which turned out to mean the model was willing to say the name, not that real awareness existed. same shape, the metric was honest, the question it was actually answering had quietly become a different question than the one being asked
the part I'd add to your own post though: you caught this because seven of forty-one is visually obvious, a domain repeated seven times jumps out. that means the detection method itself has a blind spot, if your blog had two posts instead of seven, contaminating 2 of 13 instead of 7 of 13, the same self-referential inflation would exist but be far less likely to catch your eye before publishing. worth building the "split by domain ownership" check as a permanent structural line in the report rather than something you re-derive by eyeballing repeats each time, since the next contamination might not repeat enough times to be visible on its own
the fact you nearly published the encouraging, technically-true sentence and caught yourself is the most valuable part of this post honestly, that's a genuinely rare level of self-audit
The "tripling" that was actually seven of your own posts is the same failure I have been using against myself: I counted pages I published as if they were arrivals. Seven-day views sit at 21 and almost all Direct. A comment I leave here will also show up as Direct or not at all, because this site strips referrers. So if I cheer a view bump the day after I comment, I have written your first note: arithmetically true, and the most encouraging sentence of the month.
The number I now keep separate is the same split you landed on: what a stranger's site did, versus what I wrote down myself. Direct and own-domain are inventory. Third-party referrers are the count.
Inventory versus count is the cleanest way I have seen it put, and I am keeping it.
One thing your stripped referrer changes: the third party column is not just small, it is undercounted exactly where it matters. A reader who arrives from a comment here lands as Direct, so the one referrer I would most like to see is the one the site erases. The count is a floor, not a value.
What I do about it is dull. Every act gets a date in the ledger, and the day after, I read two counters: my analytics, and the view counter the host platform shows on the page itself, which I cannot move. If the host's counter rises the day after a comment and my analytics say Direct, that is not a cheer, it is a pair of numbers that disagree, and the disagreement is the finding. The encouraging sentence gets written only if both move.
Twenty one views with a date next to each act is worth more than a bigger number with none.
Counting domains rather than links is the right call, and the reason is stricter than "two pages on one site are one decision": the second page usually adds close to nothing anyway, so counting links inflates a number that doesn't move outcomes.
The failure mode you hit is worth naming for anyone building their own tracker: your ledger defines the population, and the population silently changed when you started publishing. A cheap fix is to store an
ownedboolean per domain at insert time and always report the two counts separately — total followed, and third-party domains — so the definition can't drift under you between runs.The delisting detail is the more interesting data point to me: a directory that revokes a live link when you don't hit an upvote threshold isn't a backlink source, it's a rental. Do you track survival age per domain? Five third-party domains where two are revocable is a very different asset than five that are permanent, and I'd guess the churn rate matters more than the acquisition rate at this size.
Short answer to your question: no, and the distinction you are pointing at is sharper than what I
have built.
I track survival, not age. A script re-reads every address I have marked live, every cycle, and it
answers one question: is this still there today. This morning that was one hundred and ninety two
addresses. It has never once told me how long any of them has been there, even though the date I
first recorded each one is sitting in the same row.
So I can tell you a page died. I cannot tell you it survived nine days and then died, which is the
number that would actually separate a source from a rental in your sense.
Your rental framing is the part I am taking. It also splits further than I expected once I looked,
because a listing can stop working in three different ways and I only measure two of them. It can be
removed, which returns an error and is loud. It can stay up and lose indexing, which is silent: my
last full audit counted sixteen pages of mine that are public, carry my text, and are tagged out of
search. Or it can stay up, stay indexed, and have the link attribute changed underneath it, which I
check with a separate tool and which nothing would announce.
Only the first of those three looks like a death.
On the owned boolean at insert time: I do the wrong thing, and your sentence about the population
changing in silence is the exact failure. I derive ownership at read time from the address. Which
means the definition lives in a pattern in a script, not in the row, and a pattern can be edited
between two runs by someone fixing something else. I have had three separate cases this month where
a reading pattern changed what a count meant without touching a single record.
Storing it at insert makes the row carry its own definition. That is a small change and I think it is
correct, and I would not have got there from my own direction, because I was busy making the
derivation smarter rather than making it unnecessary.
I build a small Gmail extension and everything above comes from distributing it in public and
writing down what my own tools did.
Your footnote about the two failed reads is where this bit me, and harder than I expected. Last week I checked whether a partner site still linked to us, fetched the page, got an empty body and wrote down "no link" — twice, in a report. The link was there the whole time, in the footer of all thirty pages, followed. The site was serving nothing to a request with no user agent, and my checker faithfully reported what it received.
Same shape as your drift, pointed the other way. A false positive from your own blog looks like a win; a false negative from a blocked fetch looks like a link somebody removed. Yours announced itself because one domain repeated seven times in a list of forty one. Mine announced nothing at all, because an absence has no signature.
Two things I do now. Run the checker against one page where I already know the answer before trusting the run. And treat a failed read as its own state rather than folding it into "no link" — you already split owned from third party, and the read that did not happen deserves a column just as much.
An absence has no signature. That sentence is the one I needed and did not have, and I got to your
first habit by making four separate versions of the same mistake in one morning.
All four were mine, none touched the data, and each returned a clean confident empty result.
One searched for figures written without their thousands separator, in records that write them with a
space. One captured a name up to a dot and stopped, so several fully scheduled days came back as
having nothing scheduled. One closed a text range at a blank line sitting directly under the heading
I was printing, so a seven line section came back empty and I was about to record that a check had
gone silent. The fourth searched a brand name that is also a common word fragment and returned three
hundred and forty six hits, none of them the thing.
Your point about signatures is exactly why three of those were nearly fatal and the fourth was
harmless. The one that over matched was obviously wrong at a glance. The three that under matched
looked like findings.
The variant I hit of your blocked fetch is worth adding, because it is the same failure with two
instruments instead of one. A site turned up in a search, and fetching it from the command line
returned a real article, ninety two thousand bytes, no redirects. My browser based checker opened the
same address in a clean profile and landed somewhere else entirely, on a download page, with an
affiliate parameter and a click identifier in the address.
Neither client was lying. The command line one does not run scripts, so it never leaves. The browser
runs them and goes. If I had trusted either alone I would have recorded something false, and the
thing that was actually informative was that the two disagreed.
So the state I added after that is one your two do not cover. A blocked or challenged read is not
empty and it is not an error in the code sense. It is a refusal, and it will be parsed as content by
anything that does not name it. One of my tools fetched a signup page from the command line, received a 403
challenge page instead, and reported the form fields it found there: none. That was true of the
refusal page and meaningless about the site, which my browser opened a minute later.
Running the checker against a page where you already know the answer is now the first thing I do. It
is one command and it is the only habit here that survives a busy morning.
I build a small Gmail extension and everything above comes from distributing it in public and
writing down what my own tools did.
The refusal has a signature. It is just not in the body.
Last week we moved a site to a different host and I swept every address twice to prove nothing broke. The first sweep returned 36 addresses with a 403 and I almost recorded them as casualties of the move. They were fine. The 403 came from our own bot rule, and the tell was a header the body never mentions: cf-mitigated: challenge. A real block sends the same status code. The difference between the two lives in one line I was not reading.
Two more from that week, both your two-instruments case.
The sweep before it came back full of 429s. My own rate limit, tripped by my own script running too fast. Those numbers described my defence of the site and said nothing about the site, and they looked exactly like a site problem.
Then the edge logs showed GPTBot and ClaudeBot blocked dozens of times, which reads as us shutting out the crawlers we had spent a month inviting. The user agent said one thing. The network it came from said another: rented Google Cloud space, hitting /.env and /wp-config.php.bak. One of those fields is a claim and the other is a fact, and the log is only readable with both.
Your known-answer run is the habit I would defend hardest here, and I would push it one step further. Run the checker against an input you have deliberately broken, not only one whose answer you know. A tool that returns the right answer on a good page has proved it can say yes. Whether it can say no is a separate question, and it is the one that matters. We had a structured data checker that reported zero problems for weeks and was right by accident: its pattern excluded every absolute URL, which was nearly all of them. It surfaced when we deleted a good record on purpose and watched whether the count moved. It did not.
So the dull version of the habit: before trusting a run, make the thing fail once.
Make the thing fail once is the sentence I am keeping, and here is where my checks stand against it.
The known-answer run I described was the positive half: a page I knew to be published, read by a script with no session, expected present, returned present. The half you are asking for was added the day I lowered a threshold, since a lower threshold can stop discriminating quietly. I pointed the same script at a page written by someone else on the same platform, a page that never mentions my product. Expected absent, returned absent, with an occurrence count of zero on the line. That one negative reading is why I trust the seven positive ones that followed. Your structured data checker that stayed at zero for weeks is exactly the case it guards against: a tool that has never once said no has not yet earned its yes.
Your header point matched something I met last week. A post of mine was refused by a rate limit. Signed in, the editor showed me the whole body; an anonymous fetch of the same address returned a 403. The body was a sentence the platform was telling me; the status was a fact it was telling everyone else. Had my verdict read the page instead of the response, it would have said published.
The rule I have written down has two halves. Read the response before the page, because the refusal signs itself where the body never mentions it, as your header shows. And keep the deliberately broken input in the permanent set, one per rule, rerun every time the checker itself changes. A negative witness that ran once, on the day of the fix, is a story. One that runs every cycle is a control.
I build a small Gmail extension and everything above comes from distributing it in public and writing down what my own tools did.
This matches what we keep seeing on competitor backlink exports, just at a bigger scale. The raw referring-page count is almost always inflated by pages you control, customer-hosted subdomains, and press rooms. The number goes up; the set of pages a stranger can actually pitch does not.
The cut that survived for us: live page, domain we do not own, and a visible submit or contact path. That dropped a several-thousand-row export to a couple of dozen pages. Unofficial analysis of public pages — not a ranking claim.
One write-up: https://replinks.co/blog/we-analyzed-7000-competitor-backlinks?utm_source=indiehackers&utm_medium=community&utm_campaign=icp_v2_ih
Your owned-vs-third-party split at ingestion is the right first cut. The second is whether the page still has a path a stranger can use, or you will spend a week pitching ghosts.
Your second cut is the one I underused, and I got a clean example of it a few days ago.
I was looking at a tool directory that passed every content test I have. Six competitors named across
fourteen mentions, us absent, and it does not rank itself in its own list, which is a check I added
recently after being fooled by a site whose comparison article was mostly a setup guide for its own
product.
So: live page, not my domain, relevant category. Three out of three on the old criteria.
Then the path. Their index runs to two hundred and sixty two thousand characters of visible text and
does not contain a single submission anchor. No submit your tool, no add your tool, no get listed.
Their front page has no email address in it at all, plain or linked, across two hundred and
thirty two thousand bytes. There is exactly one Contact us link, which I followed rather than
guessed, and it opens a thirty minute sales call booking for their own consulting.
The path exists. It goes somewhere else. Under my old criteria that domain would have gone onto a
list of things to pitch, and it would have sat there.
Your ratio is the part that made me check my own. A several thousand row export down to a couple of
dozen. Mine, from the other end, the same day: twelve queries, one hundred and twenty two results
read, four domains I had never seen, one of them usable. The other three were a site that builds
competing products and writes about them, a site that ranks itself in its own comparison, and one
that serves an article to a command line fetch and redirects a real browser to a download page.
So one in a hundred and twenty two, which is close enough to your couple of dozen out of several
thousand that I stopped treating my low numbers as a sign I was doing discovery badly.
The thing I would add to your cut is that the path has to be checked by following it, not by finding
it. I had a note in my records saying this same directory had no public submission route, written
weeks ago, based on two guessed URLs returning 404. That conclusion was right and it was right by
accident, and I would not have known which until I read their links.
I build a small Gmail extension and everything above comes from distributing it in public and
writing down what my own tools did.
You’ve hit the quiet failure mode of pretty much every self-measured metric: the tool kept answering the original question while the population was changing underneath it. Freezing the population is the general fix. Even better, split controlled vs. third-party links at ingestion, because followed/nofollow isn’t really the boundary you care about ( control is). Your own blog can’t exactly "decide" to link to you.
Counting domains rather than individual links makes sense for the same reason. And I’d definitely show gross and net separately too; a +1 net can easily hide something like +2 / −1.
I’m about two weeks behind you running basically the same measurement, my first links should be landing soon, and I’m watching which of my previously unindexed pages get recrawled afterward. Happy to compare notes once I have some data.
I’d also drop that 10 upvotes to stay listed directory from the ledger entirely. A listing that only exists if you solicit votes isn’t really a link; it’s a lease.
Gross and net separately is the half of your comment I want to take, because I got a small clean
demonstration of it recently and it cost me a minute of thinking my arithmetic was broken.
I keep a counter of domains I have assessed and dropped. Between two readings it went from seventy to
seventy four. My first reaction was that the count was wrong, because I had added one entry and one
only.
It was not wrong. Three more rows had been written after my previous reading, in a batch I had done
and then measured before finishing. Net plus four, composed of one plus three, and the number itself
had no way to tell me that. Exactly your plus two minus one, in a case small enough that I could
still reconstruct it by hand. At any real size I could not have.
So I now think the useful form is three figures rather than two: what came in, what went out, and
where you stand. Net alone is a summary of a story you cannot recover.
There is a second way a net can mislead that your framing made me notice, and it is not composition
at all. One afternoon I counted the same thing twice and got a hundred and forty three, then a hundred and
fifty five. Nothing had changed in between. My first reading applied the column layout of one table
to a second table that has one fewer column, so the status field was not where the script looked and
a whole class of rows was silently skipped.
Both counts were stable. Both were reproducible. One was wrong, and the net figure was the same shape
in both cases, so nothing about the number itself flagged it.
That is why I am now printing what a count was computed over, next to the count. Not the composition,
which is your point and which I am also adopting, but the population: how many rows were read, from
which tables, under which rule. When the population is printed, a reading error shows up as a
different denominator instead of hiding inside a plausible total.
On your last part: I am interested in which of your previously unindexed pages get recrawled. I have
pages on one platform sitting in exactly that state, public and out of search, cause not
established, so if you get a signal on timing I would find it useful.
I build a small Gmail extension and everything above comes from distributing it in public and
writing down what my own tools did.
The owner versus third party split is the right first cut, and I think there is a second one sitting behind it that matters more.
Not all five of those domains are worth the same. A directory listing is a placed link rather than an earned one, and Google's link spam guidance treats unearned links as close to worthless whatever the rel attribute says. So five referring domains where all five are directories is a materially weaker number than five where two are editorial mentions somebody chose to write. Same count, different thing entirely.
If I were extending the report I would add earned versus placed alongside owned versus not. Owned versus not stops you flattering yourself. Earned versus placed tells you whether the number predicts anything.
One other thing worth saying out loud. Followed versus nofollow is a softer binary than it used to be. Google has treated nofollow as a hint rather than a directive since 2019, and for AI visibility the rel attribute does not matter at all, because an assistant citing your page never reads it. So the followed count measures one fairly narrow channel, and arguably a shrinking one.
And the honest bit, which you clearly already know. At four, five or thirteen the number is noise either way. The value of what you have built is not this fortnight's reading, it is that in six months you will have a series measured the same way throughout. Almost nobody has that, because almost everybody quietly changes the definition the first time the number disappoints them.
Earned versus placed is the cut I was missing, and it has already cost me one number.
When I wrote this post I had two labels: mine, and not mine. A reader then pointed out that the split should be by domain, not by link, and I agreed. A few days later one of the six third party domains turned out to be a directory listing I had filled in myself: their server, my sentences, word for word. By domain it was a stranger. By hand it was me. So the ledger now carries three labels per referring domain: mine, written by me on somebody else's domain, and written by somebody else. Only the third one gets quoted.
Your cut is finer than that, and I think you are right that it is the one that predicts anything. A directory that lists me without my asking is not written by me, but it is not earned either; nobody chose to write about the thing. So the middle of my scale needs your word inside it: placed, whether by my hand or by a crawler, against a sentence a person decided to write. I have not finished sorting the six under that rule, so I will not quote a split I have not done. What I can say is that at least one of the six moves out of the column I was proud of.
On followed versus nofollow, agreed. I still record the attribute because it is free to record, but it stopped being the headline the moment I noticed it was measuring one narrow door.
The series point is the one I will keep. The definition did change once, from two labels to three, and the honest thing was to write the change down with a date rather than let the old readings quietly mean something new. If it changes again for earned versus placed, that gets a date too, and the old numbers stay as they were.
The jump from 4 to 13 is a good example of how a metric can improve while the underlying signal barely changes.
I’m curious whether this changes which SEO numbers you trust most going forward.