Same disclosure as always: I test tools firsthand on my own site (solostacklab.net) before writing about them. I'm an affiliate for some of the tools I mention, not all. For this post: Ahrefs has no affiliate program, and Semrush declined my application.
Three things from this week.
I thought I had 332 backlinks. I have one.
I signed up for Ahrefs Free (free if you verify you own the site), because Google Search Console's Links report has never shown me anything. Ahrefs showed 332 referring domains. Then I actually read the list. Only one is a real site: indiehackers.com, 5 links, none of them dofollow. Nearly all the rest are auto-generated "SEO services" style pages that inserted my domain, and Ahrefs flags them as spam itself. It added 213 new domains in the last 30 days. Google's own help says most sites don't need to disavow, so I'm doing nothing. Weird feeling to be excited about a number for a few minutes and then find out it was mostly junk.
A bug only showed up when I looked from the server's side.
I've crawled the site with Screaming Frog and Semrush, and I have my own check script. All three said fine. Then I opened Cloudflare's AI Crawl Control page to see which bots were visiting, and started poking at the URLs they were requesting. Every URL that doesn't exist on my site was returning a 200 with the homepage, not a 404. That's a soft 404. My guess is the crawlers missed it because they only follow links that exist, so they never asked for a page that isn't there. I added a real 404 page.
One of my own published claims was wrong for four weeks.
In my Semrush review I wrote that the www version of my domain redirects to the main one. I never re-tested it after writing that. It was serving a full duplicate copy of the whole site with a 200. Found it this week, added a proper 301, and put a dated correction on the page instead of quietly editing it.
Question for anyone who's been through this: how did you get your first genuine, non-spam backlink on a new domain? I keep reading "make something people want to cite," which doesn't tell me what you actually did.
That's a good reminder to actually verify the tools instead of trusting the dashboard number. What finally tipped you off that the 332 figure was wrong?
Fresh from doing exactly this on a new domain (launched this month): my first real backlinks came from startup directories, but you have to pick the right ones. AlternativeTo and SaaSHub give you a listed profile that Google indexes fast, Crunchbase adds a credibility link, and Product Hunt / Uneed leave a permanent product page. None of it brings customers (learned that the expensive way), but as clean, non-spam backlinks on a brand-new domain, they did the job: my site went from invisible to getting impressions in Search Console within days, and the only clicks come from my actual market.
So my practical answer to "make something people want to cite": before anyone cites you, get the boring foundational links from legitimate directories in one efficient pass, then go where your customers actually are.
And thanks for the soft-404 tip — checking my own site for that this week.
Clear and practical, thanks. Did anything surprise you along the way?
You're sitting on the answer and calling it a blog. "I test tools firsthand before writing about them" is one of the few genuinely citable things a new domain can publish. Three things that got my first real links on a fresh domain: 1. Original data, packaged for lazy journalists. Your tool-comparison numbers, the soft-404 discovery — each is a mini-study. Give each a headline stat and a methodology paragraph. Listicle authors updating "best SEO tools" posts need fresh sources to look current. 2. Be the correction. Publish a short "we re-tested X, here's what changed" series. Nobody links to the tenth identical review. People link to the one that admits it was wrong. 3. Answer questions where the answerers are bloggers — niche forums and blog comments where the site owner writes roundups. One thorough answer with a "full methodology here" link outperforms ten directory submissions. And on the spam domains: that's just what happens once Ahrefs knows you exist. Ignore it.
Solo founder here. The distribution problem is real. I spent months building, then realized I had zero pipeline. Now I spend 50% of my time on outreach and content. The product is table stakes — distribution is the moat.
What made you pick this stack over the alternatives?
Nice work shipping it. What has been the biggest challenge since launch?
Helpful post. How did you get your first bit of traction?
Your soft-404 and duplicate-www examples are a good reminder that “SEO progress” needs a verification checklist before chasing links. For a first genuine backlink, I’d make a narrowly useful artifact that answers a recurring question with original data, then send it one-to-one to the few writers or communities already discussing that problem—no “please link me” blast. A public crawl/redirect test checklist with examples, or a small dataset, gives someone a reason to cite it. I’d also log which page or claim earned each link so you learn what actually travels. Did your one real link come from a useful artifact, a relationship, or a directory?
You can post links to your pages if it allows on other editable websites - for instance - I posted a link to my website Score150.com on Aops.com because my content (AMC 10 mock tests, strategy articles for content math) WAS RELEVANT on Aops so they accepted it - of course, in this case the other site must allow user contributions so you can post your links. Otherwise, you create articles/resources on your site Google starts indexing and your site becomes known through google searches
On the first genuine link, the route most available to you specifically is the one you're already sitting on: you write firsthand tool reviews, and vendors keep press/reviews pages and marketing people who are starved for third-party coverage that isn't an affiliate rehash. Send the finished review to the vendor's marketing contact with no ask attached other than an accuracy check — here's what I measured, tell me if anything is wrong before it spreads. You get a correction pass either way, and a decent share of the time you get a link or a share from an account with actual reach, which is usually how links two and three arrive. The other durable shape for a site like yours is the dated verification page: "what the Ahrefs free tier actually includes, checked 21 Sep 2026, screenshot attached." Other writers cite that because re-verifying is the annoying part of their job, and it plays directly to the habit you just built by stamping a dated correction instead of quietly editing. One small push-back: doing nothing about the 213 new spam domains matches Google's guidance, but log the date of the influx anyway, because if rankings wobble in two months you'll want to know whether it preceded the drop rather than guessing. Question — of all the tools you've reviewed so far, have you ever actually sent the review to the vendor? It's the cheapest untested experiment on your list and it would tell you inside a week whether editorial links are reachable at your current size.
The gap between Ahrefs showing 332 domains and Search Console showing basically nothing is pretty wild. Did you ever figure out why GSC wasn’t even showing the Indie Hackers links? I’m curious whether it’s just slower to report them or if Google is filtering out way more links than Ahrefs does.
Never figured it out. The Links report has shown nothing at all for this site since launch, not even the Indie Hackers links, so I can't tell slow reporting from filtering.
Also, quick question since you seem to know the platform a bit better — I tried creating a post here yesterday but got the “you can’t create posts yet” message. Do you know how posting access usually gets unlocked? Is it based on account age, points/upvotes, or just being active for a while?
Honestly not 100% sure myself. Mine unlocked a while back after commenting on a few other people's posts, but I never saw an official rule, so I wouldn't treat that as confirmed.
That’s interesting. Have you checked whether the Indie Hackers page itself is indexed and whether Google has crawled it recently? I’d be curious if the link is known to Google but just never appears in the Links report.
Haven't specifically checked whether the IH page is indexed via a site: search. Good idea, I'll look.
This is the kind of gap that matters more than people expect, a claimed number and a verified number can tell completely different stories, and most people never go check. What tipped you off to look, did something downstream stop making sense, or did you just decide to audit on a whim?
Different trigger for each. Search Console's Links report has never shown me anything, so I wanted a second source, which is how I ended up reading the Ahrefs list. The soft 404 came from looking at what bots request in Cloudflare and then requesting a made-up URL myself. The www one I only found by re-testing a claim I should have re-tested four weeks earlier.
The four-weeks-wrong claim is the one that got me. I had the same thing, worse: I changed a number on my product (a free report went from 30 questions to 12) and then found 68 hand-written copies of "30 buyer questions" across my own marketing pages. Every one of them had been true when it was written.
What fixed it wasn't discipline, it was moving the number. There's now one module that owns every figure the site promises a reader — question counts, engine names, price per plan — and the pages import from it. Nothing is typed twice. Then a test asserts the pages and the product agree, so if I change the product and forget the copy, the build goes red instead of the site going stale.
I did the same thing for the class of bug you found from the server's side. I publish a robots.txt validator as a free tool, so I pointed it at my own site in a test: it takes every URL the sitemap claims and asks the validator whether robots.txt actually allows it. A Disallow: /s blocking /sample-report doesn't throw, doesn't show up in a crawl of pages that exist, and surfaces in Search Console three months later. Same shape as your soft 404 — the crawler never asks for the thing that's broken.
On the backlink question I don't have a useful answer, and I'd be suspicious of anyone who hands you a tidy one. I have zero genuine ones on a domain that's a few weeks old. The only bet I've made is building free tools that are better than the common version rather than the same as it, on the theory that the reason to cite something is that it does something the others don't. No evidence yet that it works.
Using the sitemap as the test list is a good idea, and cheaper than what I did. I ran it against mine just now: robots.txt is allow-all and all 32 sitemap URLs return 200, so nothing there, but it's going into my routine. The single-source part I only half have: prices live in one JSON file with a last-verified date and a script warns after 90 days, but claims in the prose have nothing.
Exactly — I’d treat that as a useful invariant, rather than a full crawl audit: every URL the sitemap asks a crawler to index must be allowed and return a real 200. The complementary checks are a deliberately nonexistent URL (to catch soft 404s), the same requests under representative bot user agents, and comparing the sitemap against the set of published routes. Each catches a different “the crawler never asked the broken question” failure.
Good breakdown, hadn't thought about diffing the sitemap against actual published routes separately — that'd catch a route that's live but never made it into the sitemap, which is a different bug than the ones I've found so far.
Appreciate the honesty here, most people only share the wins.
Really relatable. How much time do you put into this each week?
Curious how long it took before you saw the first real results?
Honestly, not yet. Six weeks in, one genuine backlink and zero conversions so far. Position on most pages is fine (3-8 range), impressions just haven't shown up yet.
Makes sense. Are you planning to charge for it, or keep it free for now?
Solo Stack Lab itself is free to read, always has been. The only money is from affiliate links on some reviews.
All matter is quality backlinks. Focus on that
Agreed, quality over quantity. Still working out how to get the quality ones though.
Makes sense. Are you planning to charge for it, or keep it free for now?
It’s easy to see a high backlink count in Ahrefs and assume everything looks good. But once you check the referring pages, you might find links that have disappeared, pages that no longer work, or links that have little connection to the business.
I’d start by checking a small sample: is the link still there, where does it point, and does it make sense on that page? A link in a relevant article gives you different information from one sitting in a footer or an unused profile. The total count alone doesn’t tell that story.
I also like your point about looking at content and internal linking alongside backlinks. Otherwise, it’s easy to focus on getting more links without checking what else needs attention.
When you went through the referring domains, what surprised you most? Were many of the links missing, or were they still there but less useful than the numbers suggested?
Honestly the volume was what surprised me, not any specific missing link. I expected maybe a dozen spam entries mixed in with real ones, not 331 out of 332. Almost none of them were dead links either — they're live pages, just auto-generated ones that never had anything to do with the site.
the wrong-for-four-weeks claim is the one that scares me most. on a comparison page the claim is usually about someone else's product, so it goes stale on their schedule, not yours.
what i ended up doing on mine: every competitor price is stamped with the date i checked it and the plan it sits on, the test suite fails if a comparison figure is missing either, and a monthly job re-fetches each vendor's pricing page and diffs it against what my page says. the visible date does double duty: readers can see how old it is, and i can't pretend i checked something i didn't.
no proven answer on the first real backlink either, sorry. following for the replies.
No shame in that, it's an open question for me too. The single-source-of-truth-for-numbers idea is a good one, might steal it for pricing specifically since that's the part most likely to drift.
Getting that first genuine backlink is definitely harder than it sounds. I’d love to hear what worked for others beyond the usual “create something worth linking to” advice.
Same. "Create something worth linking to" is directionally right but doesn't tell me what to build first.
What made you pick this stack over the alternatives?
Astro for static content, Cloudflare Pages for hosting — needed something that builds fast and doesn't need a database for what's basically a stack of markdown reviews. Nothing exotic.
The soft-404 catch is the best one here — 'crawlers only follow links that exist, so they never asked for a page that isn't there' is exactly why all three tools said fine. The common thread across your three findings: verification has to happen from the other side of the request (the bot's URL, the server's response, an outsider's read), not from your own dashboard. On the first-backlink question: the non-spam path that works is helping in public where your audience already gathers — most of those links are nofollow, but they generate brand searches and referral traffic, and genuine dofollow links tend to follow those. The dated correction instead of a silent edit is exactly what earns them.
That's a clean way to put it — verify from the other side of the request, not the dashboard. The brand-search-driven dofollow point is new to me, hadn't thought about the causal order that way.
Great catch on the soft 404 issue. It’s a good reminder that dashboards don’t always show the full picture — sometimes you need to test from the crawler’s perspective. Also, the dated correction was a nice touch for building trust.
Original data has been the clearest answer for me so far. I published a survey because I wanted better data for my own content, and other publications ended up using the findings and linking back to the research. It made me realize the useful part of “create something worth citing” is probably having something another writer actually needs as a source.
That's a strong data point, thanks. Matches what I keep hearing: original data beats another opinion piece.
The 332 to one breakdown is probably your best shot at the next real link. You actually opened the list and checked what was there, which is more useful than another general backlinks guide.
I’d put a few screenshots and the way you sorted real sites from junk into a short post, then send it to a couple of small SEO newsletters. “I checked all 332 referring domains and only one was real” is a pretty good story for them to link to.
I haven’t earned an editorial link for Padro yet, so I can’t claim this works. But your audit is exactly the sort of thing I’d want to read before trusting my own backlink count.
Already did — just published a short guide built from exactly this: https://solostacklab.net/guides/how-to-sanity-check-a-backlink-report/. Hadn't thought about pitching it to small SEO newsletters directly though, that's a good next step.
What made you pick this stack over the alternatives?
Astro for static content, Cloudflare Pages for hosting — needed something that builds fast and doesn't need a database for what's basically a stack of markdown reviews. Nothing exotic.
The soft 404 returning 200 with the homepage is a classic one that crawlers miss exactly because of what you described — they only follow links that exist, so they never test the wrong URL. Good catch doing it from Cloudflare's side rather than just the crawlers.
On first genuine backlinks: the pattern that's worked most consistently in my experience is being first to document something specific that doesn't exist yet. Not "how to get backlinks" — 50,000 of those. More like your specific Cloudflare AI Crawl Control + Screaming Frog combination uncovering something that standard crawlers missed. That's actually quite specific and linkable.
The test I use: could the person linking to this page say "I found this exact method nowhere else"? If yes, it'll get cited eventually. Your soft 404 discovery via Cloudflare could easily be linked by someone writing about crawl budget or server-side validation. The dated correction on the Semrush page is also the kind of editorial honesty that builds trust with people who cover SEO tools.
The frustrating honest answer to your question is: first backlink usually comes from someone in a community you're already in, who reads something you wrote and drops it in a Slack or newsletter. This post is more linkable than you might think.
Thanks. That "could they say they found it nowhere else" test is a useful filter, I'll use it on what I write next.
ask “are these backlinks real?” you’ve got something specific to point at. It’s not glamorous, but the first legit links usually seem to come from being the clearest answer to one narrow annoying question.mething specific to point at. It’s no
Sorry, browser posted only the tail of what I wrote. Full thought: I’d turn this exact experience into a small evergreen page: “how to sanity-check your first backlink report.” Include the Ahrefs/GSC mismatch, the spam patterns you saw, and the decision tree for disavow vs ignore. Then when people ask whether a backlink report is real, you have a concrete thing to share. First legit links usually seem to come from being the clearest answer to one narrow annoying question.
Thanks for posting the full version. A "how to sanity-check a first backlink report" page is a good idea and I already have the material for it. Might do it.
I track every single customer conversation and the pattern is clear: businesses don't churn because of missing features. They churn because of poor onboarding. If they don't get value in the first 48 hours, they're gone.
Interesting, though I don't have customers or onboarding yet. Mine is a review site, so I'm not sure how it maps. Thanks for stopping by.
Screaming Frog and Semrush staying green makes sense: they never asked for a path that isn't on the site, so they never saw the 200. Cloudflare's bot list is a different dataset. Those are URLs nobody on the team planned to request. That's how the soft 404 showed up.
The www copy had the same hole. A checker that only follows the sitemap, or links that already exist, keeps reporting that everything is fine.
Yeah, my crawlers only ever check what I already know exists. Cloudflare showed me what other people were asking for.
Did the same thing, put a placeholder reference and forgot to update. So silly!
Ha, mine was a claim I never re-tested rather than a placeholder, but same energy.
The Ahrefs 332 -> 1 experience is basically universal. Free-tier discovery is largely noise: PBN scrapes, directory aggregators, expired-domain listings that auto-inject you. The signal-to-noise gets a bit better on Search Console's Links report over 60-90 days because Google filters most of that garbage before it ever shows up there.
On your actual question — first non-spam link on a new domain, from what has worked for me and people I trade notes with:
Original data or a repro. Your "332 -> 1" post is exactly the kind of thing that earns links. Take one tool, one claim, actually test it, publish the raw numbers. People writing round-ups need a source to cite; you become it. Hisashi Space's traffic post that's on the front page right now is another example — he'll get links from that piece for years.
A free tool no-signup. Doesn't have to be complex. A calculator, a converter, a checker. Even a well-formatted comparison table that people bookmark. The link magnet has to solve one small painful thing in under 5 seconds.
Guest posts on niche blogs, not "SaaS growth" mega-sites. Small niche blogs run by one operator will happily take a genuinely useful 1500-word piece with a link back. Ignore the DR score, look at whether they've published in the last 3 months.
On the soft 404: nice catch. Cloudflare's crawl view is genuinely one of the underused free diagnostics. Almost every WordPress site I audit has that same issue — non-existent slug returns 200 + homepage because the redirect rule is too greedy. Worth checking robots.txt for accidental Disallow lines while you're in there.
Useful list, thanks. On the Search Console part: it's been about six weeks and the Links report still shows nothing at all for this site, not even the Indie Hackers links, so I'm not counting on 60-90 days. Checked robots.txt after your note, it's allow-all.
The backlink number is noise, the wrong claim is the real story. One bad legal threshold in a guide costs more than 332 spam domains ever cost you in rankings, because a reader who catches it quietly stops trusting the other forty things you published. The rule that has held up for me: any sentence carrying a number, a price, or a regulatory limit gets the primary source URL stored beside it, and those are the only lines worth re-checking on a schedule.
Good rule, thanks. I'll steal the source URL next to every number part.
Good prompt to audit my own site. I ran your soft-404 and www checks on mine after reading: a made-up URL returns a real 404, www 301s to the apex, canonicals are set. What I didn't catch on my own was a content claim: one of my guides gave the EU micro-enterprise limit as turnover only, when the directive says turnover or balance sheet, no more than €2M. I only found it by re-reading the primary text against the article the next day.
Curious how you decide which published claims are worth re-testing: do you keep a list, or re-check on a schedule?
Both, badly. Prices have a last-verified date and a script warns after 90 days. Everything else, like the www claim, had no list and no schedule until it burned me. Now I re-test the end state after any fix instead of assuming it worked.
Your soft 404 has a second life in analytics: a 200 that serves the homepage logs as a real pageview on a URL you never published. The only place I ever see those URLs is bot request logs, since nothing links to them. That's why we keep assistant and bot hits out of visitor counts in amami.dev.
Hadn't thought about the analytics side at all. I use Clarity, which as far as I know only records when its script runs in a browser, so plain bot hits shouldn't show up. I haven't checked whether crawlers that execute JS leaked in, though. Worth a look.
The number I would watch is not 332, it is 213 added in thirty days. That is not a one off scrape, it means the domain is in an active list that regenerates, so the count keeps climbing and every backlink figure you see from here is mostly noise. Which makes right now the moment to write down the one real link, while it is still countable by hand. Agree on doing nothing about disavow, though the reason is worth stating: those pages are already ignored rather than harmless, so disavowing mostly costs you an afternoon.
Fair point on watching the 213 instead of the 332, I'll track that number rather than the total. And "ignored rather than harmless" is a better way to put the disavow reason than mine. Thanks.
The 332 → 1 realization is probably one of the more useful SEO lessons here. It’s so easy to look at a dashboard number and assume it represents something meaningful without checking what’s actually behind it.
For a new site, I’m starting to think the first real backlink is less about “doing SEO” and more about creating a reason for another person to reference the site specifically. A useful tool, original dataset, detailed experiment, or genuinely useful comparison seems much more likely to earn a real link than another generic piece of content.
I also really like that you left the incorrect claim visible with a dated correction instead of quietly changing it. That kind of transparency probably matters more for long-term trust than having a perfectly clean-looking archive.
Thanks Dana 🙂 Agreed on giving someone a specific reason to reference the site. That's the part I'm still working on. The dated correction felt awkward to write, but it seemed worse to leave the wrong claim sitting there.
The 332 thing only means something because you opened the list. I usually stop at the count and then I repeat it like it's real. Same with the www redirect claim that was wrong for four weeks. You wrote it, felt done, never hit the URL again. I do that with my own notes way more than I want to admit. Dated correction instead of quietly rewriting it is the part I wouldn't have done.
Yeah, that one stung. I wrote it down as done and never re-tested it, so the claim just sat there for a month.
Two things in your post are more connected than they look.
On the "332 → 1" discovery: this is the single most useful lesson about link metrics. Referring-domain counts in every tool are polluted by exactly this auto-generated junk, and dofollow-vs-nofollow matters far less than whether the linking page has real traffic and editorial intent. Your indiehackers.com links are worth more than the other 327 combined — nofollow has been a hint, not a directive, since 2019, and those links still drive crawl discovery. The metric worth watching in Ahrefs isn't the count, it's how many new referring domains have their own organic traffic.
On the soft-404 catch: that's the bigger deal of the two, and your diagnosis is right. Client-side crawlers only request URLs they can find, so a "200 for everything" misconfiguration is invisible to Screaming Frog-style checks. This matters doubly right now: AI answer-engine crawlers (Perplexity, ChatGPT browsing, AI Overviews' fetchers) decide what to ingest based on clean status codes and well-formed responses. A site full of soft 404s can get quietly deprioritized in AI indexes with zero signal in Search Console. I'd extend your server-side habit into a permanent check: curl your site with a bot user-agent, request a nonexistent path, and diff the status codes against what a browser gets. Discrepancies between what you serve browsers vs. crawlers are where the expensive bugs hide.
On your actual question: "make something people want to cite" is vague advice, but there's a concrete version. Find threads — here, Reddit, niche forums — where someone says "does anyone have a source for that?" and nobody answers well. Publish the definitive, quotable answer, then go reply with it where the question was asked. Those comments are backlink requests wearing a costume: citing you becomes the cheapest way for the other person to sound credible. That's where first real links come from for most people — not outreach, not directories, just answering a question nobody had answered properly.
I did run a version of the bot user-agent check. A spoofed UA from my own machine isn't a real crawler though, so Cloudflare's IP-verified bot logs turned out to be the more useful signal. Do you have a source for AI crawlers deprioritizing sites with soft 404s? I haven't seen one.
First real backlinks I got came from answering a specific question somewhere people already search, with a link only when it actually solved the problem. Directory dumps mostly look like those fake referring domains. One useful page beats volume.
That's the pattern I keep reading about, but I can't picture the "where". Was it Reddit, a forum, Stack Overflow, something niche? I've been answering questions on IH and Reddit for a few weeks without a link and none of it has turned into a backlink yet, so I'm curious what kind of question it was.
The “332 backlinks → actually one” discovery is painful 😅
I think the hardest part with SEO numbers is that the metric can look like progress until you inspect where it actually came from.
For genuine backlinks, I’ve found that being useful in the right communities seems more realistic than chasing backlink numbers directly.
Curious — are you mainly trying to build backlinks for SEO, or are you also looking for referral traffic from the sites linking to you?
Mostly SEO. As of last week's Search Console, most of my pages sit around positions 3-8 for their queries but get almost no impressions, so I'm treating it as a trust problem more than a traffic problem. Referral traffic would be a bonus, but I'd take a real link first.
This comment was deleted 6 hours ago