I started GameVault Deals because I kept seeing the same two frustrations: game currency/account marketplaces full of scams with zero safety guidance, and gaming-gear "best of" content that's almost always sponsored rankings with numbers nobody can verify.
So I set myself one hard rule from day one: nothing fabricated. No invented benchmarks, no fake "top customer" logos, no made-up review scores, no rigged influencer polls. In practice this has actually cost me things — I've turned down easy wins (a directory site literally asked me to nominate "top products," and I had to decline because I'd never actually used them) because faking it would've been faster.
The tradeoff: content takes longer to write when every claim has to trace back to a real spec sheet or formula, and some SEO tactics that "everyone does" (fabricated stats, invented testimonials) are just off the table. But I think it's the only way a comparison site is actually worth trusting.
Curious if anyone else here has deliberately made their own growth harder by refusing to cut a corner that "everyone else" cuts — and whether it paid off long-term or you regretted it.
One way to make the evidence standard perceptible is to turn it into a product feature rather than relying on readers to infer it from the writing.
For each important claim, you could expose a lightweight provenance layer:
source and retrieval date
whether the value is manufacturer-reported or independently measured
the formula used for calculated scores
what you could not verify
a visible correction history when something changes
That would give readers something concrete to compare with sites that simply present confident conclusions.
I’d also separate “trust engagement” from normal content engagement. Expanding a source, checking the methodology, or returning after a correction may be stronger signals than page depth alone.
The interesting test is whether exposing uncertainty reduces immediate clicks but increases repeat visits, citations, and higher-intent outbound clicks over time.
Have you considered giving every comparison a visible evidence-completeness score rather than a conventional editorial rating?
This is a sharp reframe — treating the evidence standard as a literal, inspectable feature rather than something readers have to infer from tone.
Being fully honest about where I'd actually land on each piece: "manufacturer-reported vs. independently measured" would currently just say "manufacturer-reported" on every single spec, since I don't have independent testing capability — no lab, no benchmark rig. That's worth stating outright rather than implying otherwise. A retrieval date and a visible "what I couldn't verify" section are both things I could add now with basically zero new infrastructure — it's just discipline in the writing, made visible instead of implicit.
Correction history is the one that actually requires new plumbing (a changelog per page), but it's the part I like most, because it's the hardest to fake — a site that's never corrected anything is either perfect or not being watched closely.
On the completeness score vs. editorial rating: I never added a star rating precisely because I didn't want to imply a tested opinion I don't have. An evidence-completeness score is actually more honest than what I have now, which is nothing — right now a reader has no structured way to see "this comparison is missing X" at a glance.
The click-vs-repeat-visit tradeoff you're describing is the right experiment, but I don't have the traffic yet to read it cleanly — too early to separate signal from noise. I might build the provenance layer anyway, before I can measure its effect, just because it's a better description of what the site already tries to do.
What made you land on "evidence-completeness score" specifically, rather than a binary verified/unverified tag per claim?
I landed on completeness rather than a binary verified/unverified label because most comparisons are not cleanly one or the other.
A claim can be sourced but still incomplete. For example, the manufacturer may publish weight and battery life, but not explain the test conditions, regional model differences, or when the specification was last updated. Calling that simply “verified” would hide useful uncertainty, while “unverified” would be too harsh.
A completeness view can show how much of the evidence needed for a confident comparison is actually present:
source available
retrieval date recorded
measurement method known
independent confirmation available
conflicting values resolved
important unknowns disclosed
I wouldn’t rely on one score alone, though. A number such as 72% can create false precision unless readers can see what produced it. I’d use the claim-level labels as the underlying evidence, then treat the completeness score as a summary of what is present and what is still missing.
That also makes the score actionable: a comparison improves when a specific evidence gap is filled, not because an editor changed an opinion.
This maps cleanly onto what I've already got, and it clarifies something I hadn't separated out: of your six criteria, three are currently constants across my entire catalog, not per-product variables — source available (always yes, it's always the Amazon listing), retrieval date recorded (now always yes), and independent confirmation available (always no, I don't have lab/benchmark capability). Scoring those per-product would imply variance that isn't real; it's a structural fact about how a one-person site sources data, not a finding specific to any one comparison.
The three that actually vary and are worth tracking per-claim are: measurement method known (manufacturers publish a battery-life number but rarely the test conditions — genuinely inconsistent across products), conflicting values resolved (I only pull from a single source per product right now, so there's no cross-reference to reconcile — "no conflicts" would currently mean "didn't check," which I don't want to misrepresent as "checked and consistent"), and important unknowns disclosed, which I already have a field for.
So the honest version of your framework, applied to my actual constraints: state the two universal limitations once as a standing disclosure rather than repeating them as if they were a per-product finding, and reserve the completeness view for the things that genuinely differ comparison to comparison.
Does that distinction — structural limitation vs. per-claim finding — match how you'd draw the line, or would you still score the structural ones per-product for consistency?
Yes — structural limitation versus per-claim finding is exactly how I’d draw the line.
I wouldn’t score the structural items as though they were product-specific variables, because that would create artificial differences between comparisons. I’d separate the model into two layers:
A catalog-wide methodology baseline:
primary source is the Amazon listing
no independent lab or benchmark testing
retrieval dates are recorded
current cross-referencing policy
A comparison-specific completeness view:
test or measurement conditions are known
conflicting values were actually checked and resolved
important unknowns are disclosed
any product-specific source limitations
The one caution is not to let removing the structural criteria make the per-product score look more complete than the evidence really is. A comparison could score highly on the variable criteria while still relying entirely on manufacturer-reported data.
So I’d state the structural limitations once in the methodology, but also surface a compact inherited label on each comparison — something like “Manufacturer-reported data; not independently tested.” It would remain visible without pretending it is a variable finding.
Then the completeness score can honestly answer a narrower question: given this site’s stated sourcing model, how completely and transparently was this particular comparison documented?
This is the right shape — the methodology-baseline / per-comparison split resolves exactly the risk I was worried about: removing structural noise making individual comparisons look more rigorously verified than they actually are.
I'm going to build this as: a standing methodology statement (source, no independent testing, retrieval-date policy, cross-referencing policy — stated once, linked from every comparison rather than repeated as if it varies), a compact inherited label on each comparison page itself ("Manufacturer-reported data; not independently tested"), and then a completeness view scoped to just the four things that actually differ: measurement conditions known, conflicts checked, unknowns disclosed, product-specific source limitations.
One design question I haven't resolved: whether the completeness view should surface as a single rolled-up count ("3 of 4 documented") or as the four labels shown individually with no combined number at all. A count is easier to scan across many comparisons, but showing all four unrolled is more resistant to someone treating a partial count as an overall reliability signal — which is closer to what your caution is protecting against. I'm leaning toward showing the four labels first, count second. Which would you pick?
I’d make the four labels primary and the count secondary.
The labels explain what is actually present or missing, while a rolled-up “3 of 4” can easily be misread as an overall reliability score rather than documentation coverage.
A practical split could be:
comparison pages: show all four labels prominently
catalog or comparison lists: show a compact “Documentation coverage: 3 of 4” summary for scanning
If you include the count on the detail page, I’d keep it visually subordinate and explicitly call it documentation coverage — not evidence quality, confidence, or reliability.
I’d also avoid a percentage or a green/red grade, since those would imply that all four dimensions carry equal weight and that a higher number means the underlying product claims are more trustworthy.
So yes: labels first, count second. The count helps navigation; the labels carry the meaning.
Built it this way — labels primary on detail pages, compact count only in the list view, explicitly named "documentation coverage," no percentage or grade.
One wrinkle your split surfaced: a comparison card covers two products whose coverage can differ, so flattening it to one number would produce a figure true of neither. It shows a range there instead ("0–2 of 4") and the exact count per product on the detail page.
Everything reads 0 of 4 right now, which is the honest current state — the fields exist but I haven't filled them yet. Felt more useful to ship the structure showing zero than to backfill something thin just to make the number move.
That sounds like exactly the right implementation.
Using a range on the comparison card avoids creating a single number that accurately describes neither product, while the detail page can preserve the exact per-product meaning.
Shipping the structure at 0 of 4 is also the more credible choice. It gives you an honest baseline, and any future movement will reflect real documentation work rather than thin backfilling just to improve the display.
Nice job turning the distinction into something concrete so quickly.
Appreciate you working through it with me — the two-layer split is the part I wouldn't have landed on alone, and it's what makes the count defensible rather than decorative.
The real next step is filling those four fields where the evidence actually supports it, which means going back through listings rather than editing the display. Slower, but it's the only version where movement in the number means anything.
The harder part seems to be whether readers can actually perceive the difference.
Have you seen any behavior yet — repeat visits, clicks, comments, conversions — suggesting people value the evidence standard itself, rather than simply consuming the comparison like they would on any other site?
I appreciate you being candid about where things actually stand.
I'd be interested in continuing the conversation by email if you're open to it. What's the best email to reach you on?
Glad to continue there — reach me at contact [at] gamevaultdeals [dot] shop
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.