10
44 Comments

Link analytics tools hand you numbers. What marketers actually need is a verdict.

I've spent the last year staring at link analytics dashboards, ours and everyone else's and I keep coming back to the same complaint. Figured I'd say it out loud and see who pushes back.

We built the dashboard era to be complete, not to be useful.

Every tool in this space, mine included until recently, competes on how many dimensions it can track. Country, city, ISP, timezone, browser, OS, device, referrer type, UTM source/medium/campaign/term/content, it's a lot of data, and honestly, it's impressive. But nobody staring at a dashboard actually wants a 31st column. A marketing manager doesn't need more rows. They need a sentence: "LinkedIn beat Facebook 3× on conversions, mobile UK evenings drove most of it, your email campaign underperformed go fix the landing page."

That's not a dashboard. That's a verdict.

Here's the part I want to be upfront about:

Every major shortener has shipped some kind of AI feature by now. Bitly, Short.io, Rebrandly, Dub, it's table stakes, not a differentiator anymore. I'm not going to sit here and tell you trimy.io is the only one doing this, because it isn't, and somebody in the comments would call that out anyway (fairly).

What I think actually matters isn't "has AI" vs. "doesn't." It's whether the AI stops at describing what happened or commits to telling you what to do about it. Most of what I've poked at from competitors is really just an auto-generated summary of numbers you already saw on the same screen. That's still a report wearing an AI costume. A verdict is different, it tells you the next move, and it's willing to be wrong.

That's the bet I'm making with trimy.io : plain-English narratives that end with a recommendation, not just a description of the chart above it.

Genuinely don't know the answer to this one:

Would you actually trust an AI-generated verdict on your own campaign data? Or do you want to see it, then override it every single time regardless? And if it's the latter, what would change that? Seeing its reasoning? A confidence score? A track record you could go back and audit later?

posted to Icon for group AI Tools
AI Tools
on July 20, 2026
  1. 2

    The "would you trust an AI verdict on your own data" question hit close to home. We had a real version of this today: our transactional email provider hit its sending quota mid-send, and the failover logic had to decide on its own whether to reroute the rest of the batch through a second provider or hold. We let it auto-decide for reversible stuff (retry, reroute) but kept a human check on anything that would have silently dropped or duplicated sends. That's basically your decision-cost point from the other comment, reversible actions get automated trust fast, irreversible ones don't, no matter how good the model's track record looks on paper. I'd trust your verdict engine a lot faster if it exposed which bucket a given recommendation falls into.

    1. 1

      Really like this example, the email failover call is a great real-world version of the same problem. Reversible stuff getting auto-decided while anything that could silently drop or duplicate sends gets a human check, that's exactly the right split. Same idea applies to verdicts I think, so labeling which bucket a recommendation falls into is something worth adding.

  2. 2

    One angle nobody's mentioned yet: verdict confidence should probably be gated by decision cost, not just statistical confidence. If the recommendation is "pause this $10k/mo campaign," the evidence bar needs to be way higher than "try a different subject line." Same data quality, completely different verdicts. The asymmetry is structural - proving "do X" needs stronger evidence when doing X irreversibly burns time/budget vs when it's reversible. That's why marketers keep overriding: they're measuring against a different confidence threshold than the tool is calibrated for.

    1. 1

      This is a sharp way to put it, same data quality needing a completely different bar depending on how costly being wrong is. Probably explains a lot of the override behavior too, people aren't distrusting the data, they're just using a stricter bar for big calls than the tool even knows about. Want to think about tying the evidence bar to the size of the decision, not just how confident the model feels.

  3. 2

    "We built the dashboard era to be complete, not to be useful" — that line applies way beyond analytics. Most tools (and most founder decks, honestly) optimize for showing everything instead of deciding something. A verdict requires an opinion, and an opinion is a risk — which is exactly why software keeps shipping 31 columns instead.

    The test I'd apply: after someone reads your verdict, do they know what to DO next — kill the campaign, double the spend, rewrite the landing page? If the verdict doesn't change Monday morning's to-do list, it's just a number wearing a costume.

    Curious how you'll handle being wrong — a verdict engine earns trust by being right, but it earns loyalty by saying "we called this one wrong, here's why" when it misses.

    1. 1

      Good test, actually, does the verdict change what someone does Monday morning, or is it just a number in a nicer outfit. That's basically what I'm aiming for over a stat nobody acts on. On being wrong, I think you're right that owning it out loud builds more trust than being quietly right most of the time, so yeah, planning to build "we got this one wrong, here's why" in on purpose rather than hoping it never happens.

  4. 2

    the verdict vs dashboard distinction is the whole thing. been living this the last two weeks on a tiny solo project, more columns felt like rigor but was actually just avoiding the hard part, forming an actual opinion about what's happening.

    on trust: for me it wasn't a confidence score, those always feel arbitrary. what actually built trust was checking a verdict against ground truth after the fact, does "here's the real problem" hold up once I go look directly at the database. a couple times the obvious read of the numbers was wrong, traffic was actually fine, the real leak was two steps downstream in something the top-line numbers didn't show at all. being able to go verify that instead of just accepting a summary is what built confidence over time, not a score attached to the claim.

    1. 1

      Great story, and it matches what a few others are saying here too, a confidence score never really convinces anyone, actually checking it against the real data does. The "leak was two steps downstream, not where the numbers pointed" part is a good reminder that the obvious read isn't always the right one. Being able to go verify instead of just trusting the summary is probably the whole game, thank you for sharing that with me.

  5. 2

    I agree that numbers alone don't tell the whole story. For local businesses, understanding customer intent and actual conversions is often more valuable than tracking every metric. Which KPIs have you found to be the most reliable?

    1. 1

      Totally agree, most of the value is in a handful of signals, not all 31 columns. For us it's usually which channel actually converts, not just clicks, and whether a specific campaign is trending up or down versus its own past, not some generic benchmark. currently we still figuring out the full answer to this, curious what's worked on your side too.

  6. 2

    "The verdict framing is right, but I think the trust problem is upstream of the reasoning or confidence score. Most marketers have been burned by 'AI insights' that were just pattern-matching on noise — so the default is to override regardless. What might actually build trust is showing the AI being wrong on a low-stakes call first, owning it, and correcting. A tool that demonstrates intellectual honesty early earns the right to be believed later."

    1. 1

      This is a really good point, being burned by fake "AI insight" before is exactly why people default to overriding everything. Letting it be visibly wrong on something low-stakes first and owning it is a smart way to earn trust instead of just asking for it upfront. Might actually think about building that in on purpose rather than just hoping it never messes up.

  7. 2

    I'd trust the verdict in proportion to the evidence attached, not the reasoning. I build dashboards for non-technical business owners in a different niche, and what made people trust an automated conclusion was never a confidence score — it was being able to click through to the raw data behind the sentence. So: verdict on top, one tap to the underlying rows. A visible track record would do even more. "Last month's verdicts: 7 right, 2 wrong" would convert me faster than any explanation of the model's reasoning, because it's falsifiable.

    1. 1

      Evidence over reasoning matches what others are saying here too, verdict on top, one click to see the raw numbers behind it. The track record idea is even better though, "7 right, 2 wrong last month" is something people can actually check, a confidence score just sits there and you have to take it on faith. Might build something like that, thanks for sharing your real experience with me.

  8. 2

    I think the real value depends on the cost of making a wrong decision. If an AI tells me one channel is performing better I would trust it only if I can see why it reached that conclusion. A recommendation without context is still a guess.

    What would really help marketers is an AI that explains the tradeoffs behind each suggestion and learns from whether the user followed that advice.

    Over time the goal should not be better reports. It should be better decisions. I am also building a solution app right now and this made me think more about how AI should guide users instead of simply giving them more information.

    1. 1

      The cost of being wrong is a really good filter, small calls probably don't need the why, big ones definitely do. Also hadn't thought much about learning from whether people actually followed the advice, that seems like the real way to earn trust over time instead of just claiming it upfront. Good luck with your app, curious to see how you handle that part.

  9. 2

    strongly agree, and the reason almost nobody does it is that a verdict requires being opinionated, which means being WRONG sometimes. a dashboard is never wrong, its just data, so tools hide behind 31 columns to stay safe. the verdict is the value precisely because someone had the guts to make the call. two adds: 1) pair every verdict with its one-line evidence ("LinkedIn beat FB 3x, mostly mobile UK evenings") so people can sanity-check it, a verdict without the why gets distrusted the first time it feels off, and trust is the whole game here. 2) verdicts are also your distribution, nobody forwards a dashboard but people forward "your email campaign is underperforming, fix the landing page." the sentence IS the product and the marketing, the dashboard becomes the proof not the pitch.

    1. 1

      That's such a clean way to put it, a dashboard can't really be wrong since it's just data, and that's probably the whole reason nobody wants to commit to a call. Totally agree on pairing every verdict with its one-line evidence too, skip that and one bad miss is enough to lose someone's trust for good. Hadn't thought about the distribution angle either, people forward a sentence, nobody forwards a dashboard. Really good way to put it, thank you.

  10. 2

    Interesting perspective! I agree that marketers don't just need more data—they need actionable insights. The shift from reporting metrics to providing clear recommendations could save a lot of time. Looking forward to seeing how this approach performs in real-world campaigns.

    1. 1

      Thanks, that's exactly the shift I'm betting on in trimy.io, appreciate you following along, will share how it plays out once more people are using it day to day.

  11. 2

    A verdict inherits whatever bias sits in the measurement, and it hides that bias better than a table does. "LinkedIn beat Facebook 3x" might be true, or it might be an artefact of where the tracking survived: links opened in the LinkedIn in-app browser, iOS Safari, anything forwarded in a DM, EU sessions that never got recorded because analytics sat behind a consent banner. None of that lands evenly across channels.

    Staring at a table, someone notices the sample looks thin. A confident sentence gives them nothing to notice.

    I'd make every verdict carry its coverage, meaning the share of journeys actually attributed rather than total clicks, and refuse to recommend below a threshold. "LinkedIn won, and I can see a third of the journeys" is a different claim from the same sentence at 90%.

    1. 1

      Really good point, hadn't thought about it this precisely before, you're right that in-app browsers and DM shares can quietly skew which channel "wins." I like your idea a lot, showing how much of the traffic a verdict is actually based on, and just staying quiet if that's too thin instead of forcing a confident answer anyway. Might actually add that.

      1. 2

        If you do add it, coverage won't be uniform across the channels you're comparing, which is the part that trips people up. LinkedIn and Meta traffic loses more of it than email or direct does, so one global number still lets a skewed comparison through. Per-channel denominators, or the same bias just moves down a level.

        Cheap way to measure it without building anything: your redirect already counts every click server-side. Compare that with sessions that arrived with the tag intact. The gap is your loss, per channel, for free.

        1. 1

          I really like how you've explained that, it makes a lot of sense, our data is set up in a way that this would actually work pretty well. We already record every click on our server, whether or not the UTM tag was included and we also store any UTM info that does come through with each click. So, we could easily look at the difference between total clicks and clicks that still had their tag, without having to build anything new. Thanks for breaking it down like that, it's definitely something we could do, thanks again.

          1. 2

            Worth checking what that ratio actually measures before you wire it to a verdict. Your redirect logs the click every time, so on your side the click always exists. The thing that goes missing sits downstream: the session on the destination site, the purchase. Tagged over total tells you how disciplined people are about building their links. It says nothing about how much of the journey you can see.

            The coverage that should gate a recommendation is clicks with a matching downstream event. Where those events arrive through someone else's integration, the honest verdict admits it: I can see clicks, I can't see outcomes.

            Whatever you do, strip bots and prefetches from the denominator. Link scanners, Slack and Twitter unfurls and mail security proxies all hit the redirect and never arrive anywhere, and they bunch up in email and chat, so those two will look worse than they are.

            1. 1

              That's a really important correction, tagged-over-total tells you about link discipline, not visibility into what happened after. Completely agree the real gate should be clicks with a matching downstream event, and being honest that "I can see clicks, not outcomes" is exactly the right call when you don't have that. On the bot/prefetch point, we already filter out known unfurl bots like Slack and Twitter by default, but you're right that mail security scanners like Outlook Safe Links aren't fully covered, so email would look artificially worse right now. Good catch, that's something worth tightening.

              1. 2

                Safe Links is the nastier half of that family, along with Proofpoint URL Defense, Mimecast and Barracuda. They fetch from server IPs while presenting an ordinary browser user agent, so filtering on UA alone won't catch them.

                Two signals hold up better: the originating ASN, since Microsoft and the security vendors are easy to spot, and a click with no subsequent event from the same IP a few seconds later. There's a third tell in the timing. Scanners fire on delivery rather than when a human opens the mail, so a click sitting suspiciously close to send time is usually not a person.

                1. 1

                  Really appreciate you breaking this down into this level of details, super helpful.

  12. 2

    The trust question gets more interesting when you look at what happens after the verdict. Assume the AI correctly identifies "LinkedIn 3x outperforms Facebook." A marketing manager still faces: "okay, now what? Do I kill Facebook entirely or just shift budget? Do I change the creative or the timing?"

    The deeper bottleneck I've seen: marketers usually don't know whether a metric matters because they don't know causality. They can see the outcome but not the input variables that drove it. An AI verdict without that scaffolding reads as a confidence play, not a diagnosis.

    Your question about trust might actually be a question about completeness: would they trust the recommendation if it came with the "here's why" layer that shows the causal chain, not just the final call?

    1. 1

      This is the sharpest reframe so far, "They can see the outcome but not what drove it" is exactly the gap between a verdict and a useful verdict. I think you're right that it's not really a trust question, it's a completeness question. Next thing I want to try is attaching the "why" (which inputs moved the number) to every recommendation, not just the "what." Appreciate you putting it this way, going to steal "causal chain" as the framing internally.

  13. 2

    In our own tool (a software recommendation quiz for small businesses), we found the raw inputs next to the verdict mattered more than a confidence score or explanation ever did. People don't read through a full reasoning trace, but they do a two second gut check on the actual data behind it. If that data is visibly real and checkable, the verdict gets trusted fast. If it is just a number with no visible source, people override it every time, which sounds like exactly what you are running into.

    1. 1

      That matches what I'm seeing too, people don't want your reasoning, they want a fast way to sanity-check it themselves. "Two second gut check on the real numbers" is a great way to put it. That's basically the direction I'm leaning: show the verdict, but put the exact numbers right next to it so nobody has to take it on faith.

  14. 2

    Building AI recommendations at SocialPost.ai taught us the trust question has a boring answer: nobody trusts the first verdict, everybody trusts a scoreboard. Show each recommendation with its past hit rate ('we made 14 calls like this, 11 improved conversions') and let users grade the verdicts they followed, and the override behavior fades within weeks. The verdict isn't the product, the track record is.

    1. 1

      That's a great way to think about it, makes sense too, one good call means nothing, it's a pattern of good calls over time that actually gets people to stop double-checking. Might genuinely build a track record view because of this, thanks for sharing what worked for you.

  15. 2

    I like the distinction between a report and a verdict, but I would add one more requirement:

    A verdict should not only recommend an action. It should define what the available evidence can actually support.

    For example:

    “LinkedIn beat Facebook 3× on conversions” may be directly supported by campaign data.
    But “your email campaign underperformed because the landing page is weak” is a different type of claim. The data may show lower conversion, but the proposed cause may still be an inference.
    I would separate the output into four layers:
    Observed fact — what the data directly records.
    Comparison — how that result differs from a baseline.
    Inference — the most plausible explanation.
    Recommended action — the next test or decision.
    That distinction matters because users may trust a recommendation more than the evidence justifies.
    I would also avoid forcing a verdict when the data is insufficient. “Not enough evidence to determine the cause” is sometimes the most useful conclusion a system can produce.
    For me, trust would come from an auditable structure:
    the exact claim;
    the metrics supporting it;
    contradictory evidence;
    the baseline used;
    the confidence or uncertainty;
    the reason the recommendation follows.
    A confidence score alone would not be enough. An audit trail would.
    The strongest system would not only tell me what to do. It would make clear which parts are measured, which parts are inferred, and what new evidence could prove the recommendation wrong.

    1. 1

      Really like this breakdown, it's basically the missing QA layer for AI verdicts, you're right that "underperformed" and "why it underperformed" are two different confidence levels, and blurring them is how trust gets broken. I want to try separating "what the data shows" from "what we think it means" as two distinct lines instead of one blended sentence. And yeah, "not enough data to say" should absolutely be a valid output, not a failure state.

  16. 2

    There's a group where the verdict framing lands even harder, people running campaigns for someone else. A freelance marketer or a small agency already has to turn the dashboard into a sentence every week, because the client never wanted 31 columns. They want to know if it was worth the money. For them a verdict isn't replacing analysis, it's replacing the most tedious writing task of their week.

    On the trust question, I'd trust it faster in that role than for my own calls. If I'm forwarding a verdict to a client I'll check it against the chart first, and that check is quick when every claim links to the exact numbers behind it. So audit trail over confidence score for me. An 82 percent confident label doesn't tell me what to double check. A citation does.

    1. 1

      This is a really good point I hadn't fully considered, the "verdict" isn't just for the person running the campaign, it's for the person who has to explain the campaign to someone else, that reframes the whole value prop, and "citation over confidence score" makes sense; a number like 82% doesn't tell you where to look, a link to the actual data does. Going to think about this for how we surface the "why" behind each verdict.

  17. 2

    "A dashboard gives you columns. A verdict gives you a sentence." This exact gap exists across most SaaS tools I use. The difference between "here's your data" and "here's what to do" is the difference between a tool and an assistant.

    I shipped a Shopify tutorial product recently — my entire analytics stack is basically "how many people clicked the Payhip link." I don't need 31 data columns. I need one sentence: "Medium sent 4 visitors who read the whole article, Twitter sent 30 who bounced in 3 seconds, Quora sent 1 who bought."

    Is trimy tackling the "verdict" piece as a structured template (always tells you the same 3 things) or as a freeform insight engine?

    1. 1

      love your Shopify example, that's the exact shape I'm going for, to answer directly: it's not a fixed template with the same 3 labeled fields every time, it's a short paragraph the AI has to hit certain notes on: overall performance, the standout trend, and one actionable recommendation, but written in plain sentences, not a schema. We do have one fully structured output elsewhere (a campaign health score with a fixed score/grade/tips format), but the narrative itself stays prose on purpose so it reads like someone telling you what happened, not a form being filled in.

  18. 2

    This nails it - the bottleneck was never computing the verdict, it's earning enough trust that people act on it instead of re-checking every row themselves. On what changes that: in my experience it's less a confidence score and more the tool visibly refusing to overreach. The first time it asserts something that isn't in the data, trust is gone and they go back to reading rows.

    I hit the same wall from the YouTube analytics side and built a tool (AlgoLens) around three hard rules: never state a number or claim that isn't in the data, always judge against the user's own baseline instead of global benchmarks, and never contradict an earlier answer when they re-ask. The 'show the evidence' part you mentioned matters most - every verdict links back to the exact metric it came from, so it stays checkable.

    Free to try if you want to see the pattern (coupon algolensday = 100 credits). How are you handling the 'not enough signal' case in trimy - softer verdict, or stay silent?

    1. 1

      "The first time it asserts something not in the data, trust is gone" is exactly right, that's the whole risk with this feature, we handle low-signal cases by having it say so plainly rather than forcing a verdict, a soft "not enough clicks yet to call this" beats a confident guess.

  19. 2

    I like the distinction you're making between describing the data and committing to an interpretation of it.

    I'll be curious which recommendations users actually act on repeatedly. That usually reveals whether the product is becoming an analytics tool, a decision-support tool, or something in between.

    1. 1

      That's the real test, not whether it sounds smart once, but whether people keep taking the same type of recommendation over time without checking it manually, currently I don't have enough usage data yet to answer that properly, but it's the metric I care about most going forward, more than "did they read it" only, thank you for make me see that clearly.

Trending on Indie Hackers
I built a web-based vector editor from scratch and integrated an AI Agent. Need just ONE beta tester! User Avatar 65 comments I built an AI that turns an idea into a live business in under 10 minutes. Here’s what 1,000 launches taught me User Avatar 41 comments Building Noodle, a keyboard-first REST client for the terminal User Avatar 27 comments "Looks Good to Me" Is Quietly Killing Your Feedback Loop User Avatar 26 comments Launched 580 landing pages in 1 week. Solo. No team. User Avatar 18 comments I built a competitor monitor for indie founders User Avatar 8 comments