I've spent the last year staring at link analytics dashboards, ours and everyone else's and I keep coming back to the same complaint. Figured I'd say it out loud and see who pushes back.
We built the dashboard era to be complete, not to be useful.
Every tool in this space, mine included until recently, competes on how many dimensions it can track. Country, city, ISP, timezone, browser, OS, device, referrer type, UTM source/medium/campaign/term/content, it's a lot of data, and honestly, it's impressive. But nobody staring at a dashboard actually wants a 31st column. A marketing manager doesn't need more rows. They need a sentence: "LinkedIn beat Facebook 3× on conversions, mobile UK evenings drove most of it, your email campaign underperformed go fix the landing page."
That's not a dashboard. That's a verdict.
Here's the part I want to be upfront about:
Every major shortener has shipped some kind of AI feature by now. Bitly, Short.io, Rebrandly, Dub, it's table stakes, not a differentiator anymore. I'm not going to sit here and tell you trimy.io is the only one doing this, because it isn't, and somebody in the comments would call that out anyway (fairly).
What I think actually matters isn't "has AI" vs. "doesn't." It's whether the AI stops at describing what happened or commits to telling you what to do about it. Most of what I've poked at from competitors is really just an auto-generated summary of numbers you already saw on the same screen. That's still a report wearing an AI costume. A verdict is different, it tells you the next move, and it's willing to be wrong.
That's the bet I'm making with trimy.io : plain-English narratives that end with a recommendation, not just a description of the chart above it.
Genuinely don't know the answer to this one:
Would you actually trust an AI-generated verdict on your own campaign data? Or do you want to see it, then override it every single time regardless? And if it's the latter, what would change that? Seeing its reasoning? A confidence score? A track record you could go back and audit later?
This framing lands hard — dashboards full of CTR/bounce rarely answer the only question a marketer cares about: "should I keep spending here?"
Curious how you decide the threshold for a "verdict" vs "keep watching." Is it mostly relative (vs cohort/baseline) or absolute (hit a hard ROI bar)? I've found absolute thresholds feel clearer but break when volume is low.
Good question, actually leaning more relative than absolute, comparing against a campaign's own baseline feels safer than a fixed ROI bar and you're right that absolute thresholds fall apart fast at low volume, one lucky click can swing a small sample way past a hard cutoff. Still working out exactly where the line sits though.
Hello, I’m a software engineer interested in building AI-powered SaaS products. I’m here to learn from other founders and developers, exchange ideas, and connect with people working on interesting projects. Nice to meet you!
Hey, nice to meet you too. Good luck with the AI SaaS stuff, always good to have more people building in this space.
We wrestled with this exact question building SocialPost.ai, and the answer from our users was clear: they override the AI until it builds a track record they can audit, then they stop checking. Confidence scores did nothing, a simple log of past calls and outcomes did everything. Ship the verdict with a scoreboard attached and trust takes care of itself.
Same lesson twice now from your side, and it's a good one, a scoreboard beats a confidence score every time. Definitely building this in, thank you so much for sharing that with me.
This resonates. The dashboard era optimized for completeness because completeness is easy to measure internally — "we track 30+ dimensions" sounds better in a sales call
than "we tell you one thing to do next."
The trust question is the hard part. I think it breaks into two stages: (1) do I trust the data is correct (table stakes, everyone assumes yes), and (2) do I trust the
recommendation is the right move given that data. Stage 2 is where most AI features lose people — not because the AI is wrong, but because the founder knows context the AI
doesn't (a rebrand happening next week, a seasonal anomaly, a bug in Tuesday's tracking).
I'd trust a verdict more if it showed its work: "LinkedIn converted 3× better than Facebook. The gap was widest on mobile in UK evenings. Therefore → shift budget to
LinkedIn mobile ads in UK evening slots." If I disagree, I can see exactly which premise I'm pushing back on, rather than just overriding a black box.
Splitting it into "is the data right" versus "is this the right call given the data" is a really useful way to break down trust, and you're right that stage 2 is where founders' hidden context usually beats the model. The example format is exactly what I'm going for too, showing the premise, the pattern, then the recommendation, so if someone disagrees they know which part to push back on instead of just rejecting the whole thing.
"We built the dashboard era to be complete, not to be useful" — that line applies way beyond analytics. Most tools (and most founder decks, honestly) optimize for showing everything instead of deciding something. A verdict requires an opinion, and an opinion is a risk — which is exactly why software keeps shipping 31 columns instead.
That's exactly it, completeness is safe to sell because nobody can argue with "we track 30 dimensions," but a verdict puts you on the hook for being right. That risk is the whole reason most tools stop at the dashboard.
strong agree on the 31st column, but here's the push-back you asked for: the reason nobody has shipped the verdict is that a verdict needs a goal, and the tool doesn't know yours. "linkedin beat facebook 3x on conversions" is only a verdict if conversions are what you're chasing this week. if you're testing top-of-funnel reach, that exact sentence is noise. the scary failure mode is a tool confidently telling you to go fix the landing page when the real problem was the audience. so the version that wins isn't "tool hands you the verdict," it's "tool asks what you're optimizing for once, then runs everything through that one goal." a verdict without your objective is just a more confident wrong answer.
This is a really good push-back, and it's true, "LinkedIn beat Facebook 3x" only means something if conversions are actually what someone's optimizing for that week. Asking the goal first and running everything through that one thing makes way more sense than assuming the metric that matters. Going to think about building that step in before any verdict gets generated.
The "would you trust an AI verdict on your own data" question hit close to home. We had a real version of this today: our transactional email provider hit its sending quota mid-send, and the failover logic had to decide on its own whether to reroute the rest of the batch through a second provider or hold. We let it auto-decide for reversible stuff (retry, reroute) but kept a human check on anything that would have silently dropped or duplicated sends. That's basically your decision-cost point from the other comment, reversible actions get automated trust fast, irreversible ones don't, no matter how good the model's track record looks on paper. I'd trust your verdict engine a lot faster if it exposed which bucket a given recommendation falls into.
Really like this example, the email failover call is a great real-world version of the same problem. Reversible stuff getting auto-decided while anything that could silently drop or duplicate sends gets a human check, that's exactly the right split. Same idea applies to verdicts I think, so labeling which bucket a recommendation falls into is something worth adding.
One angle nobody's mentioned yet: verdict confidence should probably be gated by decision cost, not just statistical confidence. If the recommendation is "pause this $10k/mo campaign," the evidence bar needs to be way higher than "try a different subject line." Same data quality, completely different verdicts. The asymmetry is structural - proving "do X" needs stronger evidence when doing X irreversibly burns time/budget vs when it's reversible. That's why marketers keep overriding: they're measuring against a different confidence threshold than the tool is calibrated for.
This is a sharp way to put it, same data quality needing a completely different bar depending on how costly being wrong is. Probably explains a lot of the override behavior too, people aren't distrusting the data, they're just using a stricter bar for big calls than the tool even knows about. Want to think about tying the evidence bar to the size of the decision, not just how confident the model feels.
"We built the dashboard era to be complete, not to be useful" — that line applies way beyond analytics. Most tools (and most founder decks, honestly) optimize for showing everything instead of deciding something. A verdict requires an opinion, and an opinion is a risk — which is exactly why software keeps shipping 31 columns instead.
The test I'd apply: after someone reads your verdict, do they know what to DO next — kill the campaign, double the spend, rewrite the landing page? If the verdict doesn't change Monday morning's to-do list, it's just a number wearing a costume.
Curious how you'll handle being wrong — a verdict engine earns trust by being right, but it earns loyalty by saying "we called this one wrong, here's why" when it misses.
Good test, actually, does the verdict change what someone does Monday morning, or is it just a number in a nicer outfit. That's basically what I'm aiming for over a stat nobody acts on. On being wrong, I think you're right that owning it out loud builds more trust than being quietly right most of the time, so yeah, planning to build "we got this one wrong, here's why" in on purpose rather than hoping it never happens.
the verdict vs dashboard distinction is the whole thing. been living this the last two weeks on a tiny solo project, more columns felt like rigor but was actually just avoiding the hard part, forming an actual opinion about what's happening.
on trust: for me it wasn't a confidence score, those always feel arbitrary. what actually built trust was checking a verdict against ground truth after the fact, does "here's the real problem" hold up once I go look directly at the database. a couple times the obvious read of the numbers was wrong, traffic was actually fine, the real leak was two steps downstream in something the top-line numbers didn't show at all. being able to go verify that instead of just accepting a summary is what built confidence over time, not a score attached to the claim.
Great story, and it matches what a few others are saying here too, a confidence score never really convinces anyone, actually checking it against the real data does. The "leak was two steps downstream, not where the numbers pointed" part is a good reminder that the obvious read isn't always the right one. Being able to go verify instead of just trusting the summary is probably the whole game, thank you for sharing that with me.
I agree that numbers alone don't tell the whole story. For local businesses, understanding customer intent and actual conversions is often more valuable than tracking every metric. Which KPIs have you found to be the most reliable?
Totally agree, most of the value is in a handful of signals, not all 31 columns. For us it's usually which channel actually converts, not just clicks, and whether a specific campaign is trending up or down versus its own past, not some generic benchmark. currently we still figuring out the full answer to this, curious what's worked on your side too.
"The verdict framing is right, but I think the trust problem is upstream of the reasoning or confidence score. Most marketers have been burned by 'AI insights' that were just pattern-matching on noise — so the default is to override regardless. What might actually build trust is showing the AI being wrong on a low-stakes call first, owning it, and correcting. A tool that demonstrates intellectual honesty early earns the right to be believed later."
This is a really good point, being burned by fake "AI insight" before is exactly why people default to overriding everything. Letting it be visibly wrong on something low-stakes first and owning it is a smart way to earn trust instead of just asking for it upfront. Might actually think about building that in on purpose rather than just hoping it never messes up.
I'd trust the verdict in proportion to the evidence attached, not the reasoning. I build dashboards for non-technical business owners in a different niche, and what made people trust an automated conclusion was never a confidence score — it was being able to click through to the raw data behind the sentence. So: verdict on top, one tap to the underlying rows. A visible track record would do even more. "Last month's verdicts: 7 right, 2 wrong" would convert me faster than any explanation of the model's reasoning, because it's falsifiable.
Evidence over reasoning matches what others are saying here too, verdict on top, one click to see the raw numbers behind it. The track record idea is even better though, "7 right, 2 wrong last month" is something people can actually check, a confidence score just sits there and you have to take it on faith. Might build something like that, thanks for sharing your real experience with me.
I think the real value depends on the cost of making a wrong decision. If an AI tells me one channel is performing better I would trust it only if I can see why it reached that conclusion. A recommendation without context is still a guess.
What would really help marketers is an AI that explains the tradeoffs behind each suggestion and learns from whether the user followed that advice.
Over time the goal should not be better reports. It should be better decisions. I am also building a solution app right now and this made me think more about how AI should guide users instead of simply giving them more information.
The cost of being wrong is a really good filter, small calls probably don't need the why, big ones definitely do. Also hadn't thought much about learning from whether people actually followed the advice, that seems like the real way to earn trust over time instead of just claiming it upfront. Good luck with your app, curious to see how you handle that part.
strongly agree, and the reason almost nobody does it is that a verdict requires being opinionated, which means being WRONG sometimes. a dashboard is never wrong, its just data, so tools hide behind 31 columns to stay safe. the verdict is the value precisely because someone had the guts to make the call. two adds: 1) pair every verdict with its one-line evidence ("LinkedIn beat FB 3x, mostly mobile UK evenings") so people can sanity-check it, a verdict without the why gets distrusted the first time it feels off, and trust is the whole game here. 2) verdicts are also your distribution, nobody forwards a dashboard but people forward "your email campaign is underperforming, fix the landing page." the sentence IS the product and the marketing, the dashboard becomes the proof not the pitch.
That's such a clean way to put it, a dashboard can't really be wrong since it's just data, and that's probably the whole reason nobody wants to commit to a call. Totally agree on pairing every verdict with its one-line evidence too, skip that and one bad miss is enough to lose someone's trust for good. Hadn't thought about the distribution angle either, people forward a sentence, nobody forwards a dashboard. Really good way to put it, thank you.
Interesting perspective! I agree that marketers don't just need more data—they need actionable insights. The shift from reporting metrics to providing clear recommendations could save a lot of time. Looking forward to seeing how this approach performs in real-world campaigns.
Thanks, that's exactly the shift I'm betting on in trimy.io, appreciate you following along, will share how it plays out once more people are using it day to day.
A verdict inherits whatever bias sits in the measurement, and it hides that bias better than a table does. "LinkedIn beat Facebook 3x" might be true, or it might be an artefact of where the tracking survived: links opened in the LinkedIn in-app browser, iOS Safari, anything forwarded in a DM, EU sessions that never got recorded because analytics sat behind a consent banner. None of that lands evenly across channels.
Staring at a table, someone notices the sample looks thin. A confident sentence gives them nothing to notice.
I'd make every verdict carry its coverage, meaning the share of journeys actually attributed rather than total clicks, and refuse to recommend below a threshold. "LinkedIn won, and I can see a third of the journeys" is a different claim from the same sentence at 90%.
Really good point, hadn't thought about it this precisely before, you're right that in-app browsers and DM shares can quietly skew which channel "wins." I like your idea a lot, showing how much of the traffic a verdict is actually based on, and just staying quiet if that's too thin instead of forcing a confident answer anyway. Might actually add that.
If you do add it, coverage won't be uniform across the channels you're comparing, which is the part that trips people up. LinkedIn and Meta traffic loses more of it than email or direct does, so one global number still lets a skewed comparison through. Per-channel denominators, or the same bias just moves down a level.
Cheap way to measure it without building anything: your redirect already counts every click server-side. Compare that with sessions that arrived with the tag intact. The gap is your loss, per channel, for free.
I really like how you've explained that, it makes a lot of sense, our data is set up in a way that this would actually work pretty well. We already record every click on our server, whether or not the UTM tag was included and we also store any UTM info that does come through with each click. So, we could easily look at the difference between total clicks and clicks that still had their tag, without having to build anything new. Thanks for breaking it down like that, it's definitely something we could do, thanks again.
Worth checking what that ratio actually measures before you wire it to a verdict. Your redirect logs the click every time, so on your side the click always exists. The thing that goes missing sits downstream: the session on the destination site, the purchase. Tagged over total tells you how disciplined people are about building their links. It says nothing about how much of the journey you can see.
The coverage that should gate a recommendation is clicks with a matching downstream event. Where those events arrive through someone else's integration, the honest verdict admits it: I can see clicks, I can't see outcomes.
Whatever you do, strip bots and prefetches from the denominator. Link scanners, Slack and Twitter unfurls and mail security proxies all hit the redirect and never arrive anywhere, and they bunch up in email and chat, so those two will look worse than they are.
That's a really important correction, tagged-over-total tells you about link discipline, not visibility into what happened after. Completely agree the real gate should be clicks with a matching downstream event, and being honest that "I can see clicks, not outcomes" is exactly the right call when you don't have that. On the bot/prefetch point, we already filter out known unfurl bots like Slack and Twitter by default, but you're right that mail security scanners like Outlook Safe Links aren't fully covered, so email would look artificially worse right now. Good catch, that's something worth tightening.
Safe Links is the nastier half of that family, along with Proofpoint URL Defense, Mimecast and Barracuda. They fetch from server IPs while presenting an ordinary browser user agent, so filtering on UA alone won't catch them.
Two signals hold up better: the originating ASN, since Microsoft and the security vendors are easy to spot, and a click with no subsequent event from the same IP a few seconds later. There's a third tell in the timing. Scanners fire on delivery rather than when a human opens the mail, so a click sitting suspiciously close to send time is usually not a person.
Really appreciate you breaking this down into this level of details, super helpful.
The trust question gets more interesting when you look at what happens after the verdict. Assume the AI correctly identifies "LinkedIn 3x outperforms Facebook." A marketing manager still faces: "okay, now what? Do I kill Facebook entirely or just shift budget? Do I change the creative or the timing?"
The deeper bottleneck I've seen: marketers usually don't know whether a metric matters because they don't know causality. They can see the outcome but not the input variables that drove it. An AI verdict without that scaffolding reads as a confidence play, not a diagnosis.
Your question about trust might actually be a question about completeness: would they trust the recommendation if it came with the "here's why" layer that shows the causal chain, not just the final call?
This is the sharpest reframe so far, "They can see the outcome but not what drove it" is exactly the gap between a verdict and a useful verdict. I think you're right that it's not really a trust question, it's a completeness question. Next thing I want to try is attaching the "why" (which inputs moved the number) to every recommendation, not just the "what." Appreciate you putting it this way, going to steal "causal chain" as the framing internally.
In our own tool (a software recommendation quiz for small businesses), we found the raw inputs next to the verdict mattered more than a confidence score or explanation ever did. People don't read through a full reasoning trace, but they do a two second gut check on the actual data behind it. If that data is visibly real and checkable, the verdict gets trusted fast. If it is just a number with no visible source, people override it every time, which sounds like exactly what you are running into.
That matches what I'm seeing too, people don't want your reasoning, they want a fast way to sanity-check it themselves. "Two second gut check on the real numbers" is a great way to put it. That's basically the direction I'm leaning: show the verdict, but put the exact numbers right next to it so nobody has to take it on faith.
Building AI recommendations at SocialPost.ai taught us the trust question has a boring answer: nobody trusts the first verdict, everybody trusts a scoreboard. Show each recommendation with its past hit rate ('we made 14 calls like this, 11 improved conversions') and let users grade the verdicts they followed, and the override behavior fades within weeks. The verdict isn't the product, the track record is.
That's a great way to think about it, makes sense too, one good call means nothing, it's a pattern of good calls over time that actually gets people to stop double-checking. Might genuinely build a track record view because of this, thanks for sharing what worked for you.
I like the distinction between a report and a verdict, but I would add one more requirement:
A verdict should not only recommend an action. It should define what the available evidence can actually support.
For example:
“LinkedIn beat Facebook 3× on conversions” may be directly supported by campaign data.
But “your email campaign underperformed because the landing page is weak” is a different type of claim. The data may show lower conversion, but the proposed cause may still be an inference.
I would separate the output into four layers:
Observed fact — what the data directly records.
Comparison — how that result differs from a baseline.
Inference — the most plausible explanation.
Recommended action — the next test or decision.
That distinction matters because users may trust a recommendation more than the evidence justifies.
I would also avoid forcing a verdict when the data is insufficient. “Not enough evidence to determine the cause” is sometimes the most useful conclusion a system can produce.
For me, trust would come from an auditable structure:
the exact claim;
the metrics supporting it;
contradictory evidence;
the baseline used;
the confidence or uncertainty;
the reason the recommendation follows.
A confidence score alone would not be enough. An audit trail would.
The strongest system would not only tell me what to do. It would make clear which parts are measured, which parts are inferred, and what new evidence could prove the recommendation wrong.
Really like this breakdown, it's basically the missing QA layer for AI verdicts, you're right that "underperformed" and "why it underperformed" are two different confidence levels, and blurring them is how trust gets broken. I want to try separating "what the data shows" from "what we think it means" as two distinct lines instead of one blended sentence. And yeah, "not enough data to say" should absolutely be a valid output, not a failure state.
There's a group where the verdict framing lands even harder, people running campaigns for someone else. A freelance marketer or a small agency already has to turn the dashboard into a sentence every week, because the client never wanted 31 columns. They want to know if it was worth the money. For them a verdict isn't replacing analysis, it's replacing the most tedious writing task of their week.
On the trust question, I'd trust it faster in that role than for my own calls. If I'm forwarding a verdict to a client I'll check it against the chart first, and that check is quick when every claim links to the exact numbers behind it. So audit trail over confidence score for me. An 82 percent confident label doesn't tell me what to double check. A citation does.
This is a really good point I hadn't fully considered, the "verdict" isn't just for the person running the campaign, it's for the person who has to explain the campaign to someone else, that reframes the whole value prop, and "citation over confidence score" makes sense; a number like 82% doesn't tell you where to look, a link to the actual data does. Going to think about this for how we surface the "why" behind each verdict.
If the person explaining the campaign turns out to be your real user, the bar I'd design against is whether the verdict survives being pasted into a client email with zero editing. That was my test for every report I wrote. If I had to translate it before forwarding, the tool only did half the job. Short declarative sentences, each one carrying its citation, nothing that sounds like a model hedging.
One thing that might save you some work. My clients almost never clicked through to the numbers. They just needed to see that they could.
That's a great bar to design toward, pasteable without edits, and the last part is useful too, the link matters more as a safety net than something people click, just knowing it's there is enough.
"A dashboard gives you columns. A verdict gives you a sentence." This exact gap exists across most SaaS tools I use. The difference between "here's your data" and "here's what to do" is the difference between a tool and an assistant.
I shipped a Shopify tutorial product recently — my entire analytics stack is basically "how many people clicked the Payhip link." I don't need 31 data columns. I need one sentence: "Medium sent 4 visitors who read the whole article, Twitter sent 30 who bounced in 3 seconds, Quora sent 1 who bought."
Is trimy tackling the "verdict" piece as a structured template (always tells you the same 3 things) or as a freeform insight engine?
love your Shopify example, that's the exact shape I'm going for, to answer directly: it's not a fixed template with the same 3 labeled fields every time, it's a short paragraph the AI has to hit certain notes on: overall performance, the standout trend, and one actionable recommendation, but written in plain sentences, not a schema. We do have one fully structured output elsewhere (a campaign health score with a fixed score/grade/tips format), but the narrative itself stays prose on purpose so it reads like someone telling you what happened, not a form being filled in.
This nails it - the bottleneck was never computing the verdict, it's earning enough trust that people act on it instead of re-checking every row themselves. On what changes that: in my experience it's less a confidence score and more the tool visibly refusing to overreach. The first time it asserts something that isn't in the data, trust is gone and they go back to reading rows.
I hit the same wall from the YouTube analytics side and built a tool (AlgoLens) around three hard rules: never state a number or claim that isn't in the data, always judge against the user's own baseline instead of global benchmarks, and never contradict an earlier answer when they re-ask. The 'show the evidence' part you mentioned matters most - every verdict links back to the exact metric it came from, so it stays checkable.
Free to try if you want to see the pattern (coupon algolensday = 100 credits). How are you handling the 'not enough signal' case in trimy - softer verdict, or stay silent?
"The first time it asserts something not in the data, trust is gone" is exactly right, that's the whole risk with this feature, we handle low-signal cases by having it say so plainly rather than forcing a verdict, a soft "not enough clicks yet to call this" beats a confident guess.
I like the distinction you're making between describing the data and committing to an interpretation of it.
I'll be curious which recommendations users actually act on repeatedly. That usually reveals whether the product is becoming an analytics tool, a decision-support tool, or something in between.
That's the real test, not whether it sounds smart once, but whether people keep taking the same type of recommendation over time without checking it manually, currently I don't have enough usage data yet to answer that properly, but it's the metric I care about most going forward, more than "did they read it" only, thank you for make me see that clearly.
Appreciate the context.
The shift from "reading recommendations" to "acting on them repeatedly" is the interesting part here.
Would be good to discuss what you're learning as usage data comes in.
What's the best email to reach you on?
Appreciate that, happy to keep sharing as usage data comes in. Feel free to DM me here on IndieHackers and we can continue from there.
Appreciate that, Ahmed.
Happy to continue the conversation. I don't think IH has a DM feature though.
If it's easier, what's the best email to reach you on?
Fair point on IH, no DMs here. Easiest way is the contact form on trimy.io, message will come to me and I'll follow up from there.
Thanks! I’ve just sent it over at [email protected]
Looking forward to hearing your thoughts whenever you have a chance.