I've spent the last year staring at link analytics dashboards, ours and everyone else's and I keep coming back to the same complaint. Figured I'd say it out loud and see who pushes back.
We built the dashboard era to be complete, not to be useful.
Every tool in this space, mine included until recently, competes on how many dimensions it can track. Country, city, ISP, timezone, browser, OS, device, referrer type, UTM source/medium/campaign/term/content, it's a lot of data, and honestly, it's impressive. But nobody staring at a dashboard actually wants a 31st column. A marketing manager doesn't need more rows. They need a sentence: "LinkedIn beat Facebook 3× on conversions, mobile UK evenings drove most of it, your email campaign underperformed go fix the landing page."
That's not a dashboard. That's a verdict.
Here's the part I want to be upfront about:
Every major shortener has shipped some kind of AI feature by now. Bitly, Short.io, Rebrandly, Dub, it's table stakes, not a differentiator anymore. I'm not going to sit here and tell you trimy.io is the only one doing this, because it isn't, and somebody in the comments would call that out anyway (fairly).
What I think actually matters isn't "has AI" vs. "doesn't." It's whether the AI stops at describing what happened or commits to telling you what to do about it. Most of what I've poked at from competitors is really just an auto-generated summary of numbers you already saw on the same screen. That's still a report wearing an AI costume. A verdict is different, it tells you the next move, and it's willing to be wrong.
That's the bet I'm making with trimy.io : plain-English narratives that end with a recommendation, not just a description of the chart above it.
Genuinely don't know the answer to this one:
Would you actually trust an AI-generated verdict on your own campaign data? Or do you want to see it, then override it every single time regardless? And if it's the latter, what would change that? Seeing its reasoning? A confidence score? A track record you could go back and audit later?
A verdict inherits whatever bias sits in the measurement, and it hides that bias better than a table does. "LinkedIn beat Facebook 3x" might be true, or it might be an artefact of where the tracking survived: links opened in the LinkedIn in-app browser, iOS Safari, anything forwarded in a DM, EU sessions that never got recorded because analytics sat behind a consent banner. None of that lands evenly across channels.
Staring at a table, someone notices the sample looks thin. A confident sentence gives them nothing to notice.
I'd make every verdict carry its coverage, meaning the share of journeys actually attributed rather than total clicks, and refuse to recommend below a threshold. "LinkedIn won, and I can see a third of the journeys" is a different claim from the same sentence at 90%.
Really good point, hadn't thought about it this precisely before, you're right that in-app browsers and DM shares can quietly skew which channel "wins." I like your idea a lot, showing how much of the traffic a verdict is actually based on, and just staying quiet if that's too thin instead of forcing a confident answer anyway. Might actually add that.
The trust question gets more interesting when you look at what happens after the verdict. Assume the AI correctly identifies "LinkedIn 3x outperforms Facebook." A marketing manager still faces: "okay, now what? Do I kill Facebook entirely or just shift budget? Do I change the creative or the timing?"
The deeper bottleneck I've seen: marketers usually don't know whether a metric matters because they don't know causality. They can see the outcome but not the input variables that drove it. An AI verdict without that scaffolding reads as a confidence play, not a diagnosis.
Your question about trust might actually be a question about completeness: would they trust the recommendation if it came with the "here's why" layer that shows the causal chain, not just the final call?
This is the sharpest reframe so far, "They can see the outcome but not what drove it" is exactly the gap between a verdict and a useful verdict. I think you're right that it's not really a trust question, it's a completeness question. Next thing I want to try is attaching the "why" (which inputs moved the number) to every recommendation, not just the "what." Appreciate you putting it this way, going to steal "causal chain" as the framing internally.
In our own tool (a software recommendation quiz for small businesses), we found the raw inputs next to the verdict mattered more than a confidence score or explanation ever did. People don't read through a full reasoning trace, but they do a two second gut check on the actual data behind it. If that data is visibly real and checkable, the verdict gets trusted fast. If it is just a number with no visible source, people override it every time, which sounds like exactly what you are running into.
That matches what I'm seeing too, people don't want your reasoning, they want a fast way to sanity-check it themselves. "Two second gut check on the real numbers" is a great way to put it. That's basically the direction I'm leaning: show the verdict, but put the exact numbers right next to it so nobody has to take it on faith.
Building AI recommendations at SocialPost.ai taught us the trust question has a boring answer: nobody trusts the first verdict, everybody trusts a scoreboard. Show each recommendation with its past hit rate ('we made 14 calls like this, 11 improved conversions') and let users grade the verdicts they followed, and the override behavior fades within weeks. The verdict isn't the product, the track record is.
That's a great way to think about it, makes sense too, one good call means nothing, it's a pattern of good calls over time that actually gets people to stop double-checking. Might genuinely build a track record view because of this, thanks for sharing what worked for you.
I like the distinction between a report and a verdict, but I would add one more requirement:
A verdict should not only recommend an action. It should define what the available evidence can actually support.
For example:
“LinkedIn beat Facebook 3× on conversions” may be directly supported by campaign data.
But “your email campaign underperformed because the landing page is weak” is a different type of claim. The data may show lower conversion, but the proposed cause may still be an inference.
I would separate the output into four layers:
Observed fact — what the data directly records.
Comparison — how that result differs from a baseline.
Inference — the most plausible explanation.
Recommended action — the next test or decision.
That distinction matters because users may trust a recommendation more than the evidence justifies.
I would also avoid forcing a verdict when the data is insufficient. “Not enough evidence to determine the cause” is sometimes the most useful conclusion a system can produce.
For me, trust would come from an auditable structure:
the exact claim;
the metrics supporting it;
contradictory evidence;
the baseline used;
the confidence or uncertainty;
the reason the recommendation follows.
A confidence score alone would not be enough. An audit trail would.
The strongest system would not only tell me what to do. It would make clear which parts are measured, which parts are inferred, and what new evidence could prove the recommendation wrong.
Really like this breakdown, it's basically the missing QA layer for AI verdicts, you're right that "underperformed" and "why it underperformed" are two different confidence levels, and blurring them is how trust gets broken. I want to try separating "what the data shows" from "what we think it means" as two distinct lines instead of one blended sentence. And yeah, "not enough data to say" should absolutely be a valid output, not a failure state.
There's a group where the verdict framing lands even harder, people running campaigns for someone else. A freelance marketer or a small agency already has to turn the dashboard into a sentence every week, because the client never wanted 31 columns. They want to know if it was worth the money. For them a verdict isn't replacing analysis, it's replacing the most tedious writing task of their week.
On the trust question, I'd trust it faster in that role than for my own calls. If I'm forwarding a verdict to a client I'll check it against the chart first, and that check is quick when every claim links to the exact numbers behind it. So audit trail over confidence score for me. An 82 percent confident label doesn't tell me what to double check. A citation does.
This is a really good point I hadn't fully considered, the "verdict" isn't just for the person running the campaign, it's for the person who has to explain the campaign to someone else, that reframes the whole value prop, and "citation over confidence score" makes sense; a number like 82% doesn't tell you where to look, a link to the actual data does. Going to think about this for how we surface the "why" behind each verdict.
"A dashboard gives you columns. A verdict gives you a sentence." This exact gap exists across most SaaS tools I use. The difference between "here's your data" and "here's what to do" is the difference between a tool and an assistant.
I shipped a Shopify tutorial product recently — my entire analytics stack is basically "how many people clicked the Payhip link." I don't need 31 data columns. I need one sentence: "Medium sent 4 visitors who read the whole article, Twitter sent 30 who bounced in 3 seconds, Quora sent 1 who bought."
Is trimy tackling the "verdict" piece as a structured template (always tells you the same 3 things) or as a freeform insight engine?
love your Shopify example, that's the exact shape I'm going for, to answer directly: it's not a fixed template with the same 3 labeled fields every time, it's a short paragraph the AI has to hit certain notes on: overall performance, the standout trend, and one actionable recommendation, but written in plain sentences, not a schema. We do have one fully structured output elsewhere (a campaign health score with a fixed score/grade/tips format), but the narrative itself stays prose on purpose so it reads like someone telling you what happened, not a form being filled in.
This nails it - the bottleneck was never computing the verdict, it's earning enough trust that people act on it instead of re-checking every row themselves. On what changes that: in my experience it's less a confidence score and more the tool visibly refusing to overreach. The first time it asserts something that isn't in the data, trust is gone and they go back to reading rows.
I hit the same wall from the YouTube analytics side and built a tool (AlgoLens) around three hard rules: never state a number or claim that isn't in the data, always judge against the user's own baseline instead of global benchmarks, and never contradict an earlier answer when they re-ask. The 'show the evidence' part you mentioned matters most - every verdict links back to the exact metric it came from, so it stays checkable.
Free to try if you want to see the pattern (coupon algolensday = 100 credits). How are you handling the 'not enough signal' case in trimy - softer verdict, or stay silent?
"The first time it asserts something not in the data, trust is gone" is exactly right, that's the whole risk with this feature, we handle low-signal cases by having it say so plainly rather than forcing a verdict, a soft "not enough clicks yet to call this" beats a confident guess.
I like the distinction you're making between describing the data and committing to an interpretation of it.
I'll be curious which recommendations users actually act on repeatedly. That usually reveals whether the product is becoming an analytics tool, a decision-support tool, or something in between.
That's the real test, not whether it sounds smart once, but whether people keep taking the same type of recommendation over time without checking it manually, currently I don't have enough usage data yet to answer that properly, but it's the metric I care about most going forward, more than "did they read it" only, thank you for make me see that clearly.