7
23 Comments

We scored 237 apps in one Shopify category because the star ratings were useless

I spend a lot of time crawling the Shopify App Store, and this is the clearest case I've run into of ratings just not doing their job.

Picked one category — returns and exchanges apps, 237 live listings — and looked at what the storefront actually shows shoppers. Of the 111 apps that have any reviews at all, 101 sit at 4.5 stars or higher. When almost the whole category clusters at the top, the rating stops telling you anything about which one to pick.

The other 126 of the 237 have never been reviewed at all — not low-rated, just never reviewed. And five apps hold 54% of the 10,345 reviews the whole category has earned, so the review count is as lopsided as the star rating is compressed.

I re-read every shortlisted listing live on 21 September 2026 rather than trust a stale crawl, so the pricing below is current: entry prices in the category run from $4.99 to $179. And of the negative reviews that do exist, 57% are about support, not the product itself — worth knowing if you're trying to judge a tool by its complaints.

The caveat: this is one category, one snapshot, one day. I don't know yet whether returns apps are unusually skewed or whether any mature Shopify category looks like this once a handful of players have spent years compounding reviews.

If you're shipping into a crowded category with zero reviews yet, genuine question — does that actually cost you conversions, or only once a buyer has something to compare you against?

(I work on BestAppify, that's where the numbers come from)

https://merchants.bestappify.com/blog/best-shopify-apps-to-reduce-returns-52

on September 25, 2026
  1. 1

    This is a super clean breakdown. What's crazy to me is that 126 out of 237 apps have zero reviews—that's more than half the category sitting completely invisible.

    As someone building on the dev side, it really shows how much of a barrier distribution is once a few legacy apps compound thousands of reviews. The support complaint stat (57%) is a huge takeaway too—basically proves that after a certain point, a good chunk of your product's "quality" in the buyer's eyes is just how fast you reply when things break. Great data!

  2. 1

    When star ratings compress, the signals that still discriminate are operational ones — support speed, pricing clarity, and especially update cadence. A changelog or "last updated" trail is one of the few public surfaces that doesn't cluster the way stars do: shipping weekly for two years vs three quiet months is hard to fake on a listing page.

    For the 126 never-reviewed apps, I'd treat recent, specific release notes as a trust proxy until reviews exist. Merchants can't install every option, but they can skim whether the vendor is still alive and what they're actually shipping.

    Curious — in your scoring, how much weight did update cadence get relative to pricing-page completeness, and did the high-cadence apps tend to be the ones still getting installs despite low review count?

  3. 1

    The 57%-support-complaints number is the most valuable thing in this post. When ratings compress to 4.5 across a category, the differentiator moves to structured data - support responsiveness, real pricing, update cadence. I build directory software, and this is exactly the gap directories exist to fill: a flat star rating can't tell a buyer "these five answer tickets within a day," but a directory field can. On your zero-reviews question: it costs conversions when review count is the only trust proxy on the page. Give buyers a second axis to compare on and the zero-review app with complete data beats the 4.8-star mystery box for anyone doing homework. For the 126 never-reviewed apps: did anything besides review count separate the ones that still get installs - screenshots, pricing page, vendor site quality?

    1. 1

      Good push, and honestly we don't have an answer for that specific one — we don't track installs, and Shopify doesn't expose them, so we can't say what separates a zero-review app that still gets installed from one that doesn't. What we do have from the same crawl: in that category entry prices run $4.99 to $179, and 126 of 237 apps have never been reviewed. Price alone didn't obviously predict which of those made our shortlist — a few of the pricier ones scored well on things like update cadence and complete pricing pages, which is closer to the structured-data axis you're describing. But that's a correlation we noticed, not proof it moves installs. Would need traffic data we don't have to say more.

  4. 1

    This is great work — what's the biggest thing you'd do differently if you started over?

    1. 1

      If I'm honest, it'd be catching a data quirk earlier: in our crawl an app with zero reviews stores null, not 0. A naive query for number_of_ratings = 0 returns nothing — which would have quietly told us there were no never-reviewed apps at all, when in this category there were 126 of them. We caught it, but only after some manual sanity-checking that should've been automated from day one.

  5. 1

    Have you seen any evidence that merchants actually choose differently when given deeper comparison data, or is the rating problem mainly obvious analytically but weak as a buying trigger?

    1. 1

      We don't have that — no purchase-funnel or conversion data, just listing and review data pulled from the crawl. So I can't tell you whether deeper comparison data actually changes what a merchant installs. What we do have is the diagnostic side: in this category 101 of the 111 apps with any reviews at all show 4.5 stars or better, and the median rated app has seven reviews. That tells you the rating alone carries almost no signal — but whether merchants notice that and behave differently, or just default to whatever's on top of search anyway, is a real open question we can't answer from our data.

      1. 1

        That gap between a useful diagnostic and actual merchant behavior is the part I’d want to dig into. Could be useful to compare notes over email sometime, if you’re open to it.

  6. 1

    Good write-up. What would you do differently if you started again?

    1. 1

      Honestly, I'd lock the live re-check in earlier. We scored the 237 apps from our own crawl, then went back and manually re-read every shortlisted listing on the live App Store on 21 September to get today's prices and ratings — that second pass is what caught the stuff our crawl was already a few days stale on. Next time I'd build that live re-check into the pipeline from day one instead of bolting it on at the end. (I work on BestAppify, that's where the numbers are from.)

  7. 1

    Makes sense. Are you planning to charge for it, or keep it free for now?

    1. 1

      The writeups like this one stay free — that's how people find us. The paid side is the tracking underneath: we crawl 27,115 live Shopify apps and 2,495 keywords, and that ongoing monitoring is the actual product. A one-off category score like this is basically a demo of what the tool sees day to day. (I work on BestAppify.)

  8. 1

    Curious how long it took before you saw the first real results?

    1. 1

      We don't actually log a "time to first result" number for something like this — it's not a metric we track. What I can tell you is the crawl date: the scoring ran off our data and every shortlisted listing was re-read live on 21 September 2026, so what's published is only a few days old. Beyond that I'd just be guessing, so I won't. (I work on BestAppify.)

  9. 1

    How did you decide this was worth building in the first place?

    1. 1

      The trigger was looking at the raw listings and realizing the stars told us almost nothing: 101 of the 111 returns apps with any reviews at all sit at 4.5 stars or better, and the median rated app has just seven reviews. When nearly everyone's a 5-star app with single-digit review counts, star rating stops being a real signal — that gap is what the scoring was built to fill. (I work on BestAppify.)

  10. 1

    Interesting take. Would you still recommend this approach to someone starting today?

    1. 1

      Mostly yes — the alternative is trusting star ratings that don't discriminate between apps, which is worse. The honest caveat: 126 of the 237 apps in this category have never been reviewed at all, so for over half the list you're scoring on things other than review signal, and that's softer than I'd like. If you're doing this yourself, budget real time for the manual re-read step — that took longer than the crawl did. (I work on BestAppify.)

  11. 1

    How did you decide this was worth building in the first place?

    1. 1

      Mostly it came from looking at the numbers and realizing the ratings were decorative. In this category 101 of 111 rated apps show 4.5 stars or better, and the median rated app has seven reviews — that's not enough spread to tell anyone anything. Once you see a category where the star rating can't distinguish between apps, scoring from the underlying crawl data instead of the storefront becomes the obvious next step.

  12. 1

    Rating compression is the same problem in SEO tools. Everyone's site health score clusters around 85-95 because the tools grade generously to avoid churn. When the scale only uses the top 10% of its range, the number stops being information and starts being decoration.

    The 57% negative reviews about support rather than product is the more useful finding. It means the actual differentiator in a mature category is not feature parity — it is response time after something breaks. That shifts the competitive advantage from engineering to operations, which most solo founders underestimate.

    On the zero-reviews question: we sidestep it at UtilitySEO with a free scan that requires no signup. If someone can try the tool in 30 seconds, the review count matters less because they form their own opinion before ever looking at ratings. Harder to pull off when the product requires a Shopify integration to evaluate, but worth asking whether a preview or sandbox mode could serve the same function.

    1. 1

      Same compression shows up in our own numbers for SEO apps — 593 live apps with 'seo' in the name, 228 of them with a review count on record, and the top 10 hold 68% of all the reviews those 228 have earned. So the rating itself isn't doing much differentiating work there either, which lines up with what you're describing. On the sandbox idea: agree in principle, but most Shopify apps can't really be evaluated without a live store connected, so a 30-second no-signup preview is harder to pull off than it would be for a standalone SEO tool. The support-response point is the one I'd bet on more — our returns-category numbers showed 57% of negative reviews were about support, not features, so that's probably where the operational advantage you're describing actually shows up first.