2
12 Comments

Published an honest AI vs Studio comparison for jewellers (not a sales pitch)

Added a new post to the JewelViz blog today, aimed at actual jewellers researching their options.
Tried to keep it balanced — covers real studio shoot costs (photographer, styling, turnaround time) alongside where AI photography genuinely helps (catalog-scale work, budget, speed) and where it doesn't (flagship pieces, highly intricate designs where AI can still drift slightly).
Figured a comparison piece that isn't just marketing copy might actually build more trust than a features list would.
jewelviz.com/blog

on July 16, 2026
  1. 1

    Smart, but who is reading it right now? I've worked closely with Jewellers and they usually aren't researching this through blogs, they're on Google or asking someone.

    1. 1

      Fair pushback. I'd separate two things though: nobody's browsing my blog directly, but that's not really the bet — it's built to catch the Google searches jewellers are already doing ("jewelry photoshoot cost India", that kind of thing). If someone's already searching, the blog is what shows up instead of a generic listicle.
      The "asking someone" part is real though, and a blog post doesn't solve that — that's word-of-mouth/referral, which I'm not really addressing with content at all right now. Probably a gap worth thinking about more than the blog itself.

  2. 1

    The useful line is that intricate flagship pieces still drift, but "intricate" needs an observable boundary. I'd add a blinded set of 20 pieces, show which details changed, and report jeweller acceptance by category and stone setting. That turns a balanced article into a buying rule instead of a softer sales page.

    1. 1

      Fair callout — "intricate" is doing a lot of work in that sentence without backing it up.
      I actually have some informal signal on this already: in testing, simpler/symmetric designs (chokers, single-pendant necklaces) held up well, while multi-tier asymmetric pieces (like layered earrings) drifted more — occasionally missing a tier or adding an element that wasn't in the original. But that was eyeballed comparison, not a structured blinded test with real acceptance data.
      A proper version of what you're describing — 20 pieces, category-tagged, jeweller-scored — is genuinely the right next step if I want this to be a buying rule instead of a vibe. Don't have it yet, but noting it as the actual next test to run rather than another prompt-engineering phase.

      1. 1

        Missing a tier and adding a nonexistent element shouldn't average together with minor shape drift. I'd tag each failure as reject, repairable, or acceptable before asking jewellers to score the 20 pieces.

        1. 1

          Thanks for the thoughtful feedback. I'm actively working on improving this. Based on my own testing, the AI is currently achieving around 70–80% jewellery accuracy across many designs, and my goal is to push that significantly higher over the next month.
          I'm also spending more time speaking directly with jewellers and learning from their feedback, because I believe first impressions matter the most. If a jeweller doesn't trust the very first result, nothing else matters. That's the standard I'm building toward, and feedback like yours helps me get there.

          1. 1

            70–80% is only useful if the unit is explicit. If one missing tier is a reject and a small shape shift is repairable, headline accuracy can improve while first-result trust stays flat. For the next month I would track first-result reject rate by design class, then repair time for the rest.

            1. 1

              You're right, and I'll be straight about it — my last reply dodged your actual point. I threw out the same 70-80% number again without doing what you'd already told me to do: separate reject from repairable.
              Concretely, starting this: every generation gets tagged as one of three — reject (unusable, redo needed), repairable (minor shape/color drift, fixable), or acceptable (matches close enough). Then I track reject rate specifically by design class — necklace vs earrings vs bangles — since that's where the real gap seems to be. Repair time for the repairable bucket is separate.
              Will report back with actual numbers once I have a real batch tagged this way, not another blended percentage.

              1. 1

                Freeze the rubric before tagging the batch, then have two people independently label the first 20 results. Otherwise “repairable” will drift as you learn, and the class comparison becomes subjective. Report n and rejects per design class, median repair minutes, and the disagreements; that will make the next percentage defensible.

                1. 1

                  Fair enough on the rigor gap — appreciate you not letting me settle for "good enough."
                  Since you've clearly thought about this in detail, would you be open to actually looking at some real outputs? I have the full set for a few necklace pieces — original photo plus all 3 generated angles.
                  There's a WhatsApp button on jewelviz.com that goes straight to me — feel free to message there if you want to eyeball the drift yourself instead of just taking my word for it.

                  1. 1

                    Put the set in this thread as a blinded contact sheet instead of moving it to WhatsApp: original, three generated angles, and only a piece ID. Then publish the frozen rubric and both labelers’ decisions after the reveal. That keeps the test reviewable instead of turning one eyeball check into another private percentage.

                    1. 1

                      Two ways I can do this, your call:
                      I already have generated sets from earlier testing (a few necklace pieces) — I can post those now as a blinded contact sheet with piece IDs, no delay.
                      If you'd rather rule out any chance of cherry-picking, send me 2-3 jewelry photos of your choice and I'll generate fresh live sets from them, screen-recorded end to end, then post those instead.
                      Either way I'll follow your format — frozen rubric first, piece IDs only, rubric and labels published after the reveal.