3
5 Comments

Idea for a Relevance-Aware Content Reposter, Feedback Appreciated

Hello Hackers. With this post I'm taking my first step from lurker to Indie Hacker. I would greatly appreciate any feedback you have on this idea that I'll call ContentRefresher for now.

ContentRefresher gets more value out of your existing content (text, images, audio, video) by making relevance-based reposting suggestions based on current events and posts from influencers. ContentRefresher uses natural language recognition to gain a deep understanding of what your content is about. You don't need to categorize or tag anything.

Imagine that

  • A story about a Major World Leader ducking out on a restaurant bill is trending in the news. You made a series of jokes about a scenario like this on your podcast months ago. ContentRefresher would notice the similarity and suggest you post a link to that episode on your social media accounts.

  • An influential Twitter user tweets a picture of their cat wearing a party hat. You did a blog post last year about cats with hats. ContentRefresher would suggest that you tweet a link to them.

  • A blog post on Medium about software as a business is trending on Hacker News. You happen to run a popular site centered around independent hackers. ContentRefresher would suggest that you comment on this blog post with a relevant backlink.

What are your thoughts on the usefulness and marketability of this sort of AI fueled reposting strategy?

  1. 1

    This sounds like absolutely awesome tech, I think the tech part of it would be the most difficult. From your examples you have voice recognition (assuming no podcast transcript), social media monitoring, and a ton of different content sources. The content extraction alone would be a lot of work, then unifying the data types and feeding it into AI. If you can get it to work though, it sounds amazing.

    One thought would be rather than finding popular posts, finding up and coming posts. A lot of success in the comments of reddit or HN comes from being an early comment. The tech sounds like it could also make really good underpinnings of a search engine if it could accurately index content.

    I do think there is a lot of value in finding relevant social media posts, even if you don't go full AI fueled to do it. Even for promoting a site rather than back linking to blogs. For example, I have a site for tracking personal relationships, I would love to have something that found people complaining about forgetting people's names or not having access to the data that companies collect.

    1. 1

      Thanks for the feedback. Your idea about getting in early on comments is great. Being able to get early comments by detecting 'hotness' is a potential killer feature.

      From my point of view, the tech is actually the most straightforward part as this is what I do for a day job. Several big players offer voice recognition SaaS that's so accurate it's scary. The rest basically reduces down to a document classification problem. I worked on a project over a decade ago that was already getting classification results so accurate that you couldn't tell if the documents had been classified by human or computer (in their domain, humans only agreed about 85% of the time, IIRC). Acquiring customers, getting publicity and all that business-y stuff is the real wizardry.

      1. 1

        How are you going to scale the tech? For this to work you'll need to monitor almost the entire internet. This is a lot of content. Plus, each new indexed document will slow down your matching algorithms so you'll have to throw a lot of hardware at the problem.

        This will be super hard to do as a one man show. This is a google scale tech problem.

        1. 1

          Lucky for us, Google scale tech is now freely available all over the place thanks to the Big Data movement. The new stream processing systems that have come out in the last few years (thank you Apache project) are semi-miraculous. If you use Apache Storm on top of Kafka and don't do anything dumb like require sequential processing at some point, linear scalability is no longer hard to achieve.

          And you really only need to monitor a few info sources to pick up on what The Internet is interested in. Google News, Google Trends, Facebook, Twitter, Hacker News, and a few dozen notable others. The amount of current events is kept small by only maintaining a sliding window of data; news is old news after a week or so. Matching is kept fast by sharding content data by user, so processing time stays constant because you can scale horizontally just by adding more resources. If you set up reactive scaling with a management layer like Kubernetes (actual Google tech) and run it on a cloud provider, it will even scale itself as needed.

          I wouldn't call it easy, but definitely not super hard with the right technology under the hood.

          1. 1

            Ah, okay. If you are starting with a handful of sources that won't be a problem.