3
19 Comments

What do you use to know when a background task fails?

If you have a task that has to run reliably, what do you do to make sure it's working?

  1. 2

    A task that has to run continuously or periodically or is dispatched at random?

    1. 1

      good question -- periodically

      1. 2

        Using cronjobs (or Laravel scheduled tasks) and decent error reporting helps a lot. I rely heavily on sentry.io for error reporting in combination with messages to a private discord channel which I'm immediately made aware of. This helps in being reactive and able to respond promptly to misbehaving code.

        1. 1

          If the cronjob itself doesn't run, will this system tell you?

          1. 5

            @tigran made cronhub.io does tell you.

            1. 1

              Thanks for sharing the word, @shree. Appreciate it. :)

              1. 1

                @tigran I use heroku apps for app.accelerlist.com and want to onboard onto cronhub.io. My jobs are regular interval jobs scheduled on heroku clock processes. Any clue if that's compatible with cronhub? Any starter tutorial I could use?

            2. 1

              Nice. Thanks for the link

          2. 2

            I forgot to add ;) I also have a grafana+prometheus machine which is able to do monitoring, but also track metrics from my database. So when too many items in my database become stale it triggers alerting, which is the indication my background scripts are not working as expected.

      2. 1

        That is, logs seem to be ok for tracking failures. They're useful, but not as useful for tracking missing scheduled tasks.

        1. 2

          Knowing when a task hasn't run is much harder. That's why decent exception handling makes a lot of sense. Running without that in production to me is a huge no-no.

          1. 1

            What qualifies as decent exception handling?

            1. 2

              That is a good question and entirely depends on your business. I'd say that anything unexpected should be logged, you can then screen these from your dashboard, fix them or ignore them. That way you have full control of the state of your application.

  2. 1

    For failed tasks, there's a bunch of services: https://deadmanssnitch.com/ comes to mind.

  3. 1

    You can use a queue system with automatic retries when they fail. Though, it takes some tweaking so that you aren't automatically retrying invalid requests/data over and over. You could set a max retry and then manually handle any retries that exceed the maximum.

  4. 1

    At https://diffy.website we have multiple cron jobs running and different level of alerts.

    Queue system. Based on MySQL. Workers might stuck while processing it. That is why we have a separate cron job that checks the queue and if there are some jobs that are too old (we expected them to be processed) -- send an email.

    Job that pulls workers results. Runs via supervisor. If it fails -- similar to above, separate monitoring job checks and sends email.

    We do use pingdom for general uptime of the main site. Mainly to catch failed deployments.

    For you case I think your cron job should submit some results somewhere upon completion. Can be record in logs, update timestamp in some file. So you could create another job (preferrably on different environment) that will check if that record was updated. And if not -- send you a notification.

  5. 1

    Hey Doug, funny to run into you here :)
    At caddyserver.com and relicabackup.com, we use https://uptimerobot.com/. It was designed for your problem.

    1. 1

      Hey Cory, good to hear from you. Relica looks slick.

  6. 1

    You need a monitoring system that can alert you.

    • Generate metrics
    • Send them to a metrics platform that stores them (InfluxDB, Prometheus, etc.)
    • Generate dashboards and alerts from stored data (Grafana, or the tools built into the metrics platforms themselves)

    For a high-level diagram, see the one at the top of my site: https://hostedmetrics.com/

    If your task generates artifacts, such as logs, that can me monitored for symptoms of failure, than you can just scrape that data and send it to the metrics platform. InfluxDB's Telegraf component is able to pull data from many types of systems and applications.

    The alternative is to instrument your code with one-liners to generate metrics and send them to the metrics platforms using StatsD. For example, your task can generate a task_run data point every time it runs. Your alerting tool can easily keep an eye on that metric for error scenarios. The example at https://hostedmetrics.com/documentation/instrumentation/ shows you the quick bare bones of generating metrics for a repetitive task.

    I'd be glad to have a quick chat with you if you'd like more detailed advice (regardless of whether HostedMetrics.com is of interest to you).