I've seen lots of discussions on which software stack (PHP/Laravel, Ruby on Rails, Python/Flask, etc.) to use and what hosting provider (AWS, DigitalOcean, Heroku,etc.) to use.
But how did you learn to run your service? How did you learn how to:
EDIT: Thanks to everyone for the responses. It seems that most folks learned from working on live systems with other engineers. So did I. Possibly a blind spot in our training of new programmers...
In the end it will all come down to trial and error.
We learned to deploy builds very systematically. We heard about Docker and gave it a shot. We started by building our first images, and swapping services with Docker deployments in stages. We run a bunch of services, so this gave us a very gentle learning curve to running services with Docker instead of SystemD.
Once we got control over using Docker to run our services, we started looking at orchestrating our deployments more easily. I gave a presention about that a year ago that you can read here https://speakerdeck.com/briandeheus/a-docker-swarm-love-story
It's worthy to note that as our team grew, we made the switch to Kubernetes. However to get started with orchestration I would still encourage people to give Swarm a whirl as it's the easiest to get started with.
Again it's all a learning process. When you get your first call from a customer at midnight that something isn't working, you'll go rummage through your logs to see if you can find any hints. You'll soon realise you weren't logging half as much as you should, and hopefully at this point you'll realise you should log as much as you can, and use logging levels to get better insights once shits the fan. We generally
infolog at crucial moments, when a call is made, when talking to external systems, and when a call is finished. Because this does does generate a lot of logs, it's important that you think about ways to aggregate and search through your logs easily. It's up to you and your team to find out what works best.As for deciding what to monitor, again: it's all a learning process. When you notice a service discrepancy, be it either via customers complaining or via your logs you'll see some symptoms. For example a queue not clearing, high response times, or down times. These are excellent points to set up alerts for so you can be more proactive.
Oh boy. There is never a good time to learn this. Start now. Start today. Start a minute ago. Keep backups every hour. And test your backup strategy frequently. When you learn you should've kept backups it's too late. Start now. Others have learned these lessons for you.
Can you guess? It's a learning process again. The first time you'll have an outage you'll realise that you need to figure out how many customers have been affected, your total downtime, the recovery steps, and possibly restitution. What works for us doesn't work for everyone, so try to come up with something that works for you. In the end communication is going to be key.
I agree, it is definitely trial and error. There are of course many books, boot camps, video courses and other learning resources. But trying to learn everything you might need at some point in the future will just slow you down. It might also give you the impression that you have to certain things (in a certain way). But you don't. Every application and service is different and you should only look for solving actual problems you encounter.
On the other hand, having experience in solving things like deployments, logging, monitoring might give you a head start. Just as in other parts in life, more experience can enable you to move faster. If you solved a problem once, you can solve it again. But again, you shouldn't implement a solution just because you know it. You should implement it because it solves an actual problem you have.
Mostly by experience (12+ years) working for companies and from more experienced colleagues.
Before starting my product and company Tideways, I worked as a developer for 8 years at an agency, where the stream of new projects allowed for a lot of experimentation.
After that as a consultant for 4 years I went into >20-30 different companies to see what they were doing and learning a lot in the process.
I have also read a lot of books on programming and operations, and worked on a large open source project for several years.
I already built many side projects 15 years ago and still have code of them, but it was just not very professional, scalable or maintainable. I couldn't have built what I do today without the experience I got working for others.
https://i2.wp.com/www.feedough.com/wp-content/uploads/2016/11/the-internet-meme.jpg?resize=300%2C248&ssl=1
I'm curious to hear what people have to say about this!
As soon as you do something you will eventually come to these questions:
etc...
As soon as you have the question you can find an answer. It's not as hard as someone can think, Google at your service...