Code quality and System reliability: Key strategies to maintain them
August 25, 2024 · Engineering Management

Recently, I was asked a question from a colleague …
What strategies have you employed to increase feature release velocity while maintaining code quality and system reliability?
and I found out that I was doing a lot of stuff, but I have never had a well-thought strategy behind them. I just reacted on things, or made lots of unconscious decisions, but most of them well educated.
Realistic and tangible
There are some high-level strategies that you have choose, in order to keep high standards in terms of code quality and system reliability. For me, code quality equals to maintainability and system reliability equals to system uptime.
The next step is to make them tangible, like a KPI that everyone knows how to correlate with.
Code quality and maintainability
If the codebase is in good shape, then it should be understandable and highly maintainable. Why do you need to have these characteristics?
- You need to fast develop new features, while maintaining the old ones in order for your business to thrive
- You need to easily onboard new engineers to the code base
So, if you put a notion of deliverability and measure it, you have a good indicator for code quality. One good candidate that I have lots of experience with, and it works, is the OKR framework.
It is a framework that you put some goals in place and you measure their results. If you are more than 60–70%, and your goals are quite ambitious, then typically you are in a good shape. Everyone understands this and you just connected a very technical characteristic like “code quality” with a set of KPIs that everyone in your company understands.
So, now that we have defined its success criteria, which are the key pillars to increase our possibility for success?
- Software development process You have to have in place a well-thought software development process, supported by automation that run tests and ensure that features will be deployed safely in production
- Technology stack Which programming languages and/or framework should the tech team invest in?
- Inner-sourcing Source code should be open to all engineers (in principle), with clear documentation and proper method to contribute (if needed), but with clear ownership. This approach will not only eliminate dependencies, but it will create good code hygiene.
- Architectural board A team of experienced engineers that will push the architecture of the system and its principles to the next level.
- Documentation In a practical manner, this should not become an obsession, there should be easily accessible documentation describing the system.
In addition, it is really important to pick your battles, especially if you have a small team. You cannot build all technologies, but you cannot also create a Frankenstein architecture. A rule of thumb is to build (if possible) technologies that are critical to the business, because there is the potential to create more value and opportunities.
System reliability and uptime
How easy it is to correlate with business KPIs? What it really means for a system to be not available? In some cases nothing, but the most obvious KPIs is the loss of revenue.
So what are the key factors to fortify and maximise the system uptime?
- Observability Proper monitoring of the system. Graphs, logging and alerting on abnormal behaviour of the system.
- Incident management process A well-establish process that minimize the downtime and the potential loss of customers and revenue
- Secure software development and awareness for the engineering teams. This should be first-class citizen, in order to minimize security risks.
- Infrastructure-as-code approach. Automate everything and manage efficiently your infrastructure.
- Cloud-first architecture You should not have everything on the cloud, but also benefit from ready-made proven services from cloud providers.
- Cost Always should be taken into consideration, because if your infrastructure is very expensive, it is the first thing that you will need to optimise in a crisis. When you do this without any planning, then usually reliability is heavily affected. As a friend told me, “You can cut expenses only once”.
- Testing Unit and integration tests should be first-class citizen on your pipelines.
Epilogue
These are some key strategies that should be adopted, in order to have an efficient technology landscape that you can build upon. If you have all the above in check, then you can easily increase the team’s capacity by hiring in an efficient way. You will also be sure that feature can be delivered with the minimum amount of risk, but remember, the risk is never going be zero.