Understanding Recruitment
Senior Site Reliability Engineer
Requirements
Requires strong experience in SRE or DevOps with deep Linux systems knowledge and cloud infrastructure expertise. Experience with high-performance, high-throughput systems and tools like AWS, Terraform, or Ansible is highly valued.
Job Description
📍 London
💰 £150,000 - £200,000+ Base + Bonus + Equity
We're partnered with a technology company building high-performance infrastructure for decentralised financial markets.
They're looking for a Senior Site Reliability Engineer to improve the reliability, observability and operational tooling behind a latency-sensitive production platform.
The role covers production infrastructure, monitoring and alerting, incident diagnosis, deployment workflows, infrastructure automation and developer tooling. There is also a strong Linux and systems element, particularly around networking, host performance and running high-performance services in production.
The platform is still relatively early, so there is plenty of scope to improve how things are operated, introduce better automation and help set the standards the wider engineering team works to.
Responsibilities
- Improve the reliability and operability of production systems.
- Build and improve monitoring, logging, tracing, dashboards and alerting.
- Improve incident diagnosis, root cause analysis and operational workflows.
- Build safer and more repeatable deployment and rollback processes.
- Automate repetitive operational and infrastructure work.
- Improve CI/CD pipelines and release processes.
- Develop internal tooling that helps engineers operate production systems more effectively.
- Improve the developer experience from local development through to production.
- Work with Linux systems, networking, host configuration and resource contention.
- Contribute to infrastructure security, access controls, secrets management and system hardening.
The systems are latency-sensitive, so the role can extend into areas such as host-level tuning, kernel settings, CPU isolation and networking behaviour.
Skills & Experience
- Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering.
- Experience operating production infrastructure in cloud environments.
- Strong Linux systems knowledge and understanding of networking fundamentals.
- Experience with monitoring, observability and alerting.
- Strong troubleshooting and root cause analysis skills.
- Experience with CI/CD and infrastructure automation.
- AWS, Terraform or Ansible experience would be advantageous.
- Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable.
- Comfortable taking ownership of problems and driving improvements independently.
Benefits
- £150,000 - £200,000+ base salary.
- Significant performance-based bonus + Equity
- Private healthcare.
- UK visa sponsorship available.
- Engineering-led organisation - built prioritising engineering culture
- Direct influence over reliability, tooling and engineering practices.
- Opportunity to work alongside a small, elite team.
- Exposure to complex, latency-sensitive production systems.
Interested?
Contact Chris Williams with any questions.
Skills
About Understanding Recruitment
The Understanding Recruitment Universe Understanding Recruitment is the go-to destination for technology recruitment with headquarters in St. Albans, England. Specialising in Biotechnology, Artificial Intelligence, and Web3, our team of over 90 recruiters is adept in navigating the dynamic landscapes of Blockchain & Cryptocurrency, Java, JavaScript, Python, Rust, Golang, .NET, DevOps, Product Management, and other tech roles within the Software Development Lifecycle. As your total talent solution partner, we seamlessly connect organisations with top-tier talent and empower tech professionals to discover their perfect fit within the ever-evolving tech industry. We offer unparalleled matches and comprehensive support across the UK, Europe, and the USA. In 2023, Understanding Recruitment became a 60% employee-owned company. This exciting development empowers our dedicated team to share in the financial rewards of our ongoing success. In the same year, we were recognised in Recruiter's annual FAST 50 listing, as the No.1 fastest-growing privately-owned recruitment business in the UK. Our in-house training programme also won a prestigious Princess Royal Training Award that is awarded to employers in the UK and Ireland who can prove that their outstanding training and skills development programmes have resulted in exceptional benefits for their business. With over a decade of success, 2022 marked the year we secured the much coveted Best Companies 3-star accreditation with a remarkable BCI score of 738 or higher, signifying 'world-class' workplace engagement. We've also been honoured as the 'Best Staffing Firm to Work For' for three consecutive years (2016-2018), and were named 'Business of the Year' at the 2017 SME Hertfordshire Business Awards.