Job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior SRE Engineer based in United Kingdom.
This is a high-impact SRE role focused on building reliable, scalable infrastructure for an international gaming platform.
You’ll help shape DevOps and SRE practices while contributing to technical strategy and a strong engineering culture.
The role combines hands-on infrastructure engineering with observability, automation, cloud architecture, and incident management.
You’ll work closely with development teams to improve reliability, scalability, deployment processes, and CI/CD practices.
As a senior engineer, you’ll also mentor teammates and promote Kubernetes-native approaches across the engineering organization.
The environment values autonomy, proactive problem-solving, continuous improvement, and practical automation over bureaucracy.
It’s an opportunity to take meaningful ownership of reliability engineering while working with modern cloud-native technologies.
Accountabilities
-
Build and evolve DevOps and SRE practices, helping establish scalable approaches to reliability, infrastructure management, and operational excellence.
-
Contribute to the team’s technical strategy and engineering culture, promoting robust, maintainable, and automation-first practices.
-
Develop, maintain, and continuously improve the gaming platform infrastructure, with a strong focus on availability, scalability, and resilience.
-
Design, operate, and enhance monitoring and observability systems, including metrics, dashboards, exporters, alerting, and reliability indicators.
-
Participate in paid on-call rotations, strengthening incident response processes and helping ensure fast, effective resolution of production issues.
-
Automate infrastructure operations and repetitive workflows to improve engineering efficiency and reduce manual intervention.
-
Partner closely with development teams to improve system reliability, scalability, CI/CD, and change delivery processes.
-
Mentor engineers and share Kubernetes-native practices, helping raise technical standards across the wider engineering community.
-
Contribute to architectural decisions and cloud infrastructure evolution, balancing performance, reliability, scalability, and operational simplicity.
-
Investigate incidents and technical issues systematically, focusing on root-cause analysis and long-term prevention rather than short-term fixes.
-
3+ years of professional experience in DevOps, SRE, or a closely related infrastructure engineering role.
-
Deep hands-on expertise with Kubernetes and container orchestration, including operating and scaling production environments.
-
Strong experience with Terraform and Infrastructure as Code, with an emphasis on reproducible and maintainable infrastructure.
-
Proven experience designing and supporting high-availability infrastructure and production systems.
-
Practical experience working with both SQL databases, particularly PostgreSQL, and NoSQL technologies.
-
Strong knowledge of observability stacks, including Prometheus, Grafana, exporters, and alerting systems.
-
Confident scripting and automation skills using Python and Bash.
-
Solid understanding of CI/CD, GitOps, and platform engineering principles and practices.
-
Strong systems thinking, with the ability to identify root causes, anticipate reliability risks, and develop sustainable solutions.
-
A proactive approach to improving processes, automation, infrastructure, and overall reliability engineering practices.
-
Strong communication and collaboration skills, with experience working closely with software engineering teams.
-
Experience with Oracle Cloud is a plus.
-
Certified Kubernetes Administrator (CKA) certification is desirable.
-
Experience with AWS or GCP is beneficial.
-
Ability to read and understand JavaScript/TypeScript or Ruby code is an advantage.
-
Fully remote working model, with the option to work from dedicated hubs in Kyiv or Warsaw.
-
20 paid working days off, public holidays, and paid sick leave to support sustainable work-life balance.
-
Medical insurance with access to leading healthcare providers.
-
Coverage for psychological support through the Pleso platform.
-
Monthly Benefit Café allowance that can be used toward hobbies, sports, and personal interests.
-
Access to team events, workshops, team-building activities, and company celebrations.
-
Individual learning and development budget for courses and professional growth.
-
Corporate English-language training, workshops, and access to an online learning library.
-
Clear career development and performance review framework, alongside mentoring and training programs.
-
High level of autonomy and freedom of action, with a culture designed to minimize unnecessary bureaucracy.
-
Opportunity to work on technology solutions built from the ground up and contribute to modern cloud-native engineering practices.
-
Collaborative, open, and proactive environment where technical ideas, feedback, and continuous improvement are encouraged.