Talent.com
Hackajob Ltd
Lead SRE - AWS PlatformHackajob Ltd • Glasgow, SCT, GB
Search for other jobs
Lead SRE - AWS Platform

Lead SRE - AWS Platform

Hackajob Ltd • Glasgow, SCT, GB
3 days ago
Job type
  • Full-time
Job description

Salary: £100,000 - 100,000 per year

Requirements:

  • Formal training or certification on site reliability engineering concepts and advanced applied experience
  • Demonstrated hands-on experience with Amazon Web Services (AWS), including deploying, operating, and maintaining resilient, highly available workloads in a cloud environment
  • Demonstrated proficiency in reliability, scalability, performance, and enterprise system architecture, with hands-on experience conducting resiliency design reviews and implementing resiliency best practices
  • Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET
  • Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments
  • Proficiency with continuous integration and continuous delivery practices and tooling
  • Proficiency with container technologies and container orchestration
  • Experience troubleshooting common networking technologies and issues
  • Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, with a demonstrated ability to evaluate and recommend suitable new technologies
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve site reliability engineering workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
  • Experience with cloud platforms and infrastructure-as-code tooling in an enterprise environment
  • Familiarity with chaos engineering principles and proactive resiliency testing practices
  • Experience contributing to communities of practice, internal knowledge-sharing forums, or engineering guilds
  • Exposure to advanced observability platforms and distributed tracing in large-scale production environments

Responsibilities:

  • Consistently champion site reliability culture and practices, documenting and sharing knowledge across our organization through internal forums and communities of practice
  • Drive initiatives to improve the reliability and stability of our teams applications and platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks
  • Collaborate with our team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets
  • Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response
  • Serve as the primary point of contact during major incidents for our application, applying strong diagnostic skills to identify and resolve issues quickly and minimize business impact
  • Apply deep technical expertise within one or more technical domains, sharing knowledge and providing guidance to peers across the team
  • Drive reuse-first adoption of AI-assisted reliability workflows across the software development lifecycle and toolchain practices (e.g., continuous integration/continuous delivery quality checks, test and validation automation, and operational readiness), ensuring traceability, auditability, resiliency, and security controls
  • Use enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements

Technologies:

  • AI
  • AWS
  • Cloud
  • Support
  • Java
  • Marketing
  • Python
  • Security
  • Spring
  • Spring Boot
  • Web
  • ASP.NET

More:

We are JPMorganChase, partnering directly with hackajob to hire for this Lead Site Reliability Engineer role within Infrastructure Platforms. This role gives you the opportunity to help define the future of a globally recognized firm, make a direct impact on reliability outcomes, and act as a technical authority for medium to large-sized products. We offer a first-class business in a first-class way approach, a strong culture of diversity and inclusion, and a corporate functions environment where our teams support finance, risk, human resources, marketing, and other essential business areas.

last updated 36 week of 2026

#J-18808-Ljbffr

Create a job alert for this search

Lead SRE - AWS Platform • Glasgow, SCT, GB