Talent.com
CGI
Site Reliability EngineerCGI • London, United Kingdom
Site Reliability Engineer

Site Reliability Engineer

CGI • London, United Kingdom
30+ days ago
Job type
  • Full-time
Job description

Position Description:

We are seeking an experienced and proactive Site Reliability Engineer (SRE) to join a team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, and operational performance of critical data-driven platforms and services across complex production environments.

The successful candidate will work closely with engineering, platform, and support teams to strengthen monitoring and alerting capabilities, improve logging and traceability, troubleshoot production incidents, support deployments, and automate operational processes wherever possible. The environment includes Kubernetes, Helm, the ELK stack, and a strong focus on modern Site Reliability Engineering practices across cloud and platform services.

This is a hands-on technical role suited to someone who thrives in fast-paced operational environments and is passionate about reliability engineering, automation, and continuous improvement. The role requires strong collaboration with both client stakeholders and engineering teams to ensure platform stability, operational excellence, and high service availability

Your future duties and responsibilities:

- Support, maintain, and improve highly available production platforms and services across cloud and containerised environments.
- Manage and support Kubernetes clusters and Helm-based deployments across multiple environments.
- Implement and enhance monitoring, alerting, logging, and observability solutions to improve platform reliability and operational visibility.
- Investigate incidents, analyse logs, identify root causes, and drive timely resolution of production issues.
- Participate in incident response, post-incident reviews, and continuous operational improvement initiatives.
- Automate operational tasks and repetitive support activities to reduce manual effort and improve platform efficiency.
- Work closely with engineering and data platform teams to improve system resilience, scalability, deployment reliability, and operational maturity.
- Develop and maintain operational documentation, support procedures, runbooks, and troubleshooting guides.
- Contribute to reliability engineering practices including proactive monitoring, service health management, and operational readiness.
- Support deployment activities, release processes, and production change management activities.

Required qualifications to be successful in this role:

- Strong commercial experience in Site Reliability Engineering, Platform Engineering, DevOps, or Production Support environments.
- Strong hands-on experience with Kubernetes and Helm in enterprise or production environments.
- Proven experience supporting mission-critical production platforms and operational support functions.
- Strong hands-on experience with the ELK stack (Elasticsearch, Logstash, Kibana) for logging, monitoring, troubleshooting, and operational analysis.
- Demonstrated capability in log analysis, incident investigation, troubleshooting, and root cause analysis.

- Strong understanding and practical experience with core SRE practices including:
Monitoring and alerting
Incident management and response
Root cause analysis and post-incident reviews
Automation and operational improvement
Production support and reliability engineering

-Experience working with data platforms, analytics platforms, or data product teams would be highly advantageous.
- Experience with scripting and automation tools such as Bash, Python, or similar technologies is desirable.
- Exposure to CI/CD pipelines, Infrastructure as Code, and cloud-native environments would be beneficial.
- Strong communication, stakeholder engagement, and collaboration skills.
- Ability to work effectively in fast-paced support environments and manage competing priorities under pressure.

Security Clearance
- Resource must be willing and able to work onsite at the client location five days per week.
- Candidate must already hold current HLC clearance (mandatory requirement).
- Previous experience working within secure, government, defence, or highly regulated environments will be highly regarded.
- Due to client security requirements, only candidates meeting the required clearance criteria will be considered.

#LI-CGISDI

Skills:

  • Amazon Elastic Cloud Compute
  • Elastic Stack & Elasticsearch
  • Helm
  • Linux
  • BASH
  • Kubernetes
  • Python
  • Windows
Create a job alert for this search

Site Reliability Engineer • London, United Kingdom

Similar jobs

Site Reliability Engineer

Reward GatewayGreater London, England, GB
Full-time

Due to expansion, an opportunity has become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach.Integrating tightly... Show more

 • Promoted

Site Reliability Engineer – NS London

BAE Systems Digital IntelligenceLondon, England, GB
Full-time

BAE Systems Digital Intelligence.BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts.We work collaboratively across 10 countries to collect, connect and unders... Show more

 • Promoted

Site Reliability Engineer

NatoboticsCity Of London, England, GB
Full-time

Join to apply for the Site Reliability Engineer role at Natobotics.A Site Reliability Engineer is responsible for transforming the SDLC environment with engineering-focused role that emphasizes sys... Show more

 • Promoted

Site Reliability Engineer

Signify TechnologyGreater London, England, GB
Full-time

Site Reliability Engineer, London, Global Exchange Operator.Days Per week in London Office.We are working exclusively with one of the world's largest exchange operators on an SRE hire in London.Thi... Show more

 • Promoted

Site Reliability Engineer

Autonomai RecruitmentGreater London, England, GB
Full-time

A high-performing trading technology firm is seeking an SRE Engineer to drive reliability, scalability, and operational excellence across critical production systems.This role is suited to a hands-... Show more

 • Promoted

Site Reliability Engineer / Production Support

MonumentLondon, England, GB
Permanent

Site Reliability Engineer/ Production Support .Location London (Oxford Circus) | Hybrid: 2 days per week | Reports to Head of Cloud Operations.We're building something genuinely rare: a financial b... Show more

 • Promoted

Site Reliability Engineer

AlpacaGreater London, England, GB
Full-time

Alpaca is a US-headquartered self‑clearing broker‑dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.Our recent Series D funding round broug... Show more

 • Promoted

Site Reliability Engineer

Signal AIGreater London, England, GB
Full-time

We're on a mission to change the way businesses make decisions with our cutting-edge AI technology.To achieve that, we’re looking for passionate people to join our open and inclusive workplace.Our ... Show more

 • Promoted

Site Reliability Engineer

NominetGreater London, England, GB
Full-time

Select how often (in days) to receive an alert: Create Alert.Hybrid, with a minimum of 20% in the London office per month.We’re Nominet – a world-leading domain name registry operating at the heart... Show more

 • Promoted

Site Reliability Engineer

HelsingCity Of London, England, GB
Full-time

Helsing is a defence AI company.Our mission is to protect our democracies.We aim to achieve technological leadership so that open societies can continue to make sovereign decisions and control thei... Show more

 • Promoted

Site Reliability Engineer

Artificial LabsGreater London, England, GB
Full-time

Help shape the future of specialty insurance.At Artificial, we’re building the next generation of technology for the specialty (re)insurance market.Our mission is to transform how brokers and carri... Show more

 • Promoted

Site Reliability Engineer

ReapitGreater London, England, GB
Full-time

Could you be our next Site Reliability Engineer?.Reapit is the original, end-to-end business technology provider for estate agencies of all sizes.We’ve been helping sales and lettings agents to bui... Show more

 • Promoted

Engineer - Site Reliability

Cedar Cares, IncGreater London, England, GB
Full-time

The Site Reliability Engineer (London) is a role served by experienced technologists with a diverse set of skills ranging from software development to systems, network, application, and database ma... Show more

 • Promoted

Senior Site Reliability Engineer

OmiliaGreater London, England, GB
Temporary

Senior Site Reliability Engineer.This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collabora... Show more

 • Promoted

Lead Site Reliability Engineer

JPMorgan Chase & Co.Greater London, England, GB
Full-time

Our trading technology stack is undergoing a multi‑year convergence and modernization journey.You will play a pivotal role in shaping our next‑generation SRE patterns, reliability frameworks, obser... Show more

 • Promoted

Site Reliability Engineer

INTAPP LIMITEDGreater London, England, GB
Full-time

The Intapp Cloud Platform is a rapidly growing collection of cloud services.As part of a global team, the ideal candidate will be able to quickly move between architecture, design, and daily operat... Show more

 • Promoted

Site Reliability Engineer (Applications)

H&R TalentGreater London, England, GB
Permanent

An amazing Global Investment Client of ours located in Central London are looking for a Site Reliability Engineer to join their team on a permanent basis.This is a rare opportunity and the package ... Show more

 • Promoted

Lead Site Reliability Engineer

J.P. MorganLondon, England, GB
Full-time

Our trading technology stack is undergoing a multiâyear convergence and modernization journey.You will play a pivotal role in shaping our nextâgeneration SRE patterns, reliability frameworks, obser... Show more

 • Promoted

Site Reliability Engineer

Tyk TechnologiesGreater London, England, GB
Full-time

Who are Tyk, and what do we do?.The Tyk API Management platform is helping to drive the connected world and power new products and services.We’re changing the way that organisations connect any num... Show more

 • Promoted

Senior Site Reliability Engineer

Carta HealthcareGreater London, England, GB
Full-time

At Carta, our employees set out on a mission to unlock the power of equity ownership for more people in more places.We believe that the problems we solve today unlock the opportunities of tomorrow.... Show more