Talent.com
Google
Senior Software Engineer, SRE, Cloud Incident ResponseGoogle • London, England, GB
Senior Software Engineer, SRE, Cloud Incident Response

Senior Software Engineer, SRE, Cloud Incident Response

Google • London, England, GB
30+ days ago
Job type
  • Full-time
Job description

Senior Software Engineer, SRE, Cloud Incident Response

Google place London, UK

Apply

  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 5 years of experience with software development in one or more programming languages.
  • 5 years of experience with data structures or algorithms.
  • 3 years of experience in designing, analyzing, and troubleshooting distributed systems, and 2 years of experience leading projects and providing technical leadership.
  • Experience in SRE or incident management/response environments.

Preferred qualifications:

  • Experience working in computing, distributed systems, storage, or networking.
  • Experience in telemetry systems, incident and risk management.
  • Experience in designing, analyzing, and troubleshooting large-scale distributed systems.
  • Ability to debug, optimize code, and automate routine tasks.
  • Excellent problem-solving skills, with strong verbal and written communication abilities.

About the job

Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both internally critical and externally visible—maintain reliability, uptime appropriate to customer needs, and a rapid rate of improvement. Additionally, SREs monitor system capacity and performance continuously.

Much of our development focuses on optimizing existing systems, building infrastructure, and automating tasks. On the SRE team, you'll tackle the unique challenges of scale in Google Cloud, leveraging your expertise in coding, algorithms, and large-scale system design. Our culture emphasizes curiosity, problem solving, and openness, fostering collaboration and innovation in a supportive environment.

Responsibilities

  • Ensure Google Cloud Platform (GCP) stability and reliability through incident support, driving customer outcomes, and cross-team collaboration.
  • Create training and processes for incident management, collaborating with Cloud Support leadership.
  • Develop systems and tools to improve incident visibility, issue detection, and communication with stakeholders.
  • Identify and escalate risks, reducing major incident probabilities through pragmatic approaches.
  • Support system scalability and reliability throughout their lifecycle via design consulting, platform development, capacity planning, and automation.

Google is an equal opportunity employer committed to diversity and inclusion. We value a workforce that reflects our users and fosters a culture of belonging. We provide equal employment opportunities regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy, or related conditions. See Google's EEO Policy and related resources.

As a global company, English proficiency is required for all roles unless otherwise specified.

Note: Google does not accept resumes from recruitment agencies and is not responsible for fees related to unsolicited resumes.

#J-18808-Ljbffr
Create a job alert for this search

Senior Software Engineer, SRE, Cloud Incident Response • London, England, GB

Similar jobs

Senior Cloud Platform Engineer & Tech Leader

JPMorgan Chase & Co.Greater London, England, United Kingdom
Full-time

Cloud Engineering position within the XLR8 Team to accelerate delivery of public cloud solutions.The role focuses on designing platform services, automating CI/CD, and maintaining secure production... Show more

 • Promoted

Senior Cloud SRE - Kubernetes, GCP & CI/CD

Test TriangleGreater London, England, GB
Full-time

A leading technology firm in Greater London is seeking a Senior Cloud Engineer to contribute technical expertise within a cloud engineering team.This role involves architecting scalable Kubernetes ... Show more

 • Promoted

Senior Cloud Reliability Engineer

CartaGreater London, England, GB
Full-time

Carta is seeking a Senior Site Reliability Engineer to build and scale internal platform services, ensuring reliability and performance across Carta's applications.You will design monitoring and in... Show more

 • Promoted

Senior SRE — Global Platform & Observability

Carta, Inc.Greater London, England, GB
Full-time

Carta is hiring a Senior Site Reliability Engineer to help build and scale its internal platform offerings, focusing on compute, storage and networking services to ensure reliability and performanc... Show more

 • Promoted

SRE: Application Reliability & Incident Response (Hybrid)

ZILOCity Of London, England, GB
Full-time

A technology company is seeking a Site Reliability Engineer to join their SRE team.The ideal candidate should have experience with application debugging in Java, Golang, or Python, solid PostgreSQL... Show more

 • Promoted

Senior Software Engineer II - FinCrime - Scam Prevention Team

hackajobGreater London, England, United Kingdom
Full-time

Wise is a global technology company, building the best way to move and manage the world’s money.Whether people and businesses are sending money to another country, spending abroad, or making and re... Show more

 • Promoted

Senior Software Engineer II - FinCrime - Scam Prevention Team

WiseGreater London, England, United Kingdom
Full-time

Join us as a Software Engineer in our critical Scam Prevention team.This is a high‑impact engineering role with a strong product focus, dedicated to solving complex, real‑world adversarial problems... Show more

 • Promoted

Lead Security Incident Response Engineer: Cloud & Forensics

Checkout.comGreater London, England, United Kingdom
Full-time

London is seeking a senior security incident response lead to own the technical direction of incident response across the company.You will guide investigations, containment, eradication, and recove... Show more

 • Promoted

Platform Engineering Leader: SRE, Cloud & DevOps

SterlingsGreater London, England, United Kingdom
Full-time

Sterlings is partnering with a global, highly regulated financial services organisation to appoint a Director / Head of Platform Engineering.The role leads Platform Engineering, SRE, Cloud, Databas... Show more

 • Promoted

Senior Cloud Platform Engineer – Security & IaC Lead

CatapultGreater London, England, GB
Full-time

Catapult in London is seeking a Principal DevOps Engineer to enhance their cloud platform.You will own infrastructure management and ensure compliance with security standards, while collaborating c... Show more

 • Promoted

Senior Platform Engineer: Observability, SRE & IaC

FocusedGreater London, England, GB
Full-time

Focused is looking for an experienced Senior Platform Consultant to join their London team.In this role, you will be responsible for augmenting infrastructure with observability solutions, implemen... Show more

 • Promoted

Senior AI Reliability Engineer – ML Infra & Incident Lead

AnthropicGreater London, England, United Kingdom
Full-time

Anthropic is seeking a reliability-minded software engineer to enhance the reliability of its AI systems.This role involves developing Service Level Objectives for large language model serving syst... Show more

 • Promoted

Senior AI Infrastructure & DevOps Lead

DNEGGreater London, England, United Kingdom
Full-time

Brahma AI is building state‑of‑the‑art generative AI platforms, powering AI models and GPU‑heavy workloads across multi‑cloud environments.We seek a Lead Infrastructure/DevOps Engineer to drive a 7... Show more

 • Promoted

Remote Lead Incident Response Consultant

Palo Alto Networks, Inc.Greater London, England, United Kingdom
Remote
Full-time

Palo Alto Networks is seeking a client-facing Principal Consultant to lead reactive services engagements and manage incident response from start to finish.You will guide clients through containment... Show more

 • Promoted

Elite DevOps & SRE Engineer - Hybrid London

Hunter BondGreater London, England, United Kingdom
Full-time

A leading fintech firm in London is seeking a DevOps/Site Reliability Engineer to join its elite team.This role involves hands-on engineering, improving large-scale infrastructure, and automating t... Show more

 • Promoted

On‑Site GCP SRE Lead: Incidents & Cloud Ops

WALT LabsGreater London, England, GB
Full-time

WALT Labs is looking for a Site Reliability Engineer in Greater London.The role involves maintaining cloud infrastructure and managing critical incidents, with a strong focus on Google Cloud Platfo... Show more

 • Promoted

SRE Platform Engineer Lead — High-Availability Cloud

Goldman Sachs Group, Inc.City Of London, England, GB
Full-time

Lead Site Reliability Platform Engineer (SRE) in London.This role emphasizes system reliability, performance, and collaboration across teams.Your mission is to design secure, scalable systems on AW... Show more

 • Promoted

Observability Platform SRE - Greenfield, Cloud-Native

AalyriaGreater London, England, United Kingdom
Full-time

Aalyria is seeking an experienced Site Reliability/Platform Engineer to build the core observability stack for satellite and deep-space platforms.You will design centralized monitoring across metri... Show more