GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms
You will also work closely with our development teams to improve reliability, scalability, and operational efficiency of our systems, while ensuring code can be deployed safely and efficiently across the organization
You will also play a key role in responding to production incidents, troubleshooting complex issues across the infrastructure stack, and driving improvements based on what we learn from them
GIPHY serves internet-scale traffic every day, so the ideal candidate must be comfortable with operating large and distributed systems and solving problems at scale. Your work will directly impact billions of daily users, alongside some of the world’s largest social media platforms and apps
Operate and evolve GIPHY’s cloud and CDN presence, ensuring resources are reliable, performant, scalable and cost-efficient
Continuously improve and optimize our CI/CD platform to enhance developer experience and deployment reliability
Drive the implementation of multi-region Kubernetes clusters while operating and improving our existing Kubernetes environments
Partner with other Engineering teams to troubleshoot complex production and infrastructure issues, taking a holistic view across our platforms to ensure the reliability and availability of GIPHY’s systems
Stay current with emerging technologies, particularly agentic development and AI-assisted engineering, and identify opportunities to incorporate them into our platforms and engineering workflows
Benefits
Flexibility: We offer a hybrid model that’s centered around flexibility and enabling collaboration, innovation, and accountability—no matter where we are.
Connection: Whether online or on site, we make sure there are plenty of ways to connect and have fun.
Recognition: We never miss an opportunityto celebrate our achievements and commend great work through our global recognition program.
Empowerment: We offer generous and competitive compensation packages, PTO, wellness initiatives, on-and-off working hours, tuition reimbursement, and an employee referral bonus.
Growth: Our SkillUP Program lets you take ownership of your personal development through virtual learning offerings.
Belonging: Our goal is to build a workforce that’s representative of the diverse global community we serve, and where all employees can come to work as their authentic selves.
Expert-level knowledge and significant professional experience with Infrastructure as a Service (IaaS) operations and Infrastructure as Code (IaC), particularly Terraform and Atlantis
Demonstrate a high degree of autonomy and initiative, proactively identifying areas for improvement and independently driving work around optimization, cost reduction, patching, upgrades, and technical debt
Experience with agentic engineering and effectively applying AI-assisted development practices
Ability to design and implement cost-effective solutions and proactively identify cost-optimization opportunities
Proficiency in one or more programming languages, such as Python, Java, Go or Rust
Experience with CI/CD tools such as Jenkins or Github Actions and deployment tooling such as Helm, Spinnaker and ArgoCD
Strong systems engineering fundamentals, with deep knowledge of Linux and networking and the ability to diagnose complex reliability and performance issues across the OS, network, container, and cloud infrastructure layers
Strong problem-solving and communication skills, and the ability to work effectively in a team environment
Experience with observability and monitoring platforms such as Datadog and AWS CloudWatch, including diagnosing production reliability and performance issues
Extensive experience with containerization and orchestration technologies, mainly Docker and Kubernetes
Strong experience with operating, designing and troubleshooting AWS infrastructure, especially EKS, S3, EC2 and VPC
#J-18808-Ljbffr
Create a job alert for this search
Senior Site Reliability Engineer (GIPHY) • York and North Yorkshire, ENG, GB