Talent.com
OpenAI
Software Engineer, GPU Infrastructure- ChatGPT EngineeringOpenAI • London, England, UK
Software Engineer, GPU Infrastructure- ChatGPT Engineering

Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • London, England, UK
26 days ago
Job type
  • Full-time
Job description

About the Team

ChatGPT Engineering builds and operates the compute platform powering one of the worlds largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability efficiency and performance.

As our GPU fleet continues to grow were investing in the infrastructure that operates it. Our team builds the tooling automation and intelligent systems that make GPU infrastructure scalable observable and increasingly autonomous. We work across production engineering distributed systems capacity management and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.

This is a unique opportunity to work on infrastructure at the frontier of AI where small improvements in fleet efficiency reliability and automation have an outsized impact on the development and deployment of AGI.

About the Role

Were looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.

Youll design and build the systems that manage GPU clusters at scalefrom fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. Youll partner closely with infrastructure research and product engineering teams to improve reliability developer productivity and overall compute utilization.

This role is ideal for engineers who enjoy solving complex operational challenges building internal platforms and working on infrastructure that directly powers frontier AI.

In This Role You Will

  • Design build and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.

  • Build internal platforms tooling and AI-powered agents that automate fleet operations and reduce operational overhead.

  • Improve observability reliability and operational efficiency across thousands of GPUs.

  • Develop systems for capacity planning scheduling fleet health monitoring and incident response.

  • Identify infrastructure bottlenecks and implement solutions that improve utilization scalability and performance.

  • Partner closely with research platform networking and systems teams to continuously improve our compute platform.

  • Help establish engineering best practices around operational excellence automation and infrastructure reliability.

You Might Thrive in This Role If You

  • Have experience operating large-scale production infrastructure preferably GPU clusters or other compute-intensive distributed systems.

  • Have a background in Production Engineering Site Reliability Engineering (SRE) Infrastructure Engineering or Platform Engineering.

  • Have built software that automates operational workflows rather than relying on manual processes.

  • Have experience with Kubernetes Linux systems container orchestration or distributed infrastructure.

  • Understand infrastructure observability monitoring capacity planning and incident management.

  • Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity.

  • Are comfortable working across software engineering and systems operations owning problems end-to-end.

  • Thrive in fast-moving environments with significant technical ambiguity.

Qualifications

  • 5 years of software engineering experience building production infrastructure.

  • Strong programming skills in Go Python C Rust or similar systems languages.

  • Experience designing and operating highly available distributed systems.

  • Experience with GPU infrastructure high-performance computing ML infrastructure or large-scale compute platforms.

  • Experience with Kubernetes cloud infrastructure Linux networking and observability tooling.

  • Excellent debugging systems design and operational problem-solving skills.

  • Strong communication skills and experience collaborating across engineering organizations.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We build AI systems that are capable aligned and broadly beneficial. Our infrastructure teams power the research and products that bring these systems to millions of users worldwide.

We believe diverse perspectives make stronger teams and better technology. Were committed to creating an inclusive workplace where people from all backgrounds can do their best work.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core and to achieve our mission we must encompass and value the many different perspectives voices and experiences that form the full spectrum of humanity.

We are an equal opportunity employer and we do not discriminate on the basis of race religion color national origin sex sexual orientation age veteran status disability genetic information or other applicable legally protected characteristic.

For additional information please see OpenAIs Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws including the San Francisco Fair Chance Ordinance the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct adverse and negative relationship with the following job duties potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary confidential and non-public addition job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI we believe artificial intelligence has the potential to help people solve immense global challenges and we want the upside of AI to be widely shared. Join us in shaping the future of technology.


Required Experience:

IC


Employment Type : Full-Time
Experience: years
Vacancy: 1
Create a job alert for this search

Software Engineer, GPU Infrastructure- ChatGPT Engineering • London, England, UK

Similar jobs

Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAIGreater London, England, GB
Full-time

ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products.Every ChatGPT conversation relies on massive GPU clusters serving inference workloads wi... Show more

 • Promoted

Infrastructure Engineer

Third Nexus Group LimitedGreater London, England, GB
Permanent

Salary: £65,000 - 90,000 per year.Experience with Infrastructure as Code, such as Terraform.Experience with automation scripting, such as Python or Shell.Understanding of the software development l... Show more

 • Promoted

Senior C++ Mobile Infrastructure Engineer - Remote

Randstad Technologies RecruitmentGreater London, England, GB
Remote
Full-time

Randstad Technologies is seeking a Senior C++ Engineer (Mobile Infrastructure) for a contract role that is 100% remote.The project focuses on modernizing a mature C++ core shared library across iOS... Show more

 • Promoted

Senior Software Engineer - GPU Capture and Replay

PlayStation GlobalGreater London, England, GB
Full-time

Senior Software Engineer - GPU Capture and Replay.Why Sony Interactive Entertainment?.Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work.Sony Intera... Show more

 • Promoted

Infrastructure Engineer

Extech 2000 Limited.Croydon, England, GB
Full-time

Looking for two infrastructure engineers to join the newly formed infrastructure engineering team.Production use of either AWS or configuration management tools - can be either/or because of the tw... Show more

 • Promoted

GPU Core Architect

APPLEGreater London, England, GB
Full-time

APPLE is seeking talented individuals to join the GPU Platform Architecture team.In this role, you'll innovate on GPU Core Architecture, collaborating with experienced architects and partners.Candi... Show more

 • Promoted

Cloud Infrastructure Engineer

Infinity QuestGreater London, England, GB
Full-time

Proven experience in AZURE Terraform code.Demonstrable experience in GitHub Actions.Experience of working in a controlled environment with strict change control.Good understanding and practical imp... Show more

 • Promoted

Senior Software Engineer - GPU Capture and Replay

Sony Interactive EntertainmentGreater London, England, United Kingdom
Full-time

Why Sony Interactive Entertainment?.Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work.Sony Interactive Entertainment (SIE) is the company behind th... Show more

 • Promoted

Founding GPU Engineer

Fuse EnergyGreater London, England, GB
Full-time

Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale... Show more

 • Promoted

Senior GPU Software Engineer - Remote UK

QualcommGreater London, England, GB
Remote
Full-time

Qualcomm Technologies International Ltd is seeking a Senior Engineer for the UK with remote work options.You will architect, implement, and optimize GPU software, drivers, and tools, collaborating ... Show more

 • Promoted

Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates LimitedGreater London, England, United Kingdom
Full-time

Hamilton Barnes Associates Limited is seeking a senior infrastructure architect to design and own a landmark GPU deployment.You will specify cluster topology, networking, and power/cooling for a ne... Show more

 • Promoted

Software Engineer

Oriole NetworksGreater London, England, GB
Full-time

We are looking for Software Engineers to develop embedded and host software to manage and monitor our high-speed network.These engineers will be part of the team building solutions to connect GPU s... Show more

 • Promoted

AI Infra Engineer — GPU & MLOps

VEEDGreater London, England, GB
Full-time

VEED, a pioneering generative AI company, is seeking a professional to manage GPU infrastructure and deploy AI models.The role involves ensuring reliability for enterprise customers and managing in... Show more

 • Promoted

Founding GPU Engineer

Fuse Energy, LLCGreater London, England, GB
Full-time

Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale... Show more

 • Promoted

Compute Platform Engineer — Multi-Cloud GPU & Kubernetes

Reflection AIGreater London, England, GB
Full-time

A technology company in the UK is seeking a skilled team member to enhance their Compute Platform.The role focuses on managing a K8s-based platform, ensuring system health, and improving performanc... Show more

 • Promoted

Senior GPU Systems Engineer: Large-Scale Inference & RL

ReflectionGreater London, England, GB
Full-time

Reflection is seeking an experienced professional to design and operate large-scale GPU infrastructure for model inference and reinforcement learning.The role involves developing systems for high-p... Show more

 • Promoted

Software Engineer/Platform Engineer - Go/Linux in London - Quant Capital

Golang WorksGreater London, England, GB
Full-time

Software Engineer / Platform Engineer – Go/Linux.Location: London, United Kingdom.Quant Capital is urgently seeking a Software Engineer / Platform Engineer with expertise in the Go programming lang... Show more

 • Promoted

Senior Serverless AI Engineer - GPUs & Cloud, Hybrid

NebiusGreater London, England, United Kingdom
Full-time

Nebius is building a full-stack AI cloud platform and is hiring a Senior Software Engineer to help scale Nebius Serverless AI in a hybrid London/Amsterdam setting.You will own architecture decision... Show more

 • Promoted

Infrastructure Software Engineer

Coram AIGreater London, England, GB
Full-time

The Infrastructure Software Engineer will be responsible for a large part of our edge and cloud stack underpinning our portfolio of IoT products - not just in terms of infrastructure, but also to b... Show more

 • Promoted

Customer Engineer, Infrastructure Modernization, UKI at Google – London, UK; Manchester, UK

VictraysGreater London, England, GB
Full-time

Customer Engineer, Infrastructure Modernization, UKI at Google – London, UK; Manchester, UK.Customer Engineer, Infrastructure Modernization, UKI at Google – London, UK; Manchester, UK.Bachelor’s de... Show more