Talent.com
JPMorganChase
Lead Software Engineer – LLM Ops Platform ReliabilityJPMorganChase • Glasgow, Scotland, UK
Lead Software Engineer – LLM Ops Platform Reliability

Lead Software Engineer – LLM Ops Platform Reliability

JPMorganChase • Glasgow, Scotland, UK
30+ days ago
Job type
  • Full-time
Job description
Description

Help shape how AI systems run reliably in this role youll build and operate large language model serving infrastructure bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. Youll work hands-on with cloud and Kubernetes-based deployments deep observability and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better youll find meaningful impact and growth here.

As a Lead Software Engineer at JPMorgan Chase in the AI and Machine Learning Platform team you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability performance and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software improve stability and lead incident response and continuous improvement.

Job Responsibilities

  • Design develop troubleshoot and deliver secure high-quality production software and services for AI infrastructure
  • Build backend services and APIs that enable reliable operation of AI infrastructure in production
  • Operate and scale LLM serving infrastructure (such as vLLM and llm-d) including model hosting request routing continuous batching and KV-cache optimization
  • Deploy host and lifecycle-manage open-source and proprietary LLMs on Amazon EKS and Amazon SageMaker as well as on-prem and local GPU clusters using reproducible infrastructure as code and continuous delivery pipelines
  • Implement observability (logs metrics traces) with dashboards and actionable alerting including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads
  • Tune GPU and accelerator capacity autoscaling and cost efficiency for LLM inference workloads using performance and optimization techniques (e.g. quantization parallelism speculative decoding)
  • Lead reliability engineering for LLM endpoints through capacity planning load/soak testing safe rollouts (blue/green canary) failover and incident response for outages and model-quality regressions
  • Participate in an on-call rotation lead incident triage and mitigation and produce clear post-incident root-cause analyses and follow-ups
  • Identify recurring operational issues and automate remediation to improve platform stability and developer experience
  • Build and maintain multi-agent systems with strong orchestration (planning coordination tool-calling state/memory and workflow control) where appropriate
  • Contribute to an inclusive team culture grounded in diversity opportunity inclusion and respect and help drive adoption of leading-edge technologies through communities of practice

Required Qualifications Capabilities and Skills

  • Formal training certification or equivalent practical experience in software engineering concepts
  • Hands-on experience with system design application development testing and operational stability in production environments
  • Advanced proficiency in Python for building production-grade services and tooling
  • Proficiency with automation and continuous delivery methods
  • Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management
  • Strong understanding of site reliability engineering practices including incident management root-cause analysis runbooks and reliability patterns
  • Practical knowledge of observability and instrumentation across metrics logs and traces
  • Comfort with on-call operations and production troubleshooting
  • Hands-on production experience operating LLM inference servers such as vLLM and llm-d (or directly equivalent serving stacks)
  • Hands-on experience hosting and serving LLMs on Amazon EKS and/or Amazon SageMaker and on local GPU infrastructure
  • Knowledge of LLM reliability and risk considerations including latency/throughput trade-offs model and weight versioning prompt/response logging and safe rollout patterns

Preferred Qualifications Capabilities and Skills

  • Experience developing generative AI applications AI agents vector search and retrieval-augmented generation patterns
  • Experience building AI agents using frameworks such as LangChain CrewAI LangGraph or similar orchestration platforms
  • Experience operating or integrating serving platforms such as KServe Ray Serve NVIDIA Triton Inference Server Text Generation Inference (TGI) alongside vLLM/llm-d
  • Familiarity with Amazon SageMaker JumpStart SageMaker Endpoints and Amazon Bedrock for managed model hosting
  • Experience with online LLM quality monitoring (e.g. hallucination toxicity drift detection) and tracing via OpenTelemetry conventions
  • Contributions to open-source LLM serving or inference projects (e.g. vLLM llm-d Ray KServe Triton)



Required Experience:

IC


Employment Type : Full-Time
Experience: years
Vacancy: 1
Create a job alert for this search

Lead Software Engineer – LLM Ops Platform Reliability • Glasgow, Scotland, UK

Similar jobs

AMOS Platform Product Lead

IAG Transform UKWaterside, Scotland, GB
Full-time

IAG Transform UK is seeking an AMOS Platform Manager responsible for the end-to-end product management of the AMOS maintenance platform.The role entails owning the product vision, roadmap, and back... Show more

 • Promoted

Tier 1 Platform Support Engineer

Redsquid CommunicationsScotland, GB
Full-time +1

Redsquid Communications is looking for a Tier 1 Engineer based in Scotland, Fife.This entry-level role is perfect for tech enthusiasts who enjoy troubleshooting and providing exceptional customer s... Show more

 • Promoted

Lead LLM Infra & Reliability Engineer

JPMorgan Chase & Co.Glasgow, Scotland, GB
Full-time

Lead Software Engineer to shape the AI and Machine Learning Platform.This role involves building and scaling AI infrastructure, ensuring reliability, performance, and cost-efficiency of LLM inferen... Show more

 • Promoted

Software Deployment Engineer

Motorola SolutionsGlasgow, Scotland, GB
Full-time

At Motorola Solutions, we believe that everything starts with our people.We’re a global close-knit community, united by the relentless pursuit to help keep people safer everywhere.We build and conn... Show more

 • Promoted

Platform Engineer

Pracyva ltdGlasgow, Scotland, GB
Full-time

Role: Platform Engineer, Location: Glasgow (Onsite all 5 days).Role Type: Contract (Inside IR35).Envoy Proxy (xDS/ADS, ext_authz, HTTP/2, gRPC, WebSocket) and/or Kong API Gateway (plugin developmen... Show more

 • Promoted

Remote IT Systems Engineer - MBSE & Agile

Leidos Innovations UK LimitedScotland, GB
Remote
Full-time

Leidos Innovations UK Limited is seeking a Systems Engineer with a strong background in IT and software systems.The role involves maintaining key engineering documents and collaborating with teams ... Show more

 • Promoted

Senior Software Engineer (Linux, React, IaC, Observability)

Java Script WorksGlasgow, Scotland, United Kingdom
Full-time

Familiarity with Linux, including Bash scripting and basic system administration.Familiarity with Javascript development and at least one Javascript framework.Experience writing web backend.Profici... Show more

 • Promoted

Hybrid UCaaS Deployment Engineer

GammaGlasgow, Scotland, GB
Full-time

Gamma is seeking an UCaaS Implementation Engineer to design and implement standard UCaaS solutions across our Cloud Platforms, including Microsoft Teams Operator Connect and Direct Routing, Horizon... Show more

 • Promoted

Senior Software Engineer (AIX)

OVO GroupGlasgow, Scotland, GB
Full-time

Unfortunately we are unable to offer sponsorship for this role.Top 3 qualities for this role:.Communication, Delivery Expertise, Amplification.Depending on the needs of your business area, we expec... Show more

 • Promoted

Senior Systems Engineer – Automation & Low-Latency Infra

ProvnScotland, GB
Full-time

A leading technology solutions provider in the United Kingdom is seeking a Senior System Engineer.The ideal candidate will have extensive experience in Linux administration, automation (especially ... Show more

 • Promoted

Senior Software Engineer - Space Reliability

UK Space JobsGlasgow, Scotland, GB
Full-time

Name Your Satellite Program (NYSP).Spire is making a fundamental shift in how it operates its constellation.We are moving from a model where trained operators watch dashboards and escalate to exper... Show more

 • Promoted

Senior Software Engineer - Space Reliability

SpireGlasgow, Scotland, GB
Full-time

Spire is making a fundamental shift in how it operates its constellation.We are moving from a model where trained operators watch dashboards and upscale to experts, to one where the system is fully... Show more

 • Promoted

Site Reliability Engineer

Paritas RecruitmentGlasgow, Scotland, GB
Full-time

AWS Site Reliability Engineer (Data Platform) – Contract.Contract Length: February 2026 – January 2027.We are recruiting an AWS Site Reliability Engineer (SRE) to support a cloud-native data platfo... Show more

 • Promoted

Remote‑First Senior Full‑Stack Tech Lead (C#,.NET)

Areti Group | B CorpScotland, GB
Remote
Full-time

A leading software scale-up in Scotland is seeking to hire five exceptional Senior Software Engineers / Tech Leads to contribute to their AI-powered platform.The role involves hands-on leadership, ... Show more

 • Promoted

Cloud Engineer – ELK Specialist

KBC Technologies GroupScotland, GB
Full-time

Get AI-powered advice on this job and more exclusive features.Direct message the job poster from KBC Technologies Group.Global Talent Acquisition Specialist | US, Australia, Canada, India & EMEA Re... Show more

 • Promoted

Senior DevOps/MLOps Engineer — Remote, GCP

Kodamai LimitedGlasgow, Scotland, GB
Remote
Full-time

Engineering Glasgow / Remote Consultant Contract.As a Senior DevOps/MLOps Engineer at Kodamai, you will own the end-to-end infrastructure strategy, from setting up and maintaining our GCP-based int... Show more

 • Promoted

Platform Engineer

DNS INFO LTDGlasgow, Scotland, GB
Full-time

BS/MS degree in Computer Science, related technical field, or equivalent with 8+ years of industry experience.Envoy Proxy (xDS/ADS, ext_authz, HTTP/2, gRPC, WebSocket) and/or Kong API Gateway (plug... Show more

 • Promoted

Senior Oracle + Apex Developer / Technical Operations Lead

Behan Services LtdScotland, GB
Full-time

Senior Oracle + Apex Developer / Technical Operations Lead.Location: Remote working, based Central Belt, Scotland for occasional client/team meetings.Working Pattern: 5 days per week (flexible hour... Show more