Talent.com
Squarepoint Capital
Platform ULL - Colo - ReliabilitySquarepoint Capital • London, United Kingdom
Platform ULL - Colo - Reliability

Platform ULL - Colo - Reliability

Squarepoint Capital • London, United Kingdom
30+ days ago
Job type
  • Full-time
Job description

Position Overview:

Squarepoint is looking for a talented and highly motivated Ultra Low Latency Platform Engineer to provide solutions across Squarepoint’s global colocation (COLOs) estate consisting of 400+ servers across 30 global sites. The candidate will be responsible for project delivery, support escalations, monitoring, automation, security, documentation, and capacity management for Squarepoint’s low latency infrastructure. This will involve collaborating with our business partners, application owners, clients, vendors, and internal teams (SRE, Network, Application Support and Application Development, Quants, etc.) to deliver end to end solutions in a timely manner.

  • Manage systems efficiently at scale through standardization, automation, testing, and in-depth monitoring
  • Enforce development standards for source control, testing, and continuous integration for infrastructure, OS, patches, and configuration management
  • Manage a distributed compute environment and multiple petabyte-scale storage systems
  • Install, manage, and monitor the Linux operating system (RHEL based)
  • Troubleshoot complex hardware and software issues throughout the Squarepoint technology stack
  • Create self-healing systems and automated recovery processes
  • Respond to system incidents and participate in on-call rotations
  • Conduct root cause analysis of incidents and outages
  • Reduce operational toil through the development of user-driven automated workflows
  • Work with business owners to regularly re-prioritize the book of work, while delivering both tactical and long-term objectives

Required Qualifications:

  • 5+ years of experience working with Linux (RHEL/CentOS/Rocky preferred) in a large complex or niche environment with the following areas of focus: operations, systems engineering and systems performance.
  • Server Management and Support: HP, SuperMicro, Dell, various overclock servers.
  • Experience with Low latency network interfaces and kernel bypass (configuration and optimization): Solarflare with onload, Mellanox with VMA.
  • Experience with build and configuration management tools, specifically Chef or Ansible.
  • Experience with observability tools, specifically Grafana and Prometheus.
  • Highly motivated and a keen eye for scripting and automation in Python, Ruby, and Bash.
  • In depth knowledge of server network stack configuration, tuning and troubleshooting including TCP, UDP(unicast/multicast), NTP, PTP, wireshark/tshark
  • Critical thinking and problem-solving skills to tackle troubleshooting the unknown, glitches and the obscure.
  • Good understanding of trading venues such as Nasdaq, LSE, Euronext etc.
  • Degree in Engineering, Computer Science or related experience.

The minimum base salary for this role is $120,000 if located in New York. This expectation is based on available information at the time of posting. This role may be eligible for discretionary bonuses, which could constitute a significant portion of total compensation. This role may also be eligible for benefits, such as health, dental, and other wellness plans, as well as 401(k) contributions. Successful candidates’ compensation and benefits will be determined in consideration of various factors.

Create a job alert for this search

Platform ULL - Colo - Reliability • London, United Kingdom

Similar jobs

Senior Cloud Reliability & Platform Engineer

Carta HealthcareGreater London, England, GB
Full-time

Carta Healthcare seeks a Senior Site Reliability Engineer to build and scale internal platform services, ensuring reliability and performance for applications.You will design monitoring and inciden... Show more

 • Promoted

Senior Cloud Reliability Engineer

CartaGreater London, England, GB
Full-time

Carta is seeking a Senior Site Reliability Engineer to build and scale internal platform services, ensuring reliability and performance across Carta's applications.You will design monitoring and in... Show more

 • Promoted

Azure Platform & Reliability Engineer

Hitachi Solutions, Ltd.Greater London, England, GB
Full-time

Hitachi Solutions Europe is seeking an Azure Infrastructure and Platform Engineer to provide L2/L3 support for Azure-hosted solutions.You will operate and support Azure infrastructure, networking, ... Show more

 • Promoted

Director of Platform Engineering: Scale, Reliability & Product

YouLend LimitedGreater London, England, GB
Full-time

YouLend Limited is looking for a Director of Platform Engineering in London to lead and empower the Platform team.This role involves strategic leadership and operational excellence to ensure that p... Show more

 • Promoted

Principal SRE: Cloud Reliability Lead (Azure/AWS)

FourthGreater London, England, GB
Full-time

Fourth is seeking an experienced Principal Site Reliability Engineer to accelerate cloud adoption and build automated, highly reliable infrastructure pipelines across Azure and AWS.You will collabo... Show more

 • Promoted

Hybrid Platform Operations & Reliability Director

Publicis Groupe Holdings B.VGreater London, England, GB
Full-time

A leading global marketing technology company is seeking a System and Platform Operations Director to provide technical leadership for the support and stability of production systems.The role focus... Show more

 • Promoted

Senior DevOps Engineer: Platform & Reliability Lead

Breath HRGreater London, England, GB
Full-time

Auriga is seeking a Senior DevOps Engineer to architect, lead and scale the platform engineering and operational backbone of our solutions.You will define and maintain roadmaps for infrastructure, ... Show more

 • Promoted

Senior Site Reliability Engineer (Cloud Platform)

Salve.Inno ConsultingLondon, England, United Kingdom
Full-time

We're looking for a Senior Site Reliability Engineer to help build, operate, and continuously improve a highly available cloud platform supporting mission-critical production services.In this role,... Show more

Front-Office SRE Lead: Observability & AI-Driven Reliability

JPMorgan Chase & Co.Greater London, England, GB
Full-time

London is seeking a Lead Site Reliability Engineer to shape next‑gen SRE patterns, observability, and reliability across globally distributed trading systems.You will partner with front‑office trad... Show more

 • Promoted

Sr Site Reliability Engineer I

Accreditation Council for Graduate Medical EducationGreater London, England, GB
Full-time

At Axon, we’re on a mission to Protect Life.We’re explorers, pursuing society’s most critical safety and justice issues with our ecosystem of devices and cloud software.Like our products, we work b... Show more

 • Promoted

Platform Engineer: Scale, Reliability & Impact

AshbyGreater London, England, GB
Full-time

Ashby is hiring a Platform Engineer in Greater London.This role involves enhancing systems for a rapidly growing customer base, focusing on infrastructure-as-code, and collaborating closely with pr... Show more

 • Promoted

Director, Cloud Platform & Reliability

Sanity CMSGreater London, England, GB
Full-time

Sanity CMS is seeking a Director of Cloud Infrastructure to set the stage for our next phase of growth.You will lead crucial infrastructure projects, ensuring reliability, observability, and deploy... Show more

 • Promoted

Cloud Platform Reliability Engineer | Kubernetes & IaC

Talenzon groupGreater London, England, GB
Full-time

Talenzon group seeks a Platform Reliability Engineer to enhance reliability, scalability, and performance across modern cloud-native platforms in London.The role is focused on designing cloud infra... Show more

 • Promoted

Remote Principal SRE - Healthcare Platform Reliability Lead

MediSolutionGreater London, England, GB
Remote
Full-time

MediSolution in London is seeking a Site Reliability Engineer (SRE) to ensure the reliability of healthcare platforms.The candidate will lead efforts in automating operations and improving service ... Show more

 • Promoted

Platform Lead Engineer: Scale Reliability & Impact (Hybrid)

Trint Ltd.Greater London, England, GB
Full-time

A technology company specializing in media solutions is seeking a Lead Software Engineer to enhance platform infrastructure and ensure reliability.This role requires 5+ years of experience in softw... Show more

 • Promoted

Director, Platform Operations & Reliability

EpsilonGreater London, England, GB
Full-time

A subsidiary of a global technology company is seeking a System and Platform Operations Director to manage the reliability and support of production systems.This leadership role involves orchestrat... Show more

 • Promoted

Senior Cloud Reliability Engineer - Scale AI Infra & CI/CD

NebiusGreater London, England, GB
Full-time

Nebius is building a full‑stack AI cloud platform and continuously scales its services in a fast‑moving environment.You will join a team delivering high‑reliability infrastructure for data and mode... Show more

 • Promoted

Site Reliability Engineer - Comcast Technology Solutions

Blueface LtdGreater London, England, GB
Permanent

Comcast brings together the best in media and technology.We drive innovation to create the world's best entertainment and online experiences.As a Fortune 50 leader, we set the pace in a variety of ... Show more

 • Promoted

Azure Platform Engineer - Infra & Reliability

Hitachi SolutionsGreater London, England, GB
Full-time

Hitachi Solutions Europe is seeking an Azure Infrastructure and Platform Engineer to provide L2/L3 support for Azure-hosted solutions.The role focuses on operating and supporting Azure infrastructu... Show more

 • Promoted

DevOps Engineer: CI/CD & Platform Reliability

Sensor TowerGreater London, England, GB
Full-time

Sensor Tower is seeking a skilled DevOps engineer with a background in software development to enhance team productivity through effective CI/CD stack management.This role involves collaborating cl... Show more