Senior Site Reliability Engineer at jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States. This is an opportunity to join a critical AI Hardware SRE team responsible for the reliability of next-generation dedicated AI infrastructure. You will help scale and optimize high-density hardware and software environments across regional data centers. The role combines automation, observability, infrastructure engineering, networking, and real-time incident response. You will build Python-based tooling, infrastructure-as-code utilities, telemetry pipelines, and intelligent monitoring solutions. Your work will directly improve uptime, performance, scalability, and operational efficiency for business-critical systems. You will collaborate with engineering teams, infrastructure vendors, and field technicians to solve complex reliability challenges. This role is ideal for an experienced SRE who thrives on ownership, ambiguity, automation, and production-scale infrastructure.