Are you excited about building the infrastructure that powers the next generation of AI workloads? As a Senior Software Development Engineer on the EC2 UltraServer Provisioning team, you will lead technical strategy, design, and operation of infrastructure services — including provisioning and availability of AWS Trainium-based AI servers. You will architect large-scale systems, build microservices, and collaborate cross-functionally with capacity management, hardware engineering, and datacenter teams to manage AI/ML infrastructure at scale. Key job responsibilities - Lead the design and development of infrastructure systems that support AI workloads on UltraServers, ensuring reliability and performance at scale. - Own your team's architectural direction for provisioning systems, influencing technical decisions across dependent systems and driving alignment with partner teams. - Collaborate with capacity management, hardware engineering, and datacenter teams to improve provisioning workflows and operate efficiently at scale. - Mentor and coach engineers on the team, driving adoption of engineering best practices and raising the technical bar across the organization. - Investigate complex performance challenges, develop solutions, and publish actionable best practices to internal and external stakeholders. A day in the life You might start your morning reviewing operational metrics for provisioning pipelines, then shift into a design review with hardware engineering partners on an upcoming server configuration. After lunch, you could be deep in code — prototyping a new microservice or refactoring a critical path in the provisioning workflow. Throughout the day, you balance hands-on technical work with mentoring teammates, unblocking cross-team dependencies, and contributing to architectural decisions that shape how AI infrastructure scales. About the team The EC2 UltraServer Provisioning team is responsible for delivering AWS Trainium-based UltraServer infrastructure at scale. We manage end-to-end provisioning workflows from host ingestion through testing, repair, and recovery. Our mission is to ensure that AI/ML customers have reliable, high-performance compute infrastructure ready when they need it. We are a collaborative engineering team that values technical depth, operational excellence, and building systems that grow with customer demand.