This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a ASIC Architect, Principal (AI Inference) based in the United States. This is a highly senior architecture role focused on defining the next generation of AI inference accelerators for low-latency, high-throughput workloads. You will shape hardware architectures capable of supporting rapidly evolving frontier models, from transformers and MoE systems to sparse and long-context workloads. The role sits at the intersection of machine learning, systems architecture, performance modeling, and practical silicon implementation. You will work closely with architecture, software, compiler, runtime, firmware, ASIC engineering, and ML research teams to drive hardware/software co-design. Your analysis and architectural decisions will directly influence product specifications, roadmaps, and long-term scalability. This is an opportunity to take end-to-end ownership of complex architectural challenges and help establish platforms designed to remain competitive as AI workloads evolve. Accountabilities: Own architectural exploration across emerging AI workloads, including transformer models, mixture-of-experts architectures, sparse models, long-context inference, reasoning models, and other rapidly developing approaches. Evaluate emerging inference techniques such as speculative decoding, continuous batching, evolving KV-cache architectures, sub-quadratic and linear attention, and state-space models, assessing their implications for future accelerator designs. Investigate memory hierarchies, communication architectures, and interconnect strategies to determine how next-generation accelerators can achieve strong performance, scalability, efficiency, and cost characteristics. Develop analytical performance models and simulation frameworks to evaluate competing architectural approaches and quantify latency, throughput, bandwidth, memory and compute utilization, scaling efficiency, power, area, and cost. Partner closely with compiler, runtime, system software, firmware, ASIC engineering, and machine learning teams to ensure hardware architecture aligns with real-world workload requirements and software capabilities. Define ASIC architecture specifications, subsystem interfaces, and architectural documentation while evaluating tradeoffs across competing technical directions and translating findings into product roadmap recommendations. Continuously monitor new LLMs, reasoning models, research publications, academic developments, hardware competitors, and emerging algorithms to identify opportunities and ensure architectural decisions remain aligned with the evolving AI landscape. Leverage AI-assisted engineering techniques to accelerate research, performance analysis, architectural exploration, simulation development, and technical documentation. Help establish a scalable architectural direction that bridges hardware and software and supports AI inference products designed to remain competitive across multiple generations. Requirements: Bring 15+ years of experience in ASIC architecture, with deep expertise designing AI accelerators and high-performance SoCs. Demonstrate hands-on experience developing performance models for memory systems, interconnects, and other critical hardware subsystems. Possess strong hardware/software co-design capabilities and the ability to connect architectural decisions with compiler, runtime, systems, and workload requirements. Have excellent written and verbal communication skills, with the ability to clearly explain complex architectural tradeoffs and influence technical direction across multidisciplinary teams. Experience with GPU architecture is preferred, particularly where it contributes to a strong understanding of high-performance accelerator design. A solid understanding of ML systems, distributed inference, large-scale model serving, and the practical requirements of modern AI infrastructure is highly desirable. Demonstrate deep knowledge of transformer internals and the architectural implications of evolving model structures and inference techniques. Compiler experience and a track record of research publications are advantageous. Be comfortable operating at a highly senior level, with the technical breadth and judgment to influence long-term architecture, roadmap, and cross-functional engineering decisions. Candidates must currently be authorized to work in the United States. New visa sponsorship is not available, although H-1B transfers may be facilitated for eligible candidates. Benefits: Base salary range of $200,000–$300,000 , with flexibility to exceed the range for candidates whose specialized expertise, experience, or qualifications warrant a different compensation level. Competitive total compensation including equity , with final offers determined based on factors including experience, qualifications, internal equity, and expected impact. Fully company-paid medical, dental, and vision insurance for employees and dependents. Company-paid life and disability insurance , with voluntary options for additional coverage. Access to supplemental hospital, critical illness, and accident insurance . Unlimited paid time off and 13 paid company holidays to support meaningful time away from work. Remote-first work in the United States, with a company-provided computer and home-office setup. 401(k) with company matching , with eligibility beginning on day one. The opportunity to shape foundational AI inference architecture, influence long-term technical direction, and work across frontier AI research and silicon implementation.