Infrastructure Engineer
- Location
- New York, NY
- Work type
- Full Time · On-site
- Posted
- 2026-08-30
Job description
We are seeking an experienced Infrastructure Engineer with expertise in cloud capacity management to join our team. This role focuses on leading capacity planning and optimization for a high-demand Azure public cloud environment. The ideal candidate will ensure the availability of resilient, scalable capacity across compute, storage, network, and platform services. They will collaborate with site reliability teams to meet performance and reliability targets, while maintaining compliance with rigorous program controls. This position offers an opportunity to shape critical cloud infrastructure capabilities in a regulated environment, emphasizing evidence-based practices and continuous monitoring.
Key Responsibilities
Own and manage the end-to-end Capacity Management operating model for Azure services involved in high-criticality projects, including planning, modeling, forecasting, monitoring, tuning, and governance.
Ensure sufficient capacity and engineered buffers to meet service-level agreements (SLAs), recovery time objectives (RTOs), recovery point objectives (RPOs), and contractual or regulatory requirements, with particular attention to region-specific restrictions and ongoing monitoring.
Partner with site reliability engineers to implement capacity practices via infrastructure as code (IaC), gated change controls, performance baselines, autoscaling strategies, and resilience patterns.
Contribute to documentation and compliance evidence such as system security plans, control narratives, corrective action plans, and continuous monitoring artifacts.
Develop and maintain service-level capacity models across various Azure components, establishing buffer standards based on service criticality and validating against demand patterns and failover scenarios.
Design and tune autoscaling policies, setting guardrails on quotas and throttling to ensure performance stability.
Conduct baseline and trend analyses of utilization, throughput, and performance metrics, translating insights into tuning actions, reservations, savings plans, and architectural improvements.
Forecast future demand based on product roadmaps and business growth, translating forecasts into capacity plans and procurement strategies.
Participate in change review processes, ensuring capacity and security impacts are properly assessed and documented.
Manage cryptographic mechanisms and cryptography-related configurations under change control, maintaining versioned inventories and validation compliance.
Oversee external services supporting capacity, confirming they meet required standards and conduct ongoing oversight.
Enforce region-restriction policies for processing, storage, backups, and disaster recovery specific to high-impact systems.
Balance performance, resilience, and cost-efficiency through resource rightsizing, tiering, and scheduled scaling, ensuring proactive capacity adjustments.
Perform criticality analysis to prioritize capacity needs, aligning backup, monitoring, and security policies accordingly.
Validate disaster recovery (DR) capacity and ensure buffers are maintained for failover scenarios without impacting steady-state operations.
Define, measure, and report key capacity KPIs, including utilization, saturation, headroom, runway duration, scaling effectiveness, quota use, DR readiness, and cost-performance metrics.
Prepare dashboards and reports to monitor program compliance, support audits, and inform strategic decision-making.