Description
As a Senior Site Reliability Engineer (Capacity) - Platform Infrastructure, you will play a crucial role in managing and optimizing compute resources, ensuring that Elastic Cloud Hosted and Serverless workloads can scale seamlessly.
You will collaborate closely with control plane and cross-functional platform engineering teams, addressing cloud scaling and resource allocation challenges.
Responsibilities:
- Assess current and future capacity requirements based on workload demands to ensure seamless scaling of resources.
- Develop and maintain accurate capacity models that predict resource needs and align with business objectives.
- Collaborate with teams to implement proactive measures that prevent capacity shortages and bottlenecks.
- Implement effective strategies for optimizing resource usage across cloud environments.
- Analyze capacity metrics and trends to guide effective resource allocation decisions.
- Operate an autoscaling framework that accommodates various customer workloads seamlessly.
Requirements:
- 5+ years with cloud infrastructure and capacity management
- Knowledge of performance monitoring and optimization techniques
- Understanding of cloud scaling challenges and solutions
- Proficiency with incident investigation and troubleshooting processes
- Experience with compute auto-scaling processes and capacity reservations across major CSPs
- Solid software and platform engineering background
- Worked with major cloud service providers and navigated compute capacity scaling issues
Benefits:
- Competitive pay
- Health coverage for you and your family
- Flexible locations and schedules
- Generous number of vacation days
- Matching up to $2000 for financial donations and service
- Up to 40 hours each year to use toward volunteer projects
- Minimum of 16 weeks of parental leave
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/elastic/jobs/8155560