Description
We are currently seeking an experienced professional to join our team in the role of Associate Director, Software Engineering.
Principal responsibilities:
- Lead complex troubleshooting and root cause analysis efforts for incidents impacting production, driving rapid resolution and long-term prevention.
- Design, architect, and enhance scalable, highly available, and secure infrastructure leveraging cloud, container, and orchestration technologies (e.g., AWS/GCP/Azure, Kubernetes, Docker).
- Champion the adoption and refinement of SRE practices,defining and measuring SLIs/SLOs, establishing error budgets, and automating operational processes to minimize toil.
- Develop and maintain comprehensive monitoring, logging, and alerting systems using modern observability tools (e.g., Prometheus, Grafana, ELK, Datadog, Splunk).
- Drive advancements in deployment automation, CI/CD pipelines, infrastructure-as-code (Terraform, Ansible, Helm, etc.), and configuration management.
- Guide, mentor, and coach junior SREs and engineers, fostering a culture of knowledge sharing, reliability, and continuous learning.
- Collaborate with software development, QA, product, and operations teams to embed reliability, scalability, and security considerations throughout the software development lifecycle.
- Participate in and lead on-call rotations, review and improve incident response processes, and perform blameless postmortems.
Knowledge & Experience/Qualifications:
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent experience.
- At least 10 years of hands-on experience in IT, with significant experience focused in SRE, DevOps, Production Support, or related roles.
- Advanced hands-on expertise with containers (Docker, Kubernetes), cloud platforms (AWS, GCP, Azure), and orchestration technologies.
- Deep experience with monitoring, log management, and observability platforms.
- Fluent in at least one programming or scripting language (Python, Bash, Go, etc.).
- Experience implementing and maturing SRE principles at the organisational level.
- Excellent problem-solving, analytical, communication, and mentoring skills.
- Proven ability in high-availability, mission-critical, and/or 24x7 operational environments.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://portal.careers.hsbc.com/careers/job/563774611857128