Description
We are currently seeking an experienced professional to join our team in the role of Consultant Specialist (SRE).
Principal responsibilities:
- Build and manage robust CI/CD pipelines on Google Cloud Platform for APIs, databases, messaging, storage, and compute
- Create and manage Kubernetes workloads, including pod definitions, tagging, labelling, scaling, and auto-scaling
- Support production stability, availability, and reliability of applications and platform services
- Handle production incidents, including triage, impact assessment, service restoration, escalation, and follow-up actions
- Troubleshoot complex cross-platform production issues across cloud infrastructure, Kubernetes, application services, networking, databases, CI/CD pipelines, and integrations
- Perform root cause analysis for incidents and recurring issues, and drive permanent fixes and continuous improvement
- Define and improve SLOs/SLIs, observability, and monitoring capabilities across engineering teams
- Engage early in new features and platform changes to improve quality, operability, and release readiness
- Contribute to automation, performance, and reliability improvements that reduce operational toil
Knowledge & Experience / Qualifications:
- At least 2 years of experience in GCP (Google Cloud Platform) or similar cloud technologies
- Strong experience in microservices and APIs
- Strong experience with DevOps and platform tools such as Jenkins, Helm, Ansible, Docker, Kubernetes, JIRA, and DataDog
- Experience managing source control, branching, merging, tagging, and release versioning using tools such as GitHub
- Experience defining and supporting development, test, release, update, and support processes for DevOps operations
- Strong understanding of CI/CD, automation, cloud identity and access, and Linux-based infrastructure
- Scripting skills in Shell, Python, or Ruby for automation, monitoring, and operational support
- Excellent troubleshooting and analytical problem-solving skills
- Strong incident handling and production support experience, with the ability to manage high-severity service issues calmly and effectively
- Experience in root cause analysis, issue resolution, and driving preventive actions for recurring incidents
- Experience with observability, monitoring, and alerting tools
- Working knowledge of open-source technologies and modern reliability engineering practices
- Awareness of critical Agile principles and experience working in Agile / SCRUM environments
- Familiarity with Agile team management tools such as JIRA and Confluence
- Good communication skills and the ability to work effectively with engineering teams, Product Owners, and stakeholders
- Proactive team player, comfortable working in multi-disciplinary, self-organized teams
- Professional knowledge of English
- Google Professional Cloud DevOps Engineer certification is desirable
What additional skills will be good to have?:
- Experience on banking / HSBC and understanding of how change drives benefits for HSBC, its customers and other stakeholders will be an advantage.
- Experience on HSBC business banking system knowledge will be an advantage.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://portal.careers.hsbc.com/careers/job/563774612107524