Job Description
Job Title: Kubernetes Platform Engineer – Site Reliability Engineering
Location: Buffalo, New York (Hybrid)
Overview:
We are looking for an experienced Kubernetes Platform Engineer to join our Site Reliability Engineering (SRE) team, supporting high-availability payment services on a large-scale internal Kubernetes environment. This role is suited for someone passionate about maintaining service uptime, automating operational tasks, and improving system resilience. You’ll be part of a global operations team ensuring 24/7 support for critical production workloads.
Key Responsibilities:
- Provide Tier 1 through Tier 3 support for payment services hosted on Kubernetes clusters.
- Monitor infrastructure and application health using tools like Datadog, Splunk, Prometheus, and Grafana.
- Respond promptly to alerts, resolve production incidents, and maintain service-level objectives (SLOs).
- Participate in scheduled on-call rotations and assist during system maintenance, upgrades, or compliance-related releases.
- Design and implement automation for routine tasks, incident mitigation, and system recovery.
- Collaborate with development and DevOps teams to improve CI/CD pipelines and deployment reliability.
- Maintain documentation including runbooks, standard operating procedures (SOPs), and post-incident reports.
Qualifications:
- Proven experience working with Kubernetes platforms (ideally GKE or Anthos).
- Proficiency in monitoring and logging tools such as Datadog, Prometheus, Grafana, and Splunk.
- Deep understanding of SRE best practices including incident management, root cause analysis, and alert tuning.
- Comfortable operating in production environments under pressure and handling on-call responsibilities.
- Google Cloud Platform (GCP) certification is a strong plus.
- Experience in highly regulated sectors, particularly finance or banking, is advantageous.
Job Tags