Back to Neptunez Singapore Pte. Ltd. jobs
N

Site Reliability Engineer(Kubernetes, Dynatrace, Splunk, Docker, ELK, Terraform, JMeter, Prometheus, Grafana)

Islandwide, Singapore
ContractInformation Technology

Job Description

Responsibilities

  • Lead Site Reliability Engineering (SRE) activities to ensure high availability, stability, and reliability of enterprise applications and distributed systems.
  • Provide technical leadership for L2/L3 production support, incident management, and operational excellence across mission-critical environments.
  • Investigate and resolve complex production issues by analyzing application behavior, infrastructure, databases, and system performance.
  • Monitor application and platform health using Dynatrace, Splunk, Prometheus, Grafana, and ELK, while building dashboards, alerts, and operational visibility.
  • Perform root cause analysis (RCA) for critical incidents and drive permanent corrective and preventive actions.
  • Troubleshoot Java/Spring Boot applications, APIs, Linux-based systems, and SQL-related issues to ensure rapid incident resolution.
  • Support cloud-native environments across AWS and GCP, including Kubernetes-based container platforms and cloud infrastructure components.
  • Develop and maintain automation solutions using Shell scripting, Python, and Terraform to improve operational efficiency and reduce manual effort.
  • Support CI/CD and deployment activities using Azure DevOps, Jenkins, and modern DevOps practices.
  • Collaborate with development, infrastructure, and operations teams to improve platform reliability, performance, and service availability.
  • Mentor junior engineers, provide technical guidance, and promote SRE best practices, operational standards, and knowledge sharing.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 12+ years of experience in Site Reliability Engineering (SRE), Production Support, or DevOps within enterprise environments.
  • Strong experience in L2/L3 Production Support, Incident Management, RCA, and supporting distributed systems.
  • Hands-on expertise with Dynatrace, Splunk, Prometheus, Grafana, and ELK for monitoring, observability, log analysis, and alerting.
  • Strong proficiency in Linux administration, troubleshooting, and automation using Shell scripting or Python.
  • Experience troubleshooting Java (Java 8+), Spring Boot applications, REST APIs, and application performance issues.
  • Strong knowledge of SQL, database troubleshooting, query optimization, and performance tuning.
  • Hands-on experience with AWS, GCP, Kubernetes, Docker, and cloud-native infrastructure.
  • Experience implementing Infrastructure as Code (IaC) using Terraform.
  • Good understanding of CI/CD pipelines, build and release management using Azure DevOps and Jenkins.
  • Strong analytical, troubleshooting, and problem-solving skills with the ability to manage high-severity production incidents.
  • Excellent communication, stakeholder management, and technical leadership skills, with experience mentoring engineering teams.

About Neptunez Singapore Pte. Ltd.

First seen: August 3, 2026
Last updated: August 5, 2026