Back to Neptunez Singapore Pte. Ltd. jobs
N
Site Reliability Engineer(Kubernetes, Dynatrace, Splunk, Docker, ELK, Terraform, JMeter, Prometheus, Grafana)
Islandwide, Singapore
ContractInformation TechnologyJob Description
Responsibilities
- Lead Site Reliability Engineering (SRE) activities to ensure high availability, stability, and reliability of enterprise applications and distributed systems.
- Provide technical leadership for L2/L3 production support, incident management, and operational excellence across mission-critical environments.
- Investigate and resolve complex production issues by analyzing application behavior, infrastructure, databases, and system performance.
- Monitor application and platform health using Dynatrace, Splunk, Prometheus, Grafana, and ELK, while building dashboards, alerts, and operational visibility.
- Perform root cause analysis (RCA) for critical incidents and drive permanent corrective and preventive actions.
- Troubleshoot Java/Spring Boot applications, APIs, Linux-based systems, and SQL-related issues to ensure rapid incident resolution.
- Support cloud-native environments across AWS and GCP, including Kubernetes-based container platforms and cloud infrastructure components.
- Develop and maintain automation solutions using Shell scripting, Python, and Terraform to improve operational efficiency and reduce manual effort.
- Support CI/CD and deployment activities using Azure DevOps, Jenkins, and modern DevOps practices.
- Collaborate with development, infrastructure, and operations teams to improve platform reliability, performance, and service availability.
- Mentor junior engineers, provide technical guidance, and promote SRE best practices, operational standards, and knowledge sharing.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 12+ years of experience in Site Reliability Engineering (SRE), Production Support, or DevOps within enterprise environments.
- Strong experience in L2/L3 Production Support, Incident Management, RCA, and supporting distributed systems.
- Hands-on expertise with Dynatrace, Splunk, Prometheus, Grafana, and ELK for monitoring, observability, log analysis, and alerting.
- Strong proficiency in Linux administration, troubleshooting, and automation using Shell scripting or Python.
- Experience troubleshooting Java (Java 8+), Spring Boot applications, REST APIs, and application performance issues.
- Strong knowledge of SQL, database troubleshooting, query optimization, and performance tuning.
- Hands-on experience with AWS, GCP, Kubernetes, Docker, and cloud-native infrastructure.
- Experience implementing Infrastructure as Code (IaC) using Terraform.
- Good understanding of CI/CD pipelines, build and release management using Azure DevOps and Jenkins.
- Strong analytical, troubleshooting, and problem-solving skills with the ability to manage high-severity production incidents.
- Excellent communication, stakeholder management, and technical leadership skills, with experience mentoring engineering teams.
About Neptunez Singapore Pte. Ltd.
First seen: August 3, 2026
Last updated: August 5, 2026