SRE
name
📍 Remote🌐 Remote🕐 1mo ago🔗 himalayas
Job Description
### Are you ready to unlock intelligence?
If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
### **The Mission**
The Site Reliability Engineering team is the backbone of [Kong](https://himalayas.app/companies/kong)'s cloud services, responsible for architecting and operating the large-scale infrastructure that powers our customers' most critical applications. Our mission is to achieve world-class reliability and performance, enabling our product engineering teams to ship features with velocity and confidence. We are the guardians of uptime and the champions of developer delight.
### **What You’ll Do**
* Build and maintain our core infrastructure as code using tools like Terraform and Ansible.
* Implement robust monitoring, logging, and alerting systems to ensure our services meet and exceed 99.99% uptime.
* Resolve production incidents through systematic debugging, and drive the blameless post-mortem process to prevent recurrence.
* Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams.
* Collaborate with developers to embed reliability and scalability best practices directly into the application lifecycle.
* Contribute to our capacity planning, disaster recovery drills, and security hardening processes.
* Participate in a fair and sustainable on-call rotation to ensure our platform is always available.
### **What You’ll Bring**
* Experience operating production workloads on a major cloud provider (AWS, GCP, Azure).
* Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash.
* Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes).
* Knowledge of Infrastructure as Code principles and tools (Terraform is a plus).
* Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins).
* An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK).
### About [Kong](https://himalayas.app/companies/kong):
[Kong](https://himalayas.app/companies/kong) Inc., a leading developer of API and AI connectivity technologies, is building the infrastructure that powers the agentic era. Trusted by the Fortune 500 and startups alike, [Kong](https://himalayas.app/companies/kong)'s unified API and AI platform, [Kong](https://himalayas.app/companies/kong) Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models. For more information, visit .
Originally posted on [Himalayas](https://himalayas.app)