SRE

name

📍 Remote🌐 Remote🕐 1mo ago🔗 himalayas

Job Description

### Are you ready to unlock intelligence? If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others. ### **The Mission** The Site Reliability Engineering team is the backbone of [Kong](https://himalayas.app/companies/kong)'s cloud services, responsible for architecting and operating the large-scale infrastructure that powers our customers' most critical applications. Our mission is to achieve world-class reliability and performance, enabling our product engineering teams to ship features with velocity and confidence. We are the guardians of uptime and the champions of developer delight. ### **What You’ll Do** * Build and maintain our core infrastructure as code using tools like Terraform and Ansible. * Implement robust monitoring, logging, and alerting systems to ensure our services meet and exceed 99.99% uptime. * Resolve production incidents through systematic debugging, and drive the blameless post-mortem process to prevent recurrence. * Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams. * Collaborate with developers to embed reliability and scalability best practices directly into the application lifecycle. * Contribute to our capacity planning, disaster recovery drills, and security hardening processes. * Participate in a fair and sustainable on-call rotation to ensure our platform is always available. ### **What You’ll Bring** * Experience operating production workloads on a major cloud provider (AWS, GCP, Azure). * Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash. * Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes). * Knowledge of Infrastructure as Code principles and tools (Terraform is a plus). * Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins). * An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK). ### About [Kong](https://himalayas.app/companies/kong): [Kong](https://himalayas.app/companies/kong) Inc., a leading developer of API and AI connectivity technologies, is building the infrastructure that powers the agentic era. Trusted by the Fortune 500 and startups alike, [Kong](https://himalayas.app/companies/kong)'s unified API and AI platform, [Kong](https://himalayas.app/companies/kong) Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models. For more information, visit . Originally posted on [Himalayas](https://himalayas.app)
SRE at name | MergeJobs