Key
• Ensure the stability, performance, and fault tolerance of production systems.
• Develop and maintain infrastructure automation and observability tools.
• Monitor system health, respond to incidents, and perform root cause analysis (RCA).
• Collaborate with development teams to improve scalability and reliability of services.
• Define and manage SLIs, SLOs, and Error Budgets.
• Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
• Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
• Integrate and monitor external vendor systems, collaborating with vendor technical support when needed.
Requirements
Key
• Fluent Russian, English B1+ (comfortable with technical documentation).
• 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
• Strong understanding of observability principles (metrics, logs, traces).
• Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
• Experience working with AWS and CI/CD tools.
• Practical knowledge of SLI/SLO/Error Budget frameworks.
• Experience leading and documenting incidents and post-mortems.
• Scripting skills for automation (Python, Bash, or Go).
• Solid understanding of distributed systems and networking fundamentals.
• Experience monitoring and supporting mobile applications.
• Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
• Experience working with Kubernetes and containerized environments.
•
Fully remote work format.
• Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
• Opportunity to work in an international team on a new digital product for the Mexican market.
• A data-driven environment where your contributions have a real impact.
Originally posted on Himalayas