HirePortal

Senior Site Reliability Engineer

  • AgileEngine
  • Mexico City, Distrito Federal
  • MXN 800,000 – MXN 1,000,000

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE

We are looking for a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry. You will participate in on-call rotations, incident response, and root-cause analysis, automate infrastructure tasks using Python, Bash, or Go, and ensure system health across ESM and ECP platform environments. Experience with service mesh architectures is highly valued for ECP-focused roles.

WHAT YOU WILL DO

  • Provide core system administration for the enterprise infrastructure.

  • Focus on the maintenance, scaling, and operational stability of various on-premise and SaaS-hosted systems.

  • Manage and scale containerized environments using Kubernetes.

  • For ECP-focused roles: Lean heavily into executing monitoring and observability tasks to ensure system health.

  • Participate in on-call rotations, incident response, and root-cause analysis (RCA).

  • Automate repetitive infrastructure tasks using scripting and infrastructure-as-code (IaC).

MUST HAVES

4+ years of infrastructure management experience .

  • Experience working with Kubernetes .

  • Experience with monitoring, observability, and related SRE tasks .

  • Familiarity with Snowflake and OpenTelemetry ecosystems .

  • Strong background in Linux/Unix administration .

  • Proficiency in scripting languages (e.g., Bash, Python, or Go ).

  • Upper-intermediate English level.

NICE TO HAVES

  • For ECP SREs: Experience with Service Mesh architectures is highly ideal.

  • Experience with Infrastructure as Code tools such as Terraform or Ansible.

  • Familiarity with cloud platforms (AWS, GCP, or Azure).

PERKS AND BENEFITS

Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget

Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews

Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm

Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands

Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized

Well-being & support : access local well-being programs and people-focused support tailored to your location

Skills

  • Kubernetes
  • Site Reliability Engineering
  • Monitoring and observability
  • Python
  • Bash
  • Go
  • Incident Response

Related jobs

AgileEngineApply for this job