HirePortal

DevOps / Site Reliability Engineer

  • AgileEngine
  • Bogotá, Colombia, Colombia
  • COP 200,000,000 – COP 300,000,000

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.

WHAT YOU WILL DO

  • Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).

  • Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.

  • Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.

  • Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.

  • Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.

  • Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.

  • Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.

  • Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.

MUST HAVES

5+ years of experience .

  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles .

  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting .

  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment .

  • Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.

  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.

  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures.

  • Fully autonomous.

  • Drives the architecture of complex automated runbooks and mentors Middle-level SREs.

  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz .

  • Prior experience building platforms subject to strict financial compliance standards ( PCI-DSS, SOC2 ).

  • Upper-intermediate English level.

NICE TO HAVES

  • PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.

  • ServiceNow — familiarity with incident, problem, and change management workflows and reporting.

PERKS AND BENEFITS

Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget

Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews

Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm

Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands

Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized

Well-being & support : access local well-being programs and people-focused support tailored to your location

Skills

  • Terraform
  • Azure
  • AWS
  • GCP
  • CI/CD
  • Wiz
  • Incident Management

Related jobs

AgileEngineApply for this job