Site Reliability Engineer (Guardicore AI Platform)
Descripción del puesto
Empresa: Akamai
Provincia: Madrid
Población:
Descripción: Job Description
Are you passionate about building reliable and scalable cloud platforms?
Would you enjoy supporting cybersecurity products through advanced data and AI capabilities?
Join our Akamai Guardicore Data and AI Platform Team!
The team develops and manages a cloud-native Data and AI Platform enabling analytics, insights, and AI-driven features for Akamai Guardicore Segmentation. It processes extensive security and contextual data, supporting customers in understanding environments, minimizing risks, and preventing threat proliferation.
Partner with the best
As a Site Reliability Engineer, you will take technical ownership of the reliability, availability, performance, and operational readiness of the Guardicore Data and AI Platform.
You will be responsible for:
- Operating secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling.
- Enhancing platform reliability, observability, security, performance, and cost efficiency.
- Providing guidance to engineers and developers to increase confidence that their services are performing as expected.
- Leading complex production investigations and driving long-term improvements.
- Leverage LLMs and AI-driven automation to auto-remediate incidents and streamline operations.
- Partner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problems.
- Participating in on-call rotations, guiding restoration and repair of service-impacting issues.
Do what you love
To be successful in this role you will:
- 3+ years of experience in SRE, DevOps, or Platform Engineering, with a proven track record of mastering and troubleshooting complex system architectures.
- Demonstrate ability to design and implement a comprehensive monitoring and observability strategy using tools like Prometheus and Grafana.
- Have production experience with Kubernetes, Docker, Helm, and third-party clouds (GCP, Azure, Linode, AWS) on Linux-based infrastructure.
- Have exceptional troubleshooting and problem-solving skills across network, system, applications, and database layers.
- Have experience with GitOps, CI/CD, and Infrastructure as Code.
- Have scripting and programming proficiency in Python, Go, and Bash.
- Leverage AI tools in daily operational tasks and actively propose initiatives to improve platform automation.
- Demonstrate technical leadership and ownership in driving cross-team initiatives, defining tools, and building foundational frameworks.
About us
At Akamai, we make life better for billions of people, trillions of times a day.
Whether you´re streaming live events, scrolling social media, watching your favorite series, or managing your savings, we´re the engine behind the scenes. We provide the world´s most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the int...
Tecnologías: Kubernetes, Prometheus, Grafana, Docker, Helm, GCP,
Tipo de Contrato: Indefinido
Salario: Sin especificar
Experiencia: 3 años
Funciones: DevOps - Técnico de Sistemas
Descubre más: https://www.tecnoempleo.com/site-reliability-engineer-guardicore-ai-platform-a/kubernetes-prometheus/rf-0b8a1607c29f73984646
Información complementaria
- Categoría
- Non classé
- Referencia
- —
- Fuente
- Tecnoempleo Madrid
- Publicado el
- 26/08/2026