About VOIS (Vodafone Intelligent Solutions)
VOIS is a strategic arm of Vodafone Group Plc, delivering intelligent solutions through Talent, Technology & Transformation. We empower Vodafone's global operations with cutting-edge digital platforms, cloud infrastructure, and automation capabilities.
Role Overview: Site Reliability Engineer
We are seeking an experienced Site Reliability Engineer (SRE) to join our dynamic engineering team. You will be responsible for ensuring the reliability, scalability, and performance of our mission-critical systems and cloud-native platforms. This role bridges software engineering and operations, focusing on automation, observability, and incident response to maintain high availability for Vodafone's global services.
Key Responsibilities
- Design, build, and maintain highly available, scalable, and resilient cloud infrastructure on AWS, GCP, or hybrid environments.
- Implement and manage Kubernetes clusters, container orchestration, and service mesh technologies (Istio, Linkerd).
- Develop automation frameworks using Python, Go, or Bash for provisioning, configuration management, and self-healing systems.
- Establish comprehensive observability stacks: monitoring (Prometheus, Grafana), logging (ELK/EFK), distributed tracing (Jaeger, Zipkin), and alerting.
- Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Lead incident response, conduct blameless postmortems, and drive preventive actions to eliminate toil.
- Architect CI/CD pipelines (GitLab CI, Jenkins, ArgoCD) for safe, rapid, and reliable deployments.
- Implement Infrastructure as Code (IaC) using Terraform, Helm, or Crossplane.
- Collaborate with development teams on capacity planning, performance tuning, chaos engineering, and reliability patterns (circuit breakers, retries, bulkheads).
- Participate in on-call rotations and contribute to a culture of shared ownership and continuous improvement.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer.
- Deep hands-on expertise with Linux/Unix internals, networking (TCP/IP, DNS, HTTP/HTTPS, gRPC), and security hardening.
- Proven experience managing Kubernetes in production (EKS, GKE, AKS, or self-managed).
- Strong programming skills in Python and/or Go for tooling and automation.
- Extensive experience with Infrastructure as Code (Terraform preferred) and configuration management (Ansible).
- Solid understanding of cloud provider services (AWS: EC2, RDS, S3, IAM, VPC; GCP: Compute Engine, Cloud SQL, GCS, IAM, VPC).
- Proficiency with monitoring and observability tools: Prometheus, Grafana, Loki, Tempo, Datadog, or similar.
- Experience designing and operating CI/CD pipelines for microservices architectures.
- Familiarity with service mesh, API gateways, and GitOps workflows.
Preferred Skills & Keywords
- Certifications: CKA, CKAD, AWS Solutions Architect, GCP Professional Cloud Architect.
- Experience with Apache Kafka, Redis, PostgreSQL, Cassandra at scale.
- Knowledge of chaos engineering tools (Litmus, Chaos Mesh, Gremlin).
- Background in telecommunications or large-scale distributed systems.
- Contributions to open-source projects in the cloud-native ecosystem.
- Strong communication skills and ability to work in a global, cross-functional team.
Why Join VOIS?
Work at the heart of Vodafone's digital transformation. Access to cutting-edge technology, global scale challenges, continuous learning opportunities, and a culture that values innovation, diversity, and engineering excellence. Competitive compensation, comprehensive benefits, and flexible working arrangements.