ITRS is a global leader in automated and holistic IT observability solutions, dedicated to safeguarding critical applications and enabling innovation across industries. We are seeking a highly skilled and experienced Senior DevOps Engineer to join our dynamic engineering team and play a pivotal role in building, scaling, and maintaining the infrastructure that powers our cutting-edge monitoring platforms.
Role Overview
As a Senior DevOps Engineer at ITRS, you will be responsible for designing, implementing, and managing robust CI/CD pipelines, cloud infrastructure, and automation frameworks that support our high-availability observability solutions. You will work closely with software engineering, QA, and product teams to ensure seamless deployment, monitoring, and reliability of services that protect mission-critical technology for organizations worldwide.
Key Responsibilities
- Design, build, and maintain scalable CI/CD pipelines to accelerate software delivery and improve deployment reliability.
- Manage and optimize cloud infrastructure across platforms such as AWS, Azure, or GCP, ensuring high availability, security, and cost efficiency.
- Implement Infrastructure as Code (IaC) using tools like Terraform, Ansible, or CloudFormation to automate provisioning and configuration management.
- Develop and maintain containerized environments using Docker and orchestration platforms such as Kubernetes.
- Establish comprehensive monitoring, logging, and alerting systems to ensure proactive incident detection and rapid response.
- Collaborate with development teams to improve application performance, scalability, and resilience through DevOps best practices.
- Drive automation initiatives across the software development lifecycle, reducing manual processes and increasing operational efficiency.
- Ensure compliance with security policies, industry standards, and regulatory requirements across all infrastructure and deployment processes.
- Participate in on-call rotations and provide senior-level support for production incidents, performing root cause analysis and implementing preventive measures.
- Mentor junior engineers and contribute to the continuous improvement of DevOps culture, tooling, and processes within the organization.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Minimum of 5 years of professional experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles.
- Strong hands-on experience with cloud platforms (AWS, Azure, or GCP) including compute, networking, storage, and security services.
- Proficiency in Infrastructure as Code tools such as Terraform, Ansible, Puppet, or Chef.
- Expertise in CI/CD tools including Jenkins, GitLab CI, GitHub Actions, or CircleCI.
- Solid experience with containerization and orchestration technologies, particularly Docker and Kubernetes.
- Strong scripting and automation skills in languages such as Python, Bash, or Go.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Deep understanding of networking, security best practices, and compliance frameworks in cloud environments.
- Excellent problem-solving skills with the ability to troubleshoot complex distributed systems under pressure.
Preferred Qualifications
- Experience working in observability, monitoring, or financial technology (FinTech) environments.
- Certifications such as AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator (CKA).
- Familiarity with service mesh technologies (Istio, Linkerd) and microservices architecture.
- Experience with database administration and optimization for high-throughput systems.
- Knowledge of GitOps methodologies and tools such as ArgoCD or Flux.
- Strong communication skills with the ability to document processes and present technical concepts to diverse audiences.
Why Join ITRS?
At ITRS, you will work at the forefront of IT observability, contributing to solutions that protect the critical technology infrastructure of some of the world's most important organizations. We offer a collaborative and innovative work environment, opportunities for professional growth, and the chance to make a tangible impact on global technology resilience.