Senior DevOps Engineer
- Infrastructure
- Full-time
- Tehran
Job Description
Alibaba runs the platform behind a fast-growing, multi-brand travel business. We're hiring a Senior
DevOps Engineer to join a small, high-ownership DevOps/SRE team responsible for our production and
staging environments.
Our platform is fully self-hosted and on-premise — we run our own Kubernetes, GitLab, CI/CD,
distributed storage, databases, and messaging on our own infrastructure. This is a hands-on role for an
engineer who wants to own systems end-to-end across the full stack, and who is energised by
strengthening and modernising a large, established estate. It is not a managed-cloud position.
Responsibilities:
- Operate and evolve our Kubernetes platform — Rancher/RKE and kubespray clusters: version
upgrades, scaling, storage (Longhorn), networking, and workload security. - Operate our stateful services — PostgreSQL, Redis, MongoDB, Elasticsearch, Kafka, RabbitMQ, and
distributed storage: upgrades, backups, failover, capacity planning, and incident response. This is a
significant part of the role. - Automate provisioning, patching, upgrades, and lifecycle across our Linux estate with Ansible and
Terraform/OpenTofu — reducing toil so the team can focus on higher-value work. - Lead modernisation and lifecycle work — operating-system upgrades, patch management, and
standardisation across the estate. - Own CI/CD and GitOps — GitLab CI pipelines and ArgoCD-based delivery.
- Strengthen security and access — secrets management (HashiCorp Vault), SSO/IAM (Keycloak), access
control, and CIS hardening, all of which we are actively maturing. - Own observability — metrics, logs, traces, and alerting for always-up, always-available services.
- Troubleshoot production issues end-to-end and drive them to resolution with minimal customer
impact. - Participate in the on-call rotation and partner with development teams to adopt DevOps best
practices and move applications toward cloud-native patterns on our platform. - Remediate at the root — identify recurring issues and resolve their underlying causes.
Requirements
- Deep Linux/Unix systems administration — the core of this role. 5–8 years operating, tuning,
securing, and troubleshooting Linux servers at scale in production. - Strong TCP/IP and networking fundamentals (DNS, HTTP/S, TLS, routing, load balancing).
- Production expertise with Docker and Kubernetes, including hands-on cluster lifecycle
(Rancher, RKE, and/or kubespray). - Configuration management and IaC: Ansible, plus Terraform/OpenTofu.
- CI/CD with GitLab, and solid Git / GitFlow.
- Strong scripting in at least Bash and Python (Perl or Ruby a bonus).
- Proven experience operating stateful systems in production — relational (PostgreSQL) and
NoSQL/key-value (Redis, MongoDB), plus messaging (RabbitMQ). - Web servers and load balancers (Nginx, HAProxy).
- Observability in production: VictoriaMetrics/Prometheus, Grafana, and log/trace pipelines
(Fluentd/EFK or ELK, OpenTelemetry). - Working knowledge of at least one programming language (.NET, Node.js, or Go) to partner
effectively with application teams. - Knowledge of best practices for running an always-up, always-available service.
- Excellent troubleshooting, a fast learner, and a strong ownership mindset — comfortable operating
independently and taking a problem from first report to resolution.
فرآیند استخدام در علیبابا
هر فرصت شغلی، یک فرصت ماجراجویی
- ارسال رزومه
- بررسی رزومه
- ارزیابی تخصصی
- ارزیابی منابعانسانی
- گفت و گو با مدیر ارشد
- ارزیابی نهایی
- دعوت به همکاری
- شروع ماجراجویی