Back to results

SRE

Listed by Spinomenal

  • Verified live ·

Mentioned in this posting

  • Python
  • Bash
  • SQL
  • Microservices
  • Redis
  • AWS
  • Terraform
  • Ansible
  • Jenkins
  • CI/CD
  • Prometheus
  • Grafana

Description

Spinomenal is a dynamic force in the online casino sector, bursting onto the scene in 2014. Renowned as one of the fastest-growing content creators in the iGaming industry, Spinomenal thrives on fostering creativity and collaboration. By nurturing an environment where employees excel through teamwork and communication, Spinomenal maintains a rapid pace of innovation and development. Our dedication to collective effort and shared vision has propelled us to deliver captivating, cutting-edge gaming experiences to players worldwide. About the position: Spinomenal is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, and incident response in our high-throughput, low-latency iGaming environment. Serving as the gatekeeper to production across R&D, DevOps, QA, and Support, you will ensure high availability, secure CI/CD pipelines, and rapid anomaly diagnosis through automation. Responsibilities: Validate complex configurations, execute automated health checks, and own recovery and rollback runbooks for production deployments. Identify manual operational tasks (toil) and automate them using scripting and Infrastructure-as-Code principles to increase reliability. Maintain and optimize distributed tracing, dashboards, alerting thresholds, and log aggregation using New Relic, Grafana, and advanced SQL. Serve as a critical escalation point for complex, cross-layer production failures (Application, Network, Database, and Infrastructure). Drive deep-dive post-mortems and implement structural fixes to prevent incident recurrence. Govern infrastructure and configuration changes across environments in tight collaboration with DevOps to maintain environment parity and prevent drift. Partner with QA, Architects, DevOps, and Game Producers to champion SRE best practices, operational readiness standards, and resilient architecture. Requirements: At least 2 years of experience in a dedicated Site Reliability Engineering (SRE), Production Operations, or Release Engineering role handling high-transaction web applications. Deep experience engineering and maintaining Jenkins pipelines (Pipeline-as-Code / Jenkinsfiles) and familiarity with configuration tools (e.g., Ansible, Helm). Strong hands-on experience managing and troubleshooting cloud infrastructure (AWS: EC2, ECS, S3, Lambda, IAM policies, VPC routing) and Infrastructure as Code (IaC) using Terraform. Proficiency in Python or Bash for writing automation scripts, system utilities, and internal tooling. Advanced capability with observability stacks (New Relic, Prometheus/Grafana) and strong SQL skills for log parsing and database debugging. Mastery of web debugging (interpreting JSON payloads, analyzing API contracts, diagnosing HTTP status anomalies) and understanding of DNS, CDN/Redis caching, and proxies. Demonstrated capability in debugging distributed microservices architectures and defining Service Level Indicators/Objectives (SLIs/SLOs) and error budgets. Exceptional technical communication skills with a proactive mindset and the ability to stay calm under pressure during critical outages and on-call rotations. Show more Show less

Get these jobs by email

New jobs for “SRE” in Gush Dan — Get the new ones each morning.

No account, no password. One email a morning, only when there's something new. Unsubscribe in one click.

Back to results