Service Manager & Site Reliability Consultant
Service Manager & Site Reliability Consultant
Łódź, PL, 90-118 Kraków, PL, 30-302 Poznań, WP, PL, 61-569 Warszawa, PL, 00-839
Type of contract: B2B contract
Salary range: 205 - 275 PLN/H
What will you do?
You will manage real-time incident response, impact assessment, external communications, and coordination across production systems. Working at the intersection of Site Reliability Engineering, Incident Response, and Partner Operations, you will ensure timely, accurate, and SLA-compliant communication while supporting the scalability and reliability of global operations.
Your tasks
- Monitor and respond to production incidents
- Coordinate incident response activities across teams
- Assess impact and determine incident severity
- Manage external communications and status page updates
- Support incident reporting, RCA activities, and SLA tracking
- Collaborate with Engineering teams to improve reliability and observability
- Drive process improvements and automation initiatives
- Contribute to internal reliability tooling using Python or Kotlin
Your skills
- 5+ years of experience in Incident Operations, Site Reliability Engineering, Technical Operations, or a similar role
- Experience working in on-call environments with SLA-driven responsibilities
- Strong understanding of distributed systems and production environments
- Experience with monitoring, alerting, and incident management tools
- Familiarity with APIs, system integrations, and observability platforms
- Hands-on experience with Python or Kotlin
- Understanding of SDLC and production reliability principles
- Strong communication, stakeholder management, and decision-making skills
- Ability to work effectively in high-pressure environments and manage multiple priorities
- Strong ownership mindset and cross-functional collaboration skills
Nice to have
- Experience with Datadog or Chronosphere
- Experience with PagerDuty, Rootly, or Slack workflows
- Experience managing external status pages
- Experience with incident management automation and process improvements
- Experience contributing to reliability tooling and platform engineering