Role Overview
This role combines platform engineering and structured first-line operational responsibility.
You will work within Central Engineering, building and improving our internal developer platform - focusing on observability, load testing, resiliency, and system robustness - while also participating in a defined Level 1 incident duty rotation.
Our systems handle millions of requests daily across distributed microservices. Stability, scalability, and performance are critical. This role directly contributes to improving reliability at scale, not just feature delivery.
You will work in a structured regional shift pattern to help provide 24/7 coverage as part of our first-line response model.
Our Stack (experience in some is expected)
- Language: Java 21
- Frameworks: Spring Boot, Spring Data, Spring Cloud
- Architecture: Microservices, REST APIs, Event-Driven Systems
- Databases: MySQL, MyBatis, ShardingSphere, MongoDB
- Caching: Redis (AWS ElastiCache), Elasticsearch
- Messaging: RocketMQ
- Cloud/Infra: Docker, Kubernetes, AWS
- Observability: Grafana, Prometheus, Loki, Tempo, CloudWatch
Responsibilities
What You’ll Be Doing
Platform Engineering (Primary Focus Outside Duty Window)
- Improve observability across services (metrics, tracing, logging).
- Design and enhance load testing frameworks and resiliency tooling.
- Build reusable platform capabilities and internal libraries.
- Identify recurring operational pain points and eliminate them permanently.
- Contribute to standards that improve scalability and maintainability.
- Participate in architecture discussions and technical reviews.
Level 1 Operational Duty (During Assigned Shift Window)
- Participate in a defined regional shift pattern.
- Provide first-line incident response within your duty window.
- Triage alerts via Rootly and structured runbooks.
- Execute pre-approved mitigation steps.
- Escalate to product teams when required.
- Ensure clear documentation and handover between shifts.
This role does not own product incident resolution, but ensures structured and rapid triage.
Requirements
What You’ll Be Doing
Platform Engineering (Primary Focus Outside Duty Window)
- Improve observability across services (metrics, tracing, logging).
- Design and enhance load testing frameworks and resiliency tooling.
- Build reusable platform capabilities and internal libraries.
- Identify recurring operational pain points and eliminate them permanently.
- Contribute to standards that improve scalability and maintainability.
- Participate in architecture discussions and technical reviews.
Level 1 Operational Duty (During Assigned Shift Window)
- Participate in a defined regional shift pattern.
- Provide first-line incident response within your duty window.
- Triage alerts via Rootly and structured runbooks.
- Execute pre-approved mitigation steps.
- Escalate to product teams when required.
- Ensure clear documentation and handover between shifts.
This role does not own product incident resolution, but ensures structured and rapid triage.
Contact Sporty Group
Job Details
Ready to Apply?
Don't miss out on this opportunity. Apply now and take the next step in your career.
Apply Now