Senior Site Reliability Engineer
Shape reliability for customer-facing products at scale as a Senior SRE in Cluj. Own SLOs, observability and resilience while partnering with product teams in a flexible environment.
Core details
Company type: International Retail Group – In-House Product Engineering
Location: Cluj-Napoca, Romania
Work model: Hybrid
Engagement: Employment Contract, Permanent, Full-Time
Role level: Senior level
Experience level: 5–8 Years
Key Skills / Tech Stack: SRE (SLIs, SLOs, error budgets), AWS / GCP / Azure, Kubernetes, Docker, Terraform, GitLab CI/CD, Datadog, Python / Go / Bash
Department: Engineering
Industry: Retail – Home Improvement
About the client
Our client is a large international retail group with an in-house engineering organisation, technology built internally, not outsourced. Teams own what they design and run, end to end, on platforms that serve the group's markets and brands across Europe.
The Cluj engineering hub is part of a wider international ecosystem, and the local capability continues to grow.
About the role
This is a hands-on SRE role inside a product engineering environment. You will work directly with software engineering squads that build and own their services in production, helping them improve how those services are designed, observed and operated.
The role sits at the intersection of software engineering, cloud infrastructure and production reliability. You will have hands-on ownership around SLOs, observability, incident response, automation and resilience, while also influencing engineering teams and contributing to shared reliability practices across the organisation. You will collaborate closely with Product Engineering, Platform Engineering, Incident Management, Security, Network and Observability teams.
This is not a role focused on maintaining infrastructure or CI/CD pipelines. The focus is on engineering reliability into digital products used by customers at scale.
Responsibilities
Reliability engineering
Define and implement SLIs, SLOs and error budgets for production services.
Help product engineering teams understand and manage reliability trade-offs.
Identify reliability risks and drive improvements before they become production issues.
Help design resilient, scalable and highly available cloud-native systems.
Improve production readiness and operational ownership across engineering squads.
Observability and service health
Improve observability across services, with a strong focus on customer impact.
Develop meaningful and actionable monitoring and alerting.
Use Datadog and related capabilities to improve visibility across distributed systems.
Help teams move from reactive monitoring towards proactive reliability engineering.
Incident management
Support the response to major production incidents and participate in the on-call rotation.
Contribute to effective, blameless post-incident reviews.
Identify systemic causes rather than treating individual symptoms.
Ensure post-incident actions translate into long-term engineering improvements.
Automation and operational excellence
Identify repetitive operational work and reduce engineering toil through automation.
Improve practices around deployment, monitoring and service operations.
Contribute to shared SRE standards, tooling and engineering patterns.
Help make reliability part of everyday software engineering rather than a separate operational activity.
Qualifications
You have strong hands-on experience with production systems and understand that Site Reliability Engineering goes beyond infrastructure automation. Specifically, you bring experience with:
SRE principles and practices: SLIs, SLOs, error budgets, observability, incident response and automation
Operating production systems at scale in cloud environments, and the reliability, availability and scalability trade-offs involved in distributed systems
At least one major cloud platform: AWS, GCP or Azure
Kubernetes and Docker
Infrastructure as Code, particularly Terraform
CI/CD environments such as GitLab CI/CD (preferred), GitHub Actions or Jenkins
Observability platforms, ideally Datadog
Scripting or programming in Python, Go, Bash or JavaScript
Troubleshooting complex production issues
Just as importantly, you are comfortable working directly with software engineers, challenging existing practices where needed, and communicating reliability topics clearly across teams.
Experience defining or owning SLOs for production services, designing alerting around customer impact, taking an active role in major incident management, or introducing SRE practices across multiple engineering teams would be especially valuable. You do not need to tick every technology box ,strong SRE fundamentals and evidence that you have improved the reliability of real production systems matter more than any single tool.
About the offer
Our client offers a competitive package and room to grow:
Annual performance bonus and employee referral bonus
Private medical coverage through Regina Maria – Priority Plan, extendable to your spouse or children at no additional cost
Life insurance through Metropolitan Life
Eyeglasses vouchers and a co-funded 7Card fitness membership
Access to LinkedIn Learning, LEO Learning and Bookster
Meal vouchers, plus gift vouchers on several occasions throughout the year
21 days of annual leave, increasing with tenure up to 25 days
Additional days off when public holidays fall on weekends
Paid leave for special life events
Hybrid working model and a flexible schedule
Modern, collaborative workspace with fresh fruit and premium coffee
You would join an international product engineering organisation where teams build, own and operate their own digital products, and where reliability is treated as an engineering discipline, with real customer impact and the scope to shape SRE practices across a large-scale environment.
Why GetFrankly?
At GetFrankly we are guiding talent and creating futures.
We understand that a career move is not just about a new role , it's about finding a place where your skills, ambitions, and values align.
Give us a call to discuss what's important for your career and future, and we'll try our best to get you involved in interesting projects and provide you with fulfilling career pathways.
We will guide you towards the right environment where your abilities will thrive and have a significant impact.
- Department
- Software Engineering (IT&C)
- Locations
- Cluj-Napoca
About GetFrankly
GetFrankly is a catalyst for change! With the goal to close the gap that exists between outstanding talent and successful companies.
We are a team of dedicated professionals bound together by a shared commitment to doing the right thing. Founded by a group of experienced individuals with extensive market knowledge, a vast network, and a genuine friendship, we leverage our expertise to provide great services.