J
Staff Site Reliability Engineer
Jobgether
Netherlands
Remote
Expert
posted 3 days ago
Vacancy Summary
Join a fully remote engineering team as the first dedicated Site Reliability Engineer in the Netherlands. You will establish reliability practices across engineering teams, combining hands-on engineering with organization-wide influence.
Required skills
Kubernetes
Go
Kafka
Automation
AI Agents
AI tools
Distributed Systems
Database
Incident Response
Fraud detection
Production engineering
Testing
SRE
Planning
cloud infrastructure
Technical Leadership
System design
Security
Mentoring
Tooling
Payments
Visa
production code
Technical Communication
error budgets
Incident Investigation
Observability
Incident Management
Make
Organization
Infrastructure as code (IaC)
Growth
Teams
Amazon Web Services (AWS)
Artificial Intelligence (AI)
Driving
AWS (SES)
Communication Skills
Bootstrap (Framework)
Dashboards
Compensation
Analysis
Education
TypeScript
Support
Production Systems
influence
Investment
CAN
Clear
Culture
WELL
Safety
Development
Diagnosis
DataDog
Investigations
Telemetry
ElasticSearch
Redis
Engineering Leadership
FinOps
trade
Production Operations
Operational Risk
ACT
Engineering Standards
Site Reliability Engineering
Delivery
SLOs
Metrics
Resilience
Signals
Operational Readiness
Deployment
Progress
English level
Fluent
This position is listed on behalf of a partner company that manages all applications and next steps. Our partner is seeking a Staff Site Reliability Engineer based in the Netherlands. This is a high-impact reliability leadership role within a fully remote engineering organization operating globally. You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems. The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience. You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable. A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services. You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations. The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.
Verantwoordelijkheden
•
Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.
•
Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.
•
Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.
•
Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.
•
Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.
•
Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.
•
Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.
•
Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.
•
Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.
•
Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.
•
Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.
•
Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.
•
Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
•
Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.
•
Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.
Still searching manually?
Let us do the work for you.
TotaMatch works for you
We scan thousands of jobs daily and notify you when there is a match. No searching needed.
Anonymous, safe and free
Your profile stays anonymous. Your employer will not see it. You choose when to become visible.
Ready in 3 minutes
Answer a few questions and create your profile in minutes. No commitment.
About TotaMatch
TotaMatch helps professionals find work that truly fits their work happiness. We believe work is more than just an income. It is a source of fulfillment, growth, and pride. Instead of endlessly scrolling through job boards, TotaMatch works for you. Our platform continuously analyzes thousands of opportunities and identifies roles that align with what truly matters to you. You focus on your work and the people around you. We make sure you never miss a better opportunity.
Apply for Staff Site Reliability Engineer
In under 2 minutes.
Safe, free and anonymous via TotaMatch.
Create a free TotaMatch profile
Apply faster and enjoy extra perks.
or
Why via TotaMatch?
Chat directly with the employer when there's a match.
Automatically receive new, similar jobs.
Hired through TotaMatch? You'll receive a welcome bonus worth € 150.