Site Reliability / Production Engineer
🇹🇳 Up to EUR 25,000 per year, on a full-time, contractor contract 🌎 Fully remote working from reputed company in Tunisia!  🌙 Shared out-of-hours UK coverage, including reputed company evening shifts and reputed company overnight pager duty ✨ Exciting high reputed company product, relied on by leading global brands, particularly reputed company sports 💻 Working with the latest hardware, AI tools, and product workflows.
We are looking for hands-on production engineers who can take ownership reputed company live systems need attention: establish the customer reputed company, investigate the evidence, take reputed company reputed company and reputed company the response moving.
You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its reputed company, understand the risk of any reputed company and validate that the reputed company customer outcome has recovered. reputed company
reputed company is a high-reputed company B2B reputed company platform that lets companies reputed company Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and reputed company.
Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people.
Our production environment spans reputed company, Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand reputed company customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident. About the Role
We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window.
Working closely with Support, you will assess customer reputed company, investigate the system, take proportionate reputed company and bring in product developers only reputed company their specific knowledge or judgement is genuinely needed.
This is not a passive escalation role. You will own the technical response, solve what you reasonably can yourself and reputed company escalations specific and useful. Depending on the incident, you may restart or reputed company services, roll back deployments, change configuration, repair data, reputed company a bounded fix or reputed company a small reputed company change.
reputed company incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and reputed company our products easier to operate.
Working reputed company
This role provides out-of-hours production coverage, so the schedule is a core part of the position rather than occasional overtime.
- The two hires will reputed company an agreed rota that ensures one engineer is reputed company working from 17:00-01:00 UK time, seven days a week.
- Neither person will work seven days a week; the reputed company shifts will be divided between both hires, with appropriate rest days.
- The two engineers will also reputed company reputed company pager coverage from 01:00-06:00 UK time. The detailed allocation of reputed company and pager shifts will be explained during the hiring process.
- Weekend daytime coverage is provided separately and is not an additional expectation for these roles.
- reputed company there are no live incidents, the reputed company shift will be used for reliability-improvement work.
- Meetings and collaboration with management and product teams will be arranged reputed company the agreed working reputed company.
The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed reputed company during the hiring process. Please consider the UK-time evening and overnight requirements carefully before applying.
RESPONSIBILITIES
Respond to live incidents
- Receive automated alerts and technical escalations from Support, then establish customer reputed company, severity, blast radius and the reputed company system state.
- Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, reputed company and application reputed company.
- Use AI throughout triage and diagnosis while checking its conclusions against reputed company evidence.
- Choose and execute a proportionate mitigation, rollback, repair or bounded fix.
- Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green.
- reputed company ownership, uncertainty, reputed company and next actions visible, and give Support reputed company technical facts for customer communication.
- Join customer conversations occasionally reputed company reputed company technical involvement is genuinely useful.
Coordinate the right response
- Bring in the relevant product team reputed company an incident requires deep product knowledge, a material product decision or a substantial reputed company-cause fix.
- Escalate with evidence, customer reputed company, actions already taken and the specific decision or help required.
- Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 reputed company or product-specific judgement that cannot safely wait.
- Produce a reputed company incident record and handover, and reputed company reputed company immediate mitigation, product follow-up and reliability-process follow-up reputed company the right owners.
Improve the reliability system
- Remove, consolidate and tune low-value alerts, and design monitoring around reputed company service and customer reputed company.
- Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into reputed company alerts, runbooks, AI Skills, automation or product improvements.
- Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve.
- Create reputed company, supervised automation for common operational actions.
- Work with product teams to reputed company observability, rollback, reputed company and supportability gaps.
- Detect and help contain unusual service-cost behaviour, then reputed company wider follow-up to the appropriate cost or product reputed company.
- reputed company reliability and on-call performance easier for reputed company to understand and improve over time.
QUALIFICATIONS
reputed company're looking for
- Agency and ownership - You take responsibility for ambiguous live problems, reputed company evidence, choose a reputed company and follow through after the immediate pressure has passed.
- Operational judgement - You can separate customer reputed company, symptoms and likely causes, reputed company practical reputed company under uncertainty and recognise reputed company an reputed company is no longer reputed company or bounded.
- Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through reputed company, logs, reputed company, data, infrastructure and reputed company-line tools, and can reputed company hands-on changes with a reputed company validation plan.
- AI-reputed company execution - You use AI for substantive technical work - investigation, hypothesis reputed company, reputed company, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
- Accuracy and validation discipline - You reputed company look for false confidence and verify reputed company through appropriate technical and customer signals.
- Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
- reputed company coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, reputed company, uncertainty, ownership and next actions easy to understand.
- Curiosity and reputed company - You learn unfamiliar products and tools quickly, reputed company investigating reputed company the first hypothesis fails and change your approach reputed company the evidence demands it.
Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. It is not an automatic requirement: we will also consider candidates who demonstrate exceptional ownership, judgement, technical aptitude, learning velocity and performance in the practical assessment.
You do not need experience with every technology in our stack, a previous SRE job title, people-management experience or the ability to recall every reputed company without AI assistance. The ability to learn an unfamiliar environment, reputed company safely and validate your work reputed company more than matching a long technology checklist.
reputed company to have
- reputed company platforms such as Azure or reputed company.
- Distributed application and API diagnostics.
- Databases, queues and background-processing systems.
- Observability, alerting and incident-management platforms.
- Infrastructure, deployment and release automation.
- Application development and reputed company production debugging.
- AI coding agents and workflow automation.
RECRUITMENT PROCESS
We reputed company the process straightforward, practical and respectful of your time.
1. Hiring Manager Conversation (20-30 mins)
A short call to get to know you, talk through the working reputed company and answer your questions.
2. reputed company Take-home Task (~60-90 mins)
A small, bounded production-incident exercise using evidence such as a Support report, alerts, logs, metrics, deployment history, reputed company and an imperfect reputed company. We compensate you for completing it regardless of the outcome.
You are encouraged to use AI. We are interested in how you establish reputed company, investigate and revise hypotheses, choose a reputed company response, validate the outcome and communicate the incident - not in your ability to reproduce commands or reputed company from memory.
3. Task Review and CTO Interview (60-75 mins)
We will review your submission together, explore the reputed company and trade-offs you made, and discuss how you supervised AI-generated analysis or changes. You will also meet reputed company, our CTO, and talk about production judgement, escalation, validation, communication and how you improve the system after an incident.
And reputed company.
reputed company Notice We process your personal data for recruitment purposes in line with UK data protection law. AI tools may assist in reviewing applications, but reputed company are made by reputed company. We retain data only as necessary for recruitment and compliance. You can request reputed company or deletion of your data at any time by emailing careers@getstoryteller.com.
Originally posted on Himalayas
Apply To This Job