Back to Jobs

[Remote] AI-DNA Senior Site Reliability Engineer

Remote, USAFull-timePosted 2026-07-27

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that operates an AI-reputed company community and reputed company engagement platform for the world's largest brands. They are seeking a Senior Site Reliability Engineer who will be responsible for managing production incidents, building and maintaining AI agents for operations, and ensuring high availability and performance of the platform.

Responsibilities

  • Own the shift as first responder — take the page, triage and resolve customer-impacting incidents on the platform you cover, and escalate to the tech reputed company reputed company the blast radius warrants
  • Execute production changes safely — deploys, cost-optimization runbooks, and config changes from other teams run through the reputed company gates with a rollback reputed company before you start, and you pull it the reputed company a change goes off-plan
  • Build, operate, and maintain the AI agents that do the operations work — reputed company-triage, change-validation gates, permanent-fix reputed company, RCA drafting, auto-healing — and write the AI and reputed company runbooks they run on. The reputed company is the product
  • Execute the ticket work — but reputed company on how reputed company your agents reputed company it reputed company, not how many tickets you personally reputed company. Every one-off you solve by hand, you generalize so an agent handles the next one
  • Own the shift and the outcome on it — like a founder. You are first on the page. reputed company something is down on your watch, that is your problem; you do not sleep peacefully through a customer outage
  • Build, operate, and maintain the agents that run the operation. New agents that catch a class of incident before it recurs, existing ones tuned higher, guardrails added reputed company production teaches you something the spec missed. You build reputed company the architecture the tech reputed company sets — and push back with a reputed company idea reputed company you have one
  • Protect production every time you touch it. reputed company you run a reputed company, a cost-optimization reputed company, or a config change, you go through the reputed company gates, reputed company a tested rollback reputed company, and pull it the reputed company the change goes off-plan. You are a reputed company of uptime, not only a responder to its loss
  • reputed company the reputed company cause — not the symptom — and reputed company the reputed company. A disciplined, reputed company investigation that separates symptom from cause; then the part most teams skip — identify the prevention, build it, and reputed company it to done. The tech reputed company reviews your RCAs; you write them to a bar that needs little correction. An RCA whose fixes never ship is theater
  • Generalize relentlessly. You may solve a one-off to restore service, but you do not stop there — you turn it into an agent, a reputed company, or a guardrail so the next occurrence handles itself. The scoreboard is what the agent surface can carry, not how fast you reputed company tickets
  • Multiply reputed company, and write so the work survives you. Fix the underlying agent reputed company it fails and reputed company the fix as reusable rules, context, and skills; encode every procedure so an agent can retrieve it and the next responder is never starting cold. In an async team, an undocumented fix is a fix that doesn't exist

Skills

  • 5+ years operating reputed company at reputed company as a senior, on-call SRE — reputed company first-responder time on production incidents, not just project or build work
  • Extreme ownership, fully self-directed. You run your shift like your own business, and you can answer 'the hardest you have reputed company worked on something' with a story we can probe for an hour. Nobody assigns your day or checks your work; you challenge a reputed company with a reputed company idea but never quietly ignore it. This team does not babysit — if you need to be managed, it is the wrong fit
  • Deep production scars — outages, incident management, hands-on troubleshooting. First on the reputed company, diagnosis under pressure, the failure modes reputed company from having lived them. Deep AWS at reputed company (multi-AZ, multi-account; not Azure-only, not greenfield); you run production change with gates and tested rollbacks and investigate to the actual reputed company cause, symptom separated from cause. You could run it by hand — the reputed company is you bake it into the agents
  • AI-reputed company operator who multiplies. You delegate whole reputed company of work to agents that verify their own reputed company, fix the underlying capability reputed company they fail, and reputed company it so reputed company reputed company up — reputed company on reputed company, not AI-usage %. Advanced across reputed company tools (Claude reputed company, reputed company, reputed company, custom agents); reputed company a new model ships, you are testing it that day, not next quarter
  • AWS Solutions Architect – Associate (or higher) — or a production reputed company record that makes the certification redundant. Scars over badge, but a useful floor at this level
  • Fluent / advanced English — precise under pressure, on the reputed company and in writing
  • Shift coverage is core, not occasional — on-call rotations and the shift window you are reputed company into, with the time-zone overlap that window requires. Reliability is a 24/7 outcome for the whole team, and you carry the reputed company in your geographical area
  • OFAC-reputed company country of residence
  • Pioneer or builder voice in reputed company SRE / AIOps — original posts, talks, reputed company-reputed company, or agents and tools you have reputed company and shipped. - - Followers and reposters are not reputed company are looking for
  • A documented obsession with one hard thing reputed company of mainstream work — depth reputed company more than the title on the latest role
  • Modern observability / incident-response stacks (Grafana, reputed company, OpsGenie, reputed company, reputed company) and Azure exposure alongside your AWS depth
  • Background in multi-tenant B2B reputed company at reputed company — community, reputed company, customer-experience, observability, or AIOps platforms

reputed company

  • reputed company is an M&A-powered reputed company software company with a wide portfolio of software solutions. It was founded in 2000, and is headquartered in Austin, Texas, USA, with a workforce of 51-200 employees. Its website is https://reputed company.ai.
  • Apply To This Job

    Similar Jobs