Barcelona, Spain Remote (Global) Employment

QAD, Inc. is hiring a Sr. Site Reliability Engineer - SRE

Responsibilities

  • Create and manage systems that are highly available, scalable, and resilient to ensure excellent user experiences.
  • Act as the lead expert in observability practices such as monitoring, alerting, log management, distributed tracing, dashboarding, and synthetic transaction testing.
  • Build reliable and maintainable software solutions and self-service tools to automate operations and increase system dependability.
  • Reduce manual operational work by applying automation, refining processes, and using structured problem-solving techniques.
  • Take ownership of incident management, contribute to on-call schedules, and lead post-incident reviews focused on learning, not blame.
  • Establish, enforce, and monitor Service Level Indicators, Service Level Objectives, and error budget policies.
  • Use infrastructure-as-code methods, GitOps workflows, and CI/CD pipelines with tools like Terraform, Flux, and GitHub Actions.
  • Offer reliability guidance during design phases and help shape system architecture decisions.
  • Produce clear documentation, maintain operational runbooks, and support engineering teams through coaching.
  • Apply artificial intelligence ethically to speed up diagnostics, enhance documentation, minimize repetitive tasks, and create smart operational processes, ensuring proper human oversight, security, and compliance.

Compensation

Competitive salary and benefits package

Work Arrangement

Hybrid work model with flexibility for remote and office-based work

Team

Collaborative engineering environment focused on reliability, automation, and continuous improvement

Responsibilities (10)

  • Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences.
  • Serve as a subject matter expert for observability, including monitoring, alerting, logging, tracing, dashboards, and synthetic testing.
  • Develop robust, maintainable software and self-service tooling to automate operational tasks and improve reliability.
  • Identify and eliminate operational toil through automation, process improvements, and systematic problem solving.
  • Lead incident response, participate in on-call rotations, and drive blameless post-mortems.
  • Define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Leverage infrastructure as code, GitOps practices, and CI/CD automation using Terraform, Flux, and GitHub Actions.
  • Provide reliability expertise during system design reviews and influence architectural decisions.
  • Document processes, build runbooks, and mentor engineers across the organisation.
  • Leverage AI responsibly to accelerate investigations, improve documentation, reduce toil, and build intelligent operational workflows while maintaining appropriate human oversight, security, and governance.

May be available depending on role and candidate eligibility

About company
QAD, Inc.
QAD Redzone helps manufacturers improve employee productivity and engagement through software and coaching, enabling adaptive manufacturing enterprises to increase operational efficiency.
All jobs at QAD, Inc. Visit website
Job Details
Department Engineering
Category infrastructure
Posted 2 months ago