Mexico Remote (Global) Contract

IPSY is hiring a SRE Engineer Contractor

Responsibilities

  • Build and maintain observability across our platform in Datadog: dashboards, monitors, APM, log pipelines, and meaningful, low-noise alerting tied to user impact.
  • Define and track SLIs, SLOs, and error budgets for specific services, and use them to drive reliability conversations with service owners.
  • Participate in the on-call rotation and serve as an SRE Partner during incidents: triage, classify priority (P1-P6), validate hypotheses, and support resolution with guidance on complex situations.
  • Drive incident response per our framework, meeting target response times (P1 within ~15 minutes), and keeping clear, real-time documentation of status, findings, decisions, and next steps.
  • Manage escalation and alerting through Opsgenie, and post structured updates to incident Slack channels at the expected cadence.
  • Own clean shift handoffs and RCA continuity, ensuring notes are complete and incident ownership transfers cleanly across shifts and time zones.
  • Contribute actively to blameless post-incident reviews (RCAs), identifying root causes and action items to prevent recurrence.
  • Automate toil: build scripts, tooling, and self-healing / automated remediation to reduce manual operational work and speed up recovery.
  • Leverage AI tools (e.g., Claude, Cursor) to accelerate debugging, generate and maintain runbooks, draft RCAs, and build automation.
  • Help improve the reliability of our cloud and third-party stack (e.g., AWS, Netlify, CommerceTools, Auth0, Contentful).
  • Contribute to CI/CD reliability, deployment safety, and infrastructure-as-code practices.
  • Maintain and improve SRE runbooks and triage workflows so the whole org responds consistently.

Requirements

  • A great attitude, strong ownership, and a genuine eagerness to learn, test, and try new things.
  • A natural hunger for automation: you instinctively look for ways to remove toil, script repetitive work, and scale yourself through tooling.
  • Comfort and curiosity with AI tools, and a desire to apply them to debugging, documentation, RCAs, and day-to-day reliability workflows.
  • Hands-on experience with observability and monitoring tooling, ideally Datadog (dashboards, monitors, APM, logs); experience with other platforms (e.g., Grafana, Prometheus, New Relic, CloudWatch) translates well.
  • Experience participating in on-call and incident response, including triage, prioritization, escalation, and post-incident reviews.
  • Working understanding of SRE fundamentals: SLIs / SLOs / error budgets, reducing toil, and balancing reliability with delivery speed.
  • Working knowledge of cloud infrastructure (AWS preferred) and modern distributed / microservice and API-gateway architectures.
  • Scripting and automation skills (e.g., Python, Bash, or similar); comfort reading and reasoning about code across services.
  • Familiarity with CI/CD pipelines and infrastructure-as-code (e.g., Terraform).
  • Strong, calm communication: able to participate in an incident channel, give clear status updates, and contribute to a readable RCA.
  • A collaborative mindset and the ability to work effectively across a distributed, multi-time-zone team.

Nice to Have

  • Experience operating high-traffic consumer or e-commerce platforms, especially through peak events (flash sales, product drops).
  • Experience with Opsgenie (or PagerDuty / similar) for alerting and escalation.
  • Experience with e-commerce / subscription platforms such as CommerceTools, and identity providers such as Auth0.
  • Experience with product analytics and monitoring (e.g., Amplitude) to connect reliability to user behavior.

Work Arrangement

Remote (Worldwide) — México, Colombia

Team

Reports to: Infrastructure Engineering Manager

About the Role

We are looking for a Site Reliability Engineer Contractor to join our Foundation / SREIQ team and help keep IPSY fast, available, and resilient for millions of beauty community members. You will be at the center of how we detect, respond to, and learn from incidents, partnering with Engineering, Product, and the broader Tech org to protect the critical flows our members rely on every day. This is a hands-on reliability role. You will own observability and alerting in Datadog, participate in our on-call rotation, act as an SRE Partner during incidents, contribute to root cause analyses, and relentlessly automate the toil out of our operations. We are an automation-forward, AI-embracing team: we expect you to lean on modern AI tools to debug faster, draft runbooks and RCAs, and build smarter automation. The SRE Engineer will report to the Infrastructure Engineering Manager and will be fully remote from México or Colombia covering the PST time zone.

What We Offer

Competitive salary (USD), Paid time off and work from home flexibility

EEO Statement

We celebrate diversity and are an equal-opportunity employer. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, disability status, or any other protected characteristic. If you need reasonable accommodation in the application or employment process, please contact us.

Beware job scams!

IPSY recruiters only use @ipsy.com email addresses. We do not interview via text/message/Teams. We don't ask for software downloads (except Zoom) and we will never ask for sensitive information (like SSN/bank info). Suspect fraud? Report it to law enforcement and recruiting@ipsy.com.

Additional Information

  • Contract role
  • Fully remote position
  • Must be based in México or Colombia
  • Covering PST time zone
  • Resume/CV must be submitted in English
  • Position is remote-first
  • Monthly virtual activities, company-wide offsites, professional development, and learning sessions available
  • On-call rotation participation required
  • AI tools (e.g., Claude, Cursor) are used in workflows
Required Skills
Cloud InfrastructureMonitoring
About company
IPSY

We inspire everyone to express their unique beauty.

Through our unique synergy of expert curation, product innovation, machine learning technology, and a community-first mindset, we make the best in beauty accessible and deliver highly personalized experiences for our members. Our proprietary technology platform is the ultimate in beauty innovation, leveraging millions of member data points. To date, we’ve amassed over 200 million product reviews and built an avid beauty community that remains at the heart of all we do.

Our core competencies include merchandise innovation, personalization through machine learning, and deep community engagement with over 20 million passionate members.

All jobs at IPSY Visit website
Job Details
Department Foundation / SREIQ
Category infrastructure
Posted 6 days ago