Responsibilities
- Design, develop, and maintain scalable backend services for high-throughput transaction systems, ensuring reliability, fault resilience, and response times under one second at massive scale.
- Lead the full lifecycle design of AI-driven agent workflows using frameworks like LangGraph, AutoGen, and CrewAI.
- Deploy and manage AI orchestration systems in AWS environments using tool-calling methods and multi-agent collaboration models.
- Develop robust backend services in Python and Go, defining best practices, idiomatic patterns, and performance benchmarks used across engineering groups.
- Create and manage cloud infrastructure on AWS using services such as ECS, EKS, Lambda, MSK, RDS Aurora, DynamoDB, SageMaker, and EventBridge, following AWS Well-Architected principles.
- Define and enforce software engineering standards including code review processes, testing requirements, CI/CD implementation, and observability through distributed tracing, structured logs, and alerting systems.
- Lead technical architecture reviews for new services, integrations, and system modifications, producing detailed architecture decision records and technical documentation.
- Diagnose and eliminate performance issues in distributed systems by optimizing database queries, implementing caching strategies, and refining asynchronous processing.
- Drive proof-of-concept initiatives for emerging technologies, assessing scalability and operational readiness before production deployment.
- Support senior and mid-tier engineers through code reviews, collaborative coding sessions, and technical mentorship to strengthen team capabilities.
- Participate in on-call duties, incident management, and post-incident analysis to improve long-term system reliability.
Work Arrangement
Remote (Worldwide)
Team
engineering squads