Responsibilities
- Develop foundational platform components including API gateways, authentication, authorization, key management, and multi-tenancy safeguards.
- Create, maintain, and tune APIs and server-side logic using Python frameworks, with a focus on FastAPI, Flask, or Django.
- Design and deploy systems for tracking usage, integrating billing, and enforcing rate limits on AI inference endpoints.
- Develop resilient, scalable microservices to support data workflows and artificial intelligence integrations.
- Construct and manage a high-volume routing and proxy layer tailored for AI model inference traffic.
- Partner with multiple teams to define system architecture and ensure components work together seamlessly.
- Integrate telemetry capabilities from inception, including structured logs, distributed tracing, metrics collection, and alerting mechanisms.
- Build and manage automated pipelines for continuous integration and deployment, along with monitoring solutions for production environments.
- Lead architectural decisions around system design, data structures, and selection of technologies.
- Detect and resolve performance issues to enhance system reliability, scalability, and response speed.
- Define and enforce coding standards across the backend stack, covering testing, peer review, CI/CD, and deployment protocols.
- Uphold security, regulatory compliance, and long-term maintainability throughout the API development lifecycle.
- Work closely with machine learning infrastructure engineers to connect backend systems with model serving platforms such as NVIDIA Dynamo, vLLM, SGlang, and TensorRT-LLM.
Work Arrangement
Remote (India)