San Francisco, United States of America Hybrid Employment

Inference is hiring a Fullstack Engineer - Frontend Focus

Responsibilities

  • Own the user experience – design, build, and polish the dashboards, consoles, and customer-facing apps that let users observe, configure, and pay for inference at scale.
  • Ship end-to-end features – from Figma wireframe to React component to backend API and database migration.
  • Design a component system in React + Tailwind that supports rapid iteration and a cohesive design language.
  • Optimize performance – SSR, code-splitting, hydration, and WebSocket-driven real-time updates that hold up under millions of requests per day.
  • Collaborate across disciplines – work shoulder-to-shoulder with distributed-systems engineers, product designers, and founders to turn complex infrastructure into delightful product.
  • Level-up the team – lead design reviews, mentor junior engineers, and introduce best practices for testing, accessibility, and observability.

Requirements

  • 5+ years building production React applications
  • Deep knowledge of Tailwind CSS & modern CSS architecture
  • Typescript mastery and strong fundamentals in JS/DOM/APIs
  • Experience designing REST/JSON or gRPC backends (Node, Go, or similar)
  • AuthN/AuthZ design (OIDC, JWT)
  • Product sense: you care about UX details & accessibility

Nice to Have

  • Experience with Tanstack / Next.js
  • Data-viz libraries (Recharts, Visx, D3)
  • tRPC experience
  • Familiarity with GPU or ML tooling dashboards
  • Comfort debugging perf issues (Lighthouse, Chrome DevTools)
  • Dev-ops chops: CI/CD, Docker, Terraform

Work Arrangement

Hybrid — San Francisco

Additional Information

  • Most of us are in the office 3–4 days a week
  • remote candidates considered if time-zone compatible with Pacific hours
Required Skills
TailwindCSS
About company
Inference
We are building a real-time marketplace for AI inference that matches spare GPU capacity inside data centers with demand from developers building AI-powered applications. We currently operate the world's largest distributed GPU cluster, with over 5,000 GPUs, hundreds of individual operators, and millions of gigabytes of VRAM connected to the network at any given moment. We are a small, well-funded team working on difficult, high-impact problems at the intersection of AI and distributed systems. We primarily work in-person from our office in downtown San Francisco. Our investors include A16z CSX and Multicoin. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do.
All jobs at Inference Visit website
Job Details
Department Engineering
Category fullstack
Posted 3 days ago