
Every agent we ship is only as reliable as the infrastructure underneath it. As our AI Infrastructure Engineer, you'll own the systems that route, queue, and scale inference requests across our platform — the layer nobody sees until it breaks.
You'll work at the intersection of distributed systems and ML serving, making calls about caching, autoscaling, and failover that directly affect latency for every customer running agents on our platform.
Traffic spikes 4x during a customer's product launch, and your autoscaling policies are the reason nothing falls over. You're the one tuning the load balancers, watching GPU utilization in real time, and deciding whether to spin up new inference nodes or reroute to a warm pool.
On a quieter week, you're profiling p99 latency across our model endpoints, shaving milliseconds off the request path, and writing the runbooks that let on-call engineers resolve incidents without waking you up.




