AI Infrastructure Engineer& ML Systems Jobs

Browse active AI infrastructure jobs in GPU clusters, inference, model serving, and distributed training. Direct links to roles at AI companies, updated daily.

1,042
AI Infrastructure Jobs
Showing 50 of 1,042 jobs
Page 1 of 21
United States, Washington, Redmond +111 hours ago
Brazil, São Paulo, São Paulo17 hours ago
Hawthorne, CA +318 hours ago
M

ML Infrastructure Engineer

Mach9
|AI Application
San Franciscofull-time21 hours ago
Bellevue, Washington, USAfull-time22 hours ago
Hawthorne, CA +3Yesterday
Hawthorne, CA +3Yesterday
United States, California, Mountain View +1Yesterday
Hawthorne, CA +3Yesterday
Hawthorne, CA +3Yesterday
Pittsburgh/RemoteRemote2 days ago
M

ML Systems Engineer, ML Acceleration

Motional Singapore Pte. Limited
D05 Pasir Panjang, Hong Leong Garden, Clementi New Town, SingaporePermanent2 days ago
Mountain View, CA, USA +12 days ago
D05 Pasir Panjang, Hong Leong Garden, Clementi New Town, SingaporePermanent2 days ago
Hamina, Finland3 days ago
San Franciscofull-time3 days ago
U

Global Head of AI Compute Infrastructure

Umelife (Singapore) Pte. Ltd.
D13 Macpherson, Braddell, SingaporeFull Time3 days ago
Santa Clarafull-time3 days ago
San Franciscofull-time3 days ago
Mexico City, México4 days ago
KR - Seoulfull-time4 days ago
Germany, Multiple Locations, Multiple Locations4 days ago
Pittsburgh, Pennsylvania, US +1FULL_TIME5 days ago
Sunnyvale, California, US +3FULL_TIME5 days ago
Bellevue, Washington, US +2FULL_TIME5 days ago
Islandwide, SingaporeFull Time5 days ago
K

Global Head of AI Compute Infrastructure

Kuailu Software (Singapore) Pte. Ltd.
D13 Macpherson, Braddell, SingaporeFull Time5 days ago
Kirkland, WA, USA +15 days ago
Santa Clarafull-time1 weeks ago
United States, Virginia, Reston1 weeks ago
K

AI Infrastructure Engineer

Kaishi Partners Pte. Ltd.
Islandwide, SingaporePermanent, Full Time1 weeks ago
Remote, United StatesRemote1 weeks ago
Islandwide, SingaporePermanent1 weeks ago
D01 Marina, Raffles Place, People's Park, Cecil, SingaporeFull Time1 weeks ago
Kirkland, WA, USA1 weeks ago
Chicago, IL, USA +31 weeks ago

AI infrastructure engineers build and operate the compute systems that make large-scale model training and inference possible. As AI models grow larger and production traffic scales, the systems layer has become one of the most critical and highest-paying specializations in the industry. These roles live at the intersection of distributed systems, GPU computing, and machine learning - requiring deep expertise in all three.

Core responsibilities include GPU cluster management and scheduling, low-latency inference serving, distributed training at scale, and storage systems optimized for large datasets and checkpoints. Engineers in this space work closely with frameworks like CUDA, Triton, and Ray, and build on cloud platforms with H100 and A100 GPU capacity. Roles span both foundational model labs (Anthropic, OpenAI, xAI) and AI product companies scaling inference for millions of users.

Related searches include ML infrastructure jobs, GPU infrastructure jobs, inference infrastructure jobs, model serving jobs, distributed training jobs, LLM infrastructure jobs, and AI platform engineer jobs. Candidates with Kubernetes, networking, storage, observability, PyTorch, JAX, vLLM, TensorRT-LLM, and cluster scheduling experience are especially competitive.

Frequently Asked Questions

What does an AI infrastructure engineer do?

AI infrastructure engineers design and operate the systems that run model training and serving workloads. Day-to-day work includes managing GPU clusters and scheduling (SLURM, Kubernetes), optimizing inference latency and throughput (TensorRT, vLLM, Triton), building distributed training pipelines, and operating the storage and networking infrastructure that feeds large models. At production scale, even small efficiency gains translate to significant cost and latency impact.

What skills are required for AI infrastructure roles?

Strong distributed systems fundamentals are essential - networking, storage I/O, and fault tolerance at scale. GPU programming experience (CUDA, kernel optimization) is increasingly valued as companies move beyond off-the-shelf frameworks. Kubernetes and cloud platform fluency (AWS, GCP, Azure) is expected. Most roles also require familiarity with ML frameworks (PyTorch, JAX) and inference engines (vLLM, TensorRT-LLM). Python and Go or Rust for systems tooling are common language requirements.

How is AI infrastructure different from MLOps?

MLOps often focuses on training pipelines, experiment tracking, model registries, deployment workflows, and monitoring. AI infrastructure is usually closer to the compute layer: GPU clusters, inference serving, networking, storage, scheduling, and cost/performance optimization for large training and serving workloads.