Back to Alphabet jobs
Alphabet

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Tokyo, Japan

Job Description

Google's line of products and services to our clients never stops growing. The Partnerships Development team is responsible for seeking and exploring new opportunities with Google's partners. Equipped with your business acumen and extensive product knowledge, you are right on the front line of interacting with our partners, and helping them find ways to grow using Google's newest product offerings. Your knowledge of relevant verticals and relationships with key industry players will help shape our great applications and content for products such as YouTube, Google TV and Commerce.

Google's payments platform powers billions of transactions across our products (AdWords, Play, Cloud, Youtube, gStore, etc.). Our Payments Partnership team builds strategic relationships to ensure a reliable, cost-effective, and secure commerce experience for all Google users. As a Strategic Partner Manager, you will collaborate with cross-functional teams (Finance, Legal, etc.) to forge and manage partnerships with key players in the payments ecosystem to support our payment transaction processing for all Google products and services. Your will own management of these partnerships, and advocate on behalf of Google to engage with business, industry and regulatory leaders to optimize Google's payment processes, ensuring seamless user experiences while driving sustainability for compliance, innovation in scalability and efficiency across the industry.

The Global Partnerships organization is responsible for exploring new opportunities with Google's partners. Google’s Global Partnerships team works with a wide range of partners to bring the best of Google to power their business. The Global Partnerships team supports Google’s own Product teams with essential partnerships to help Google’s user experiences in advertising, Search, Assistant, Maps, Travel, Shopping, Payments and more. Teams create product-enabling partnerships, go-to-market strategies and incubate business growth for a variety of products. Responsibilities

  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.

Qualifications Minimum qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Preferred qualifications:

  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.

About Alphabet

First seen: October 1, 2026
Last updated: October 3, 2026