• Senior AI Cluster

    NVIDIA (Santa Clara, CA)
    …be doing: + Build internal perf/power profiling and analysis tools and platform for AI workloads at cluster scale + Build debugging tools for common ... frameworks like Pytorch, TensorFlow and etc + Knowledge of AI cluster job scheduling, storage management and...GPU cluster scale continuous profiling & analysis tools /platforms + Solid experience in large AI more
    NVIDIA (10/01/24)
    - Save Job - Related Jobs - Block Source
  • Senior AI -HPC Cluster

    NVIDIA (Santa Clara, CA)
    …performance for a variety of AI /HPC workloads. + Working knowledge of cluster configuration managements tools such as Ansible, Puppet, Salt. + Experience ... parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...join us today! As a member of the GPU AI /HPC Infrastructure team, you will provide leadership in the… more
    NVIDIA (11/06/24)
    - Save Job - Related Jobs - Block Source
  • Senior Site Reliability Engineer…

    NVIDIA (Santa Clara, CA)
    …5K GPUs cluster . + Deep understanding of GPU computing and AI infrastructure. + Passion for solving complex technical challenges and optimizing system ... NVIDIA is the leader in AI , machine learning and datacenter acceleration. NVIDIA is...Solid experience with GPU clusters, and working knowledge of cluster configuration management tools such as BCM… more
    NVIDIA (09/25/24)
    - Save Job - Related Jobs - Block Source
  • Senior AI Infrastructure Engineer

    NVIDIA (Santa Clara, CA)
    We are now seeking a Senior AI Infrastructure Engineer! NVIDIA's Compute Architecture Group is growing our team of AI focused Infrastructure Engineers who ... What you'll be doing: + Administer an NVIDIA Internal AI cluster composed of Linux systems ranging...updates, and maintenance of system availability using modern DevOps tools (Ansible, Gitlab, etc.) + Plan and maintain new… more
    NVIDIA (11/06/24)
    - Save Job - Related Jobs - Block Source
  • Senior SRE Engineering Leader - AI

    NVIDIA (Santa Clara, CA)
    NVIDIA is leading the way in the AI revolution, revolutionizing industries with our brand-new GPU technology. Our GPUs drive groundbreaking innovations, from ... in computer vision, speech recognition, and more. As "the AI computing company," we constantly push the limits of...leaders to join us on an exciting journey as Senior SRE Engineering Leader. Lead our globally distributed clusters,… more
    NVIDIA (10/08/24)
    - Save Job - Related Jobs - Block Source
  • Principal Observability Architect, AI

    NVIDIA (Santa Clara, CA)
    …, HW, and SW engineering and research teams to define a vision and roadmap for AI /HPC cluster observability. + Architect and lead teams to d evelop, test, and ... NVIDIA's Hardware Infrastructure organization is seeking a Senior or Princip al Data and Observability Architect....We serve and collaborate directly with NVIDIA's rapidly growing AI , HW, and SW engineering and research teams across… more
    NVIDIA (11/02/24)
    - Save Job - Related Jobs - Block Source
  • Senior Software Engineer, Kubernetes - DGX…

    NVIDIA (Santa Clara, CA)
    …experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have significant software engineering experience ... with kubernetes including cluster operations, operator development, node health monitoring and working...deploy leading infrastructure solutions for a broad range of AI -based applications. If you're creative, passionate about kubernetes and… more
    NVIDIA (10/11/24)
    - Save Job - Related Jobs - Block Source
  • Senior MLOps Engineer, GenAI Framework

    NVIDIA (Santa Clara, CA)
    NVIDIA is looking for a dedicated and motivated senior build and continuous integration (CI/CD) engineer for its GenAI Frameworks (NeMo, ... working on Large Language Models (LLM), Multimodal (MM), and Speech AI . NeMo provides end-to-end model training, including data curation, alignment, customization,… more
    NVIDIA (10/08/24)
    - Save Job - Related Jobs - Block Source
  • Senior Solutions Architect, NPN

    NVIDIA (Santa Clara, CA)
    …both on-premises and cloud based. + 12+ years of proven experience with cluster management and related tools , including Docker Containers, Slurm, Kubernetes, and ... part of a team that's revolutionizing the field of AI with data center scale solutions? We are looking...are the voice of experience, using Kubernetes, SaaS, infrastructure-as-code tools , network debugging, and problem solving skills to help… more
    NVIDIA (09/18/24)
    - Save Job - Related Jobs - Block Source
  • CephFS Senior Software Engineer

    IBM (San Jose, CA)
    …talk. Your Role and Responsibilities IBM's Ceph[1] engineering organization is looking for a senior software engineer to join the CephFS team. In this role you will ... to higher-level APIs for integrating with other systems (OpenStack, OpenShift, an NFS-Ganesha cluster , Samba, etc). As a member of the CephFS engineering team, you… more
    IBM (10/25/24)
    - Save Job - Related Jobs - Block Source
  • Senior Software Test Development Engineer…

    NVIDIA (Santa Clara, CA)
    We are looking for a highly experienced AI Senior Software Test development engineer in NVIDIA's Deep Learning SWQA team. The position is in NVIDIA Deep Learning ... and AI Software Quality Assurance team that defines, develops and...in validating Data Center GPU based infrastructure (multi-GPUs, multi-nodes, cluster ) + Background in validating fault tolerance infrastructure +… more
    NVIDIA (09/06/24)
    - Save Job - Related Jobs - Block Source
  • Senior Technical Program Manager - GPU…

    NVIDIA (Santa Clara, CA)
    Hardware Infrastructure is seeking a Senior Technical Program Manager to lead the strategy and execution of programs to support the bringup, operations and ... infrastructure we build and operate enables NVIDIAs most advanced AI and hardware researchers and engineers to create the...a fast paced and evolving landscape that requires a senior TPM leader to guide engineering roadmaps to be… more
    NVIDIA (11/01/24)
    - Save Job - Related Jobs - Block Source
  • Senior Product Manager, Machine Learning…

    Google (Sunnyvale, CA)
    …storage, compute hardware, databases, file systems, data analytics, cluster management, and other software infrastructure areas. Preferred qualifications: ... management and delivery. + 3 years of experience with AI /ML model or tooling development (eg, Product Management or...down technical barriers, and strengthen existing systems. As a Senior Product Manager for ML Frameworks, you will lead… more
    Google (10/23/24)
    - Save Job - Related Jobs - Block Source
  • Senior ASIC Physical Design Engineer,…

    NVIDIA (Santa Clara, CA)
    …graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is a "learning machine" that ... parallel computing! More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...high-frequency and low-power CPUs, GPUs, SoCs at block level, cluster level, and/or full chip level, with a focus… more
    NVIDIA (09/25/24)
    - Save Job - Related Jobs - Block Source
  • Senior Software Test Development Engineer…

    NVIDIA (Santa Clara, CA)
    …Deep Learning SWQA team. The position is in NVIDIA Deep Learning and AI Software Quality Assurance team that defines, develops and performs tests to validate ... healthcare, speech recognition, natural language processing, and a wide variety of other AI scenarios. We collaborate with multiple AI product teams to develop… more
    NVIDIA (09/05/24)
    - Save Job - Related Jobs - Block Source
  • Senior ASIC Timing Engineer

    NVIDIA (Santa Clara, CA)
    …graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is a "learning machine" that ... parallel computing! More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...Nvidia's GPUs, CPUs, DPUs and SoCs at block level, cluster level, and/or full chip level. + Work with… more
    NVIDIA (09/20/24)
    - Save Job - Related Jobs - Block Source
  • Senior System Reliability Engineer

    NVIDIA (Santa Clara, CA)
    …can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and build our teams with ... Hardware Reliability Engineering for Electronics/Server Systems (graphics cards, server, rack, cluster ) from Concept to End-of-Life phase. + Establish, deliver and… more
    NVIDIA (11/03/24)
    - Save Job - Related Jobs - Block Source
  • Principal Infrastructure SRE - Storage

    NVIDIA (Santa Clara, CA)
    …crowd: + Deep understanding of other infrastructure components like DNS, LDAP, NIS, Security Tools etc. + Experience with HPC cluster management tools such ... people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An...a focus on infrastructure automation. + Develop and maintain tools for collecting, analyzing, and visualizing data for reporting,… more
    NVIDIA (10/27/24)
    - Save Job - Related Jobs - Block Source