- NVIDIA (Santa Clara, CA)
- …be doing: + Build internal perf/power profiling and analysis tools and platform for AI workloads at cluster scale + Build debugging tools for common ... frameworks like Pytorch, TensorFlow and etc + Knowledge of AI cluster job scheduling, storage management and...GPU cluster scale continuous profiling & analysis tools /platforms + Solid experience in large AI … more
- NVIDIA (Santa Clara, CA)
- …performance for a variety of AI /HPC workloads. + Working knowledge of cluster configuration managements tools such as Ansible, Puppet, Salt. + Experience ... parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...join us today! As a member of the GPU AI /HPC Infrastructure team, you will provide leadership in the… more
- NVIDIA (Santa Clara, CA)
- …5K GPUs cluster . + Deep understanding of GPU computing and AI infrastructure. + Passion for solving complex technical challenges and optimizing system ... NVIDIA is the leader in AI , machine learning and datacenter acceleration. NVIDIA is...Solid experience with GPU clusters, and working knowledge of cluster configuration management tools such as BCM… more
- NVIDIA (Santa Clara, CA)
- We are now seeking a Senior AI Infrastructure Engineer! NVIDIA's Compute Architecture Group is growing our team of AI focused Infrastructure Engineers who ... What you'll be doing: + Administer an NVIDIA Internal AI cluster composed of Linux systems ranging...updates, and maintenance of system availability using modern DevOps tools (Ansible, Gitlab, etc.) + Plan and maintain new… more
- NVIDIA (Santa Clara, CA)
- NVIDIA is leading the way in the AI revolution, revolutionizing industries with our brand-new GPU technology. Our GPUs drive groundbreaking innovations, from ... in computer vision, speech recognition, and more. As "the AI computing company," we constantly push the limits of...leaders to join us on an exciting journey as Senior SRE Engineering Leader. Lead our globally distributed clusters,… more
- NVIDIA (Santa Clara, CA)
- …, HW, and SW engineering and research teams to define a vision and roadmap for AI /HPC cluster observability. + Architect and lead teams to d evelop, test, and ... NVIDIA's Hardware Infrastructure organization is seeking a Senior or Princip al Data and Observability Architect....We serve and collaborate directly with NVIDIA's rapidly growing AI , HW, and SW engineering and research teams across… more
- NVIDIA (Santa Clara, CA)
- …experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have significant software engineering experience ... with kubernetes including cluster operations, operator development, node health monitoring and working...deploy leading infrastructure solutions for a broad range of AI -based applications. If you're creative, passionate about kubernetes and… more
- NVIDIA (Santa Clara, CA)
- NVIDIA is looking for a dedicated and motivated senior build and continuous integration (CI/CD) engineer for its GenAI Frameworks (NeMo, ... working on Large Language Models (LLM), Multimodal (MM), and Speech AI . NeMo provides end-to-end model training, including data curation, alignment, customization,… more
- NVIDIA (Santa Clara, CA)
- …both on-premises and cloud based. + 12+ years of proven experience with cluster management and related tools , including Docker Containers, Slurm, Kubernetes, and ... part of a team that's revolutionizing the field of AI with data center scale solutions? We are looking...are the voice of experience, using Kubernetes, SaaS, infrastructure-as-code tools , network debugging, and problem solving skills to help… more
- IBM (San Jose, CA)
- …talk. Your Role and Responsibilities IBM's Ceph[1] engineering organization is looking for a senior software engineer to join the CephFS team. In this role you will ... to higher-level APIs for integrating with other systems (OpenStack, OpenShift, an NFS-Ganesha cluster , Samba, etc). As a member of the CephFS engineering team, you… more
- NVIDIA (Santa Clara, CA)
- We are looking for a highly experienced AI Senior Software Test development engineer in NVIDIA's Deep Learning SWQA team. The position is in NVIDIA Deep Learning ... and AI Software Quality Assurance team that defines, develops and...in validating Data Center GPU based infrastructure (multi-GPUs, multi-nodes, cluster ) + Background in validating fault tolerance infrastructure +… more
- NVIDIA (Santa Clara, CA)
- Hardware Infrastructure is seeking a Senior Technical Program Manager to lead the strategy and execution of programs to support the bringup, operations and ... infrastructure we build and operate enables NVIDIAs most advanced AI and hardware researchers and engineers to create the...a fast paced and evolving landscape that requires a senior TPM leader to guide engineering roadmaps to be… more
- NVIDIA (Santa Clara, CA)
- …graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is a "learning machine" that ... parallel computing! More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...high-frequency and low-power CPUs, GPUs, SoCs at block level, cluster level, and/or full chip level, with a focus… more
- NVIDIA (Santa Clara, CA)
- …Deep Learning SWQA team. The position is in NVIDIA Deep Learning and AI Software Quality Assurance team that defines, develops and performs tests to validate ... healthcare, speech recognition, natural language processing, and a wide variety of other AI scenarios. We collaborate with multiple AI product teams to develop… more
- NVIDIA (Santa Clara, CA)
- …graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is a "learning machine" that ... parallel computing! More recently, GPU deep learning ignited modern AI - the next era of computing. NVIDIA is...Nvidia's GPUs, CPUs, DPUs and SoCs at block level, cluster level, and/or full chip level. + Work with… more
- NVIDIA (Santa Clara, CA)
- …can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and build our teams with ... Hardware Reliability Engineering for Electronics/Server Systems (graphics cards, server, rack, cluster ) from Concept to End-of-Life phase. + Establish, deliver and… more
- NVIDIA (Santa Clara, CA)
- …crowd: + Deep understanding of other infrastructure components like DNS, LDAP, NIS, Security Tools etc. + Experience with HPC cluster management tools such ... people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An...a focus on infrastructure automation. + Develop and maintain tools for collecting, analyzing, and visualizing data for reporting,… more