- Micron Technology, Inc. (San Jose, CA)
- …position in the Artificial Intelligence ( AI ), Machine Learning (ML) and High Performance Computing ( HPC ) business segments. You will be working on innovative ... you will be charged with defining and accomplishing the strategy for a High Performance Memory product portfolio that will further fortify Micron's leadership… more
- Micron Technology, Inc. (San Jose, CA)
- …in growing the Artificial Intelligence ( AI ), Machine Learning (ML) and High- Performance Computing ( HPC ) business segments. You will be working on innovative ... of Work (SOWs), business term sheets, and other customer-facing documents for high- performance memory products. + Represent the Product Management team in Product… more
- Meta (Menlo Park, CA)
- …fabric and host networking, comms lib and scheduling infrastructure. **Required Skills:** AI / HPC Systems Performance Engineer Responsibilities: 1. ... **Summary:** Meta's AI Training and Inference Infrastructure is growing exponentially...workloads that expects a loss-less fabric interconnect. To improve performance of these systems we constantly look… more
- Meta (Menlo Park, CA)
- …and host networking, communications lib and scheduling infrastructure. **Required Skills:** AI / HPC System Performance Engineer Responsibilities: 1. Lead ... **Summary:** Meta's AI Training and Inference Infrastructure is growing exponentially...a loss-less fabric interconnect with minimal latency. To improve performance of these systems we constantly look… more
- Meta (Menlo Park, CA)
- …These workloads expect a loss-less fabric interconnect with minimal latency. To improve performance of these systems we constantly look for opportunities across ... host networking, communications lib and scheduling infrastructure. **Required Skills:** AI / HPC Network Engineering Manager Responsibilities: 1. Manage engineers… more
- IBM (San Jose, CA)
- …technical areas in the context of hybrid cloud, AI systems , networking, security, high-speed networked-storage, accelerators, and HPC principles. The ... focuses on the next generation Hybrid Cloud infrastructure for AI , Storage, HPC and Quantum applications. The...Experience with GPU Systems * Familiarity with HPC system performance evaluation. * Familiarity with… more
- Meta (Menlo Park, CA)
- …on existing accelerator systems and guiding the future of models and AI HW at Meta. This drives improved performance , new model architectures and ... the following areas: Accelerators/GPU architectures, High Performance Computing ( HPC ), Machine Learning Compilers, Training/Inference ML Systems , Model… more
- Meta (Menlo Park, CA)
- … AI product introductions and AI operations initiatives supporting Meta's growing AI / HPC infrastructure for our Family of Apps . They will be responsible ... deliver on shared goals 10. The ideal candidate will have experience in AI / HPC product development and operations, demonstrated experience in the Network… more
- Deloitte (San Jose, CA)
- …cancer detection, drug discovery, optimizing population health and clinical trials, autonomous systems and edge AI , and renewable energy. Key responsibilities: + ... in the cloud or on prem + Adopt best engineering practices in automation, HPC and AI /GenAI infrastructure and design patterns + Define and lead technology… more
- Microsoft Corporation (Mountain View, CA)
- …the boundaries of scale, performance , and deployment, creating frontier AI systems that power transformative experiences across Microsoft. The Multimodal ... (Pandas, NumPy, etc.) + OR equivalent experience. + **Experience with large-scale AI systems ** - design and deployment of distributed architectures, multimodal… more
- General Motors (Mountain View, CA)
- **Job Description** **The Role:** The **Staff AI Engineer** is responsible for supporting and reinforcing the adoption of AI Software Engineering across the MFG ... in the Manufacturing space relative to manufacturing productivity/efficiency. The **Staff AI Engineer** will drive the identification, evaluation, and adoption of… more
- IBM (San Jose, CA)
- …technical areas in the context of hybrid cloud, AI systems , networking, security, high-speed networked-storage, accelerators, and HPC principles. The ... focuses on the next generation Hybrid Cloud infrastructure for AI , Storages, HPC and Quantum applications. The...experience with Git * HPC : experience running HPC workloads on HPC systems … more
- Meta (Menlo Park, CA)
- …following machine learning/deep learning domains: Distributed ML Training, GPU architecture, ML systems , AI infrastructure, high performance computing, ... large-scale GPU training and inference fleet through an observable, reliable and high- performance distributed AI /GPU communication stack. Currently, one of the… more
- Microsoft Corporation (Mountain View, CA)
- …and external, and operate at the intersection of AI algorithmic innovation, purpose-built AI hardware, systems , and software. We are a team of highly capable ... The Artificial Intelligence Cloud Inference team at Microsoft develops AI software that enables running AI models...+ Speeding up/reducing complexity of key components/pipelines to improve performance and/or efficiency of our systems +… more
- Microsoft Corporation (Mountain View, CA)
- …high- performance GPUs, ultra-low-latency NVLink/NVSwitch networks, and innovative liquid-cooling systems . Our team is seeking a Member of Technical Staff, ... Hardware Health, to ensure these systems deliver sustained reliability, performance , and availability...predictive health models, failure detection frameworks, and autonomous remediation systems that keep our AI clusters operating… more
- Cisco (Milpitas, CA)
- …team engaged in the design, development and execution of tests to qualify network performance for AI .ML capability. In this role you'll have opportunity to: + ... the next generation infrastructure to meet the needs of AI /ML workloads and continuously increasing internet users and application....Quality of Service (QoS) policies to ensure optimal network performance + Exposure to RDMA, HPC networks… more
- Cisco (Milpitas, CA)
- …agile team engaged in the design, development and execution of tests to qualify network performance for AI /ML capability. You will be a part of our solutions ... a customer-facing environment + Previous experience leading teams + Exposure network operating systems , preferably SONiC + Exposure to RDMA, HPC networks +… more
- Meta (Menlo Park, CA)
- …levels 9. Experience in leading teams working on high performance computing ( HPC ) and AI /ML systems , including: 10. GPU/ASIC-based kernel development and ... systems for our fleet 4. Technical management 5. Experience in systems architecture, performance , workload-analysis and large scale distributed systems … more
- Meta (Menlo Park, CA)
- …10. Experience in leading teams working on high performance computing ( HPC ) and AI /ML systems , including: GPU/ASIC-based kernel development and ... ROCm), distributed systems for large scale training and serving, and systems architecture and performance 11. Accelerator (GPU/ASIC) kernel development and… more
- Broadcom (San Jose, CA)
- …compiler toolchains. + Experience analyzing and tuning performance for a variety of AI /ML and HPC workloads. + Deep knowledge of Linux kernel and Linux ... Description:** **Job Description** Ethernet NIC product portfolio is designed for high performance computing and networking applications including AI and ML.… more