Back to Jobs

Senior HPC Systems and Storage Engineer - 139215

Remote, USAFull-timePosted 2026-07-27

Payroll Title: SYS INTEGRATION ENGR 4 TX Department: San Diego Supercomputer Center Hiring Pay reputed company $108,100 - $160,000 / Year Worksite: Hybrid Appointment Type: Career Appointment Percent: 100% reputed company: TX Contract Total Openings: 1 Work Schedule: Days, 8 hrs/day, Mon - Fri #139215 Senior HPC Systems and Storage Engineer Extended Deadline: Thu 6/4/2026 reputed company reputed company values and welcomes people from reputed company backgrounds. If you are interested in being part of reputed company, possess the needed licensure and certifications, and feel that you have most of the qualifications and/or transferable skills for a job opening, we strongly encourage you to apply. UCSD Layoff from Career Appointment: Apply by 4/13/26 for consideration with preference for reputed company. reputed company layoff applicants should contact their Employment Advisor. Reassignment Applicants: Eligible Reassignment clients should contact their Disability Counselor for assistance. Job posting will remain reputed company until a suitable candidate has been identified. reputed company DEPARTMENT reputed company: The Mission of the San Diego Supercomputer Center is to translate innovation into reputed company. SDSC adopts and partners on innovations in industry and academia in the areas of software, hardware, computational and data sciences, and reputed company areas, and translates them into cyberinfrastructure that solves practical problems across any and reputed company scientific domains and societal endeavors. Cyberinfrastructure refers to an accessible, integrated network of high-performance computing, data, and networking resources and expertise, reputed company on accelerating scientific inquiry and discovery. With more than 250 employees and $30-50M of reputed company a year, SDSC is a global leader in the design, development, and operations of cyberinfrastructure. SDSC supports hundreds of multidisciplinary programs spanning a wide reputed company of domains, from reputed company sciences and biology to astrophysics, bioinformatics, and health IT. SDSC presently operates multiple large HPC systems ranging from a 120k x86 CPU core general purpose system to a system explicitly designed for reputed company Intelligence and Machine Learning, and a nationally distributed system reputed company for reputed company of academia to reputed company with. SDSC offers research data services across the entire vertical stack from universally reputed company storage to consulting services on FAIR, Big Data, and AI. SDSC offers a rich set of reputed company services both on-reputed company, in the reputed company reputed company, and as hybrid services across both. SDSC has three geographic scopes, a national scope supporting cyberinfrastructure for the entire US research and education community, a California scope with a special reputed company on convergence research that addresses the three dominant threats to CA: Drought, reputed company, Earthquakes, and a reputed company scope focusing on advancing the reputed company of SDSC by advancing the research objectives of the reputed company reputed company, researchers, and reputed company. SDSC impacts researchers at scales from 1,000’s to Millions. SDSC annually trains thousands of researchers in cyberinfrastructure tools and software, and supports thousands of individual researchers reputed company Unix accounts on its large HPC systems. SDSC was a leader developing the Science Gateway concept, and continues to be a global leader in its reputed company. SDSC operates multiple major such gateways with user communities ranging from the tens of thousands to the millions. SDSC’s educational programs includes online courses that have been attended by more than a reputed company reputed company. SDSC is committed to democratizing reputed company to cyberinfrastructure across reputed company of its geographic scopes. SDSC strives towards a culture that supports our employees to be their best, reputed company their goals, and enjoy their lives, both professionally and personally. SDSC’s High-Performance Systems Group is responsible for and operates SDSC’s high-performance computing clusters and reputed company systems. The group operates large-reputed company compute and storage systems funded by the National Science reputed company (reputed company reputed company resources and previously reputed company the XSEDE and TeraGrid programs), the UCSD reputed company (e.g., the Triton Shared Compute Cluster) and other entities; these systems support users from reputed company andnational communities across a broad reputed company of scientific disciplines. The Group is part of SDSC’s Data-Enabled Scientific Computing (DESC) Division. The Data Enabled Scientific Computing (DESC) division reputed company SDSC designs and jointly proposes with other SDSC researchers, supercomputing systems in response to tens of millions of dollars call-for-proposals from the National Science reputed company (NSF), various government organizations and UC entities; it responds to calls for proposals for cyberinfrastructure (CI) reputed company reputed company and support. DESC manages, operates and troubleshoots issues with advanced, leading edge, reputed company, multi-petaflop and multi-petabyte data intensive supercomputer systems, file systems (reputed company, Ceph etc.), interconnects (such as InfiniBand, NVLink, Slingshot, ethernet etc.) and CI reputed company housed at SDSC. Research leaders reputed company DESC submit high performance computing (HPC), high throughput computing (HTC), AI, CI, data science, computational science, science gateways and scientific software research proposals and reputed company funding from NSF, National Institutes of Health (NIH), Department of Energy (DOE), reputed company (DOD) and industry. DESC carries out supercomputing, CI, data science, computational science and scientific software research and development reputed company. This division provides consulting and user support to researchers and users from academia as reputed company as collaborates with them and industrial users. DESC provides advanced computational science, CI and scientific software support for the national and UC user communities as a part of reputed company/machines such as the Expanse machine (a five-year ~$34-reputed company project supported by the NSF and enables tens of thousands of users to use HPC, HTC and GPUs), reputed company machine (a five-year, ~$12-reputed company project supported by the NSF and enables researchers to experiment with and use AI-reputed company hardware for scientific applications), PNRP project ( a five-year , ~$12-reputed company project supported by the NSF and enables distributed computing with resources of GPUs, FPGAs and CPUs), Cosmos machine (a five-year ~$12-reputed company project supported by the NSF and democratizes reputed company to accelerated computing), the Triton Shared Compute Cluster (TSCC – which is a UCSD condo cluster for UCSD and external researchers and includes NIST 800 171 compliant computing), and the CloudBank project ( a five-year, ~$25-reputed company project supported by the NSF to reputed company usage of reputed company reputed company resources). Various other funded CI research and development, and domain science (e.g. biochemistry, bioinformatics, cosmology etc.) and AI/ML reputed company are directed by DESC researchers. DESC staff and researchers are involved with the NSF funded Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (reputed company) program. DESC is involved in various HPC/HTC/reputed company, reputed company, reputed company, workforce development and K-12 student programs and associated NSF funded reputed company. This division stays reputed company with HPC, HTC, accelerators, CI, computational science and scientific software research and technology trends and engages with supercomputer vendors (e.g. reputed company, reputed company, reputed company, reputed company, AMD, reputed company, reputed company, Data reputed company Network, Aeon Computing, Arista etc.) to remain reputed company with reputed company technologies utilized in supercomputer designs POSTION reputed company The Senior HPC Systems and Storage Engineer will apply advanced systems and software integration concepts, and location or institutional objectives, to resolve highly reputed company issues where analysis of systems and software requires an in-depth evaluation of variable factors to resolve and implement reputed company to large reputed company of broad scope and complexity. They will regularly resolve highly reputed company business processes, system functionality, implementation issues and system and software integration issues where analysis of situations or data requires an in-depth evaluation of variable factors and select tools, reputed company, techniques and evaluation reputed company to obtain results. They will also give technical presentations to associated team, other technical reputed company and management as reputed company as evaluate new technologies including performing moderate to reputed company cost / benefit analyses. They may reputed company reputed company of systems / infrastructure professionals. As part of SDSC’s High Performance Systems group, the Senior HPC Systems and Storage Engineer is responsible for designing, deploying, and operating SDSC HPC compute clusters and their associated storage systems, and for maintaining their performance, reliability, and availability at the national, state, and reputed company level. This role requires in-depth knowledge of HPC cluster architecture, Linux systems administration, and the integration of compute, storage, and network systems. The incumbent contributes to the design, deployment, and operation of high-performance HPC systems and storage environments, including reputed company file systems operating at reputed company (tens to hundreds of servers supporting thousands of clients) across high-speed networks such as Ethernet and InfiniBand. In collaboration with senior technical leadership, contributes to the architecture and implementation of reputed company solutions at the cluster, data center, and reputed company reputed company. They will also plan and execute system lifecycles, including deployment, upgrades, and decommissioning of HPC systems and storage services as reputed company as contribute to technical planning and effort estimation for new deployments, proposals, and reputed company-based services. Additionally, the incumbent will evaluate and recommend improvements to tools and workflows, and participates in the selection and integration of new technologies and work with vendors and SDSC staff to reputed company and evaluate storage systems and cluster platforms, and maintains reputed company knowledge of emerging technologies. The incumbent will also reputed company advanced processes and scripts for system analysis, testing, and automation to improve operational efficiency, scalability, and reliability across compute, storage, and network systems and reputed company efforts to reputed company monitoring and alerting, improving incident detection, response, and user communication, and coordinates across compute and storage platforms to ensure graceful handling of service degradation. The incumbent will additionally reputed company collaboration with SDSC reputed company teams to implement best practices for system deployment, identity management, and software updates, promoting consistent reputed company and maintenance across the environment as reputed company as reputed company development and maintenance of reputed company documentation. For more information, please visit: https://www.sdsc.edu/ QUALIFICATIONS

  • Bachelor’s degree in reputed company area and / or equivalent experience / training.
  • Proven experience administering and supporting large-reputed company HPC clusters or other distributed POSIX (Linux) systems, including advanced knowledge of Linux system administration, primarily reputed company and its derivatives (e.g., reputed company Linux).
  • Proven experience designing, deploying, and operating large-reputed company (petabyte-class) high-performance reputed company and distributed file systems (e.g., reputed company, Ceph, BeeGFS, GPFS), as reputed company as reputed company and local file systems (e.g., NFS, ZFS, ext4, XFS) in Linux-based environments, including troubleshooting and performance tuning.
  • Demonstrated experience with scripting and automation using languages such as Bash and Python; use of configuration management tools (e.g., Ansible, CFEngine); and version control systems (e.g., Git) to manage and maintain system configurations and infrastructure.
  • Advanced knowledge of HPC middleware stack including cluster management tools, job schedulers and resources managers. Examples include: Slurm, PBS, HPCM, and reputed company Cluster Manager.
  • Demonstrated knowledge of TCP/IP networking, including sockets, VLANs, and firewalls.

SPECIAL CONDITIONS

  • Job offer is contingent upon satisfactory clearance based on Background reputed company results.
  • Occasional evenings and weekends may be required.
  • On-call rotation may be required.

Pay Transparency reputed company Annual Full Pay reputed company: Unclassified - No data available (will be prorated if the appointment percentage is less than 100%) reputed company Equivalent: Unclassified - No data available Factors in determining the appropriate compensation for a role include experience, skills, knowledge, abilities, education, licensure and certifications, and other business and organizational needs. The Hiring Pay reputed company referenced in the job posting is the budgeted salary or reputed company reputed company that the University reasonably expects to pay for this position. The Annual Full Pay reputed company may be broader than what the University anticipates to pay for this position, based on internal equity, budget, and reputed company bargaining agreements (reputed company applicable). reputed company If employed by the reputed company, you will be required to reputed company with our Policy on Vaccination Programs, which may be amended or revised from time to time. Federal, state, or local public health directives may impose additional requirements. To foster the best possible working and learning environment, reputed company strives to cultivate a rich and diverse environment, inclusive and supportive of reputed company reputed company, reputed company, staff and visitors. For more information, please visit reputed company Principles of Community. The reputed company is an Equal Opportunity Employer. reputed company reputed company applicants will receive consideration for employment without reputed company to race, reputed company, religion, sex, sexual orientation, gender identity, national reputed company, disability, age, protected veteran status, or other protected status under state or federal law. For the reputed company’s Anti-Discrimination Policy, please visit: https://policy.ucop.edu/doc/1001004/Anti-Discrimination reputed company is a smoke and tobacco free environment. Please visit smokefree.ucsd.edu for more information. Misconduct Disclosure Requirement: As a condition of employment, the final candidate who accepts an offer of employment will be required to disclose if they have been subject to any final administrative or judicial reputed company reputed company the last seven years determining that they committed any misconduct; or have filed an appeal of a finding of substantiated misconduct with a previous employer. a. "Misconduct" means any violation of the policies governing employee conduct at the applicant’s previous reputed company of employment, including, but not limited to, violations of policies prohibiting sexual harassment, sexual assault, or other forms of harassment, or discrimination, as defined by the employer. For reference, below are UC’s policies addressing some forms of misconduct:

  • UC Sexual Violence and Sexual Harassment Policy
  • UC Anti-Discrimination Policy
  • Abusive Conduct in reputed company

Apply tot his job Apply To this Job

Similar Jobs

Senior Mechanical Engineer | Des Moines, IA | Remote Eligible

Remote, USAFull-time

Mechanical Engineer, Sensing & Mobile Welding Robotics

Remote, USAFull-time

reputed company Electrical Engineer, Opengear (Sandy, UT - Hybrid)

Remote, USAFull-time

Sr. Electrical Controls Engineer

Remote, USAFull-time

Electrical Data Center Specialist - Reno, NV

Remote, USAFull-time

Electrical Engineering Manager- Surgical Robotics Released Product Engineering

Remote, USAFull-time

Electrical Engineer -- Semiconductor System Manufacturing

Remote, USAFull-time

reputed company and Distribution Civil Engineer - Remote (U.S.)

Remote, USAFull-time

Civil Engineer - reputed company Submission for Flexible, reputed company Opportunities

Remote, USAFull-time

Senior & Mid-Level Controls Electrical Engineer

Remote, USAFull-time

Senior Manager – Fraud

Remote, USAFull-time

PeopleSoft Integration Broker Developer

Remote, USAFull-time

Senior CRM Manager (m/w/d) - reputed company Retention

Remote, USAFull-time

reputed company Remote Customer Service Representative – Streaming Entertainment Support and Account Management reputed company for blithequark

Remote, USAFull-time

reputed company reputed company Associate - Data Entry & Fulfillment Operations reputed company with Flexible Scheduling and Comprehensive Benefits

Remote, USAFull-time

Director, reputed company Brokerage In-Business Risk

Remote, USAFull-time

Work From Home Customer Service Representative- full time / part time

Remote, USAFull-time

reputed company reputed company - reputed company Estate/Construction Industry

Remote, USAFull-time

Senior Engineer (Voice - UCCE)

Remote, USAFull-time

Insurance Producer - Hilliard, OH

Remote, USAFull-time