Bioinfo-Comput Biologist Inter
University of Michigan
Ann Arbor, MI
ID: 7368172
Posted: Newly posted
Application Deadline: Open Until Filled
Job Description
Job Summary
The Parker Lab at the University of Michigan (http://theparkerlab.org) is hiring a Bioinformatician Intermediate to stand up and operate the Apache Texera platform at Michigan as part of BRIDGE, a new NSF-funded national center. Texera is an open-source, browser-based system that lets scientists build and run data analysis workflows without writing code. Our team leads the metabolic traits domain of the center.
The core of the job is deployment and operations: getting Texera running reliably on Michigan?s high-performance computing infrastructure, extending it to AWS as demand grows, and keeping it working for the metabolic trait researchers who use it. You will also port our single-cell and single-nucleus multi-omic pipelines onto the platform as reusable workflows, and answer questions from users when they run into trouble.
This is a hands-on, build-it role. Dr. Steve Parker sets scientific direction and Dr. Ha Vu provides day-to-day supervision. You will work directly with the Apache Texera engineering team at UC Irvine, who have a working reference deployment, and with Michigan?s research computing staff.
Responsibilities*
Deploy and operate Texera on Michigan HPC infrastructure, integrating its architecture with a Slurm-managed cluster.
Extend the deployment to AWS to provide elastic compute as user demand grows.
Maintain the deployment: monitoring, upgrades, storage, authentication, troubleshooting, and cost management.
Implement our single-cell and single-nucleus multi-omic pipelines (snRNA-seq, snATAC-seq, multiome) as containerized, reusable Texera workflows, and scale them to atlas-level and population-scale datasets.
Provide technical support to platform users and triage issues, escalating upstream to the Texera team where appropriate.
Contribute fixes and operators upstream to Apache Texera, and document the Michigan configuration so it is reproducible.
Required Qualifications*
Master's or PhD in computational biology, bioinformatics, computer science, or a closely related field. Equivalent research software engineering experience will be considered.
Demonstrated experience deploying and maintaining containerized services, including Docker or Singularity.
Strong Linux systems skills and substantial hands-on experience in an HPC environment, particularly Slurm.
Strong programming skills in Python and/or R, with version control, testing, and documentation as habits.
Experience building reproducible analysis pipelines, for example with Nextflow, Snakemake, or WDL.
Working knowledge of sequencing data analysis, sufficient to build and debug genomics workflows and answer user questions about them.
Evidence of independent delivery: a track record of taking a loosely specified goal to a working, documented, running result with limited supervision.
Ability to explain technical problems clearly to scientists without computational training.
English language proficiency.
Desired Qualifications*
AWS experience, including EKS, EFS, and cost-aware resource provisioning.
Kubernetes networking and configuration management, for example Helm, VXLAN overlays, Ansible, or Terraform.
Depth in single-cell or single-nucleus data analysis. We prioritize methodological understanding over familiarity with specific tools.
Experience with workflow platforms such as Texera or Galaxy.
Java or Scala, which would let you contribute native operators to the Texera codebase.
Prior open-source contribution in a public repository.
Prior work on metabolic, endocrine, or cardiometabolic disease.
The Parker Lab at the University of Michigan (http://theparkerlab.org) is hiring a Bioinformatician Intermediate to stand up and operate the Apache Texera platform at Michigan as part of BRIDGE, a new NSF-funded national center. Texera is an open-source, browser-based system that lets scientists build and run data analysis workflows without writing code. Our team leads the metabolic traits domain of the center.
The core of the job is deployment and operations: getting Texera running reliably on Michigan?s high-performance computing infrastructure, extending it to AWS as demand grows, and keeping it working for the metabolic trait researchers who use it. You will also port our single-cell and single-nucleus multi-omic pipelines onto the platform as reusable workflows, and answer questions from users when they run into trouble.
This is a hands-on, build-it role. Dr. Steve Parker sets scientific direction and Dr. Ha Vu provides day-to-day supervision. You will work directly with the Apache Texera engineering team at UC Irvine, who have a working reference deployment, and with Michigan?s research computing staff.
Responsibilities*
Deploy and operate Texera on Michigan HPC infrastructure, integrating its architecture with a Slurm-managed cluster.
Extend the deployment to AWS to provide elastic compute as user demand grows.
Maintain the deployment: monitoring, upgrades, storage, authentication, troubleshooting, and cost management.
Implement our single-cell and single-nucleus multi-omic pipelines (snRNA-seq, snATAC-seq, multiome) as containerized, reusable Texera workflows, and scale them to atlas-level and population-scale datasets.
Provide technical support to platform users and triage issues, escalating upstream to the Texera team where appropriate.
Contribute fixes and operators upstream to Apache Texera, and document the Michigan configuration so it is reproducible.
Required Qualifications*
Master's or PhD in computational biology, bioinformatics, computer science, or a closely related field. Equivalent research software engineering experience will be considered.
Demonstrated experience deploying and maintaining containerized services, including Docker or Singularity.
Strong Linux systems skills and substantial hands-on experience in an HPC environment, particularly Slurm.
Strong programming skills in Python and/or R, with version control, testing, and documentation as habits.
Experience building reproducible analysis pipelines, for example with Nextflow, Snakemake, or WDL.
Working knowledge of sequencing data analysis, sufficient to build and debug genomics workflows and answer user questions about them.
Evidence of independent delivery: a track record of taking a loosely specified goal to a working, documented, running result with limited supervision.
Ability to explain technical problems clearly to scientists without computational training.
English language proficiency.
Desired Qualifications*
AWS experience, including EKS, EFS, and cost-aware resource provisioning.
Kubernetes networking and configuration management, for example Helm, VXLAN overlays, Ansible, or Terraform.
Depth in single-cell or single-nucleus data analysis. We prioritize methodological understanding over familiarity with specific tools.
Experience with workflow platforms such as Texera or Galaxy.
Java or Scala, which would let you contribute native operators to the Texera codebase.
Prior open-source contribution in a public repository.
Prior work on metabolic, endocrine, or cardiometabolic disease.


