Somerset, NJ · Tech Lead at Regeneron · US Patent Holder

Siddhesh Salunke.

Tech Lead & Architect  /  Drug Discovery · HPC · AI · Data · Cloud

I architect the platforms that move tens of petabytes of biotech data from instrument to insight — accelerating drug discovery for the scientists who use them. Co‑inventor on US Patent 12,254,347 for the HPC Data Pipeline.

No. 01 About

Transforming biotech data landscapes through platform engineering.

I lead High Performance Computing engineering and operations at Regeneron Pharmaceuticals, where I architect enterprise-scale platforms that accelerate drug discovery. My work spans cloud-native data platforms processing more than 25 petabytes of biotech data, GxP-compliant governance frameworks, and the AI infrastructure that researchers rely on every day.

Over seven years at Regeneron I have progressed from Cloud DevOps Engineer through Senior Cloud Data Engineer and Platform Engineering Lead into my current role leading the HPC team. The platforms my teams build now serve more than 2,000 scientists and business users, with measurable outcomes: a 35% reduction in operational cost and a 25% improvement in research cycle times.

I hold a Master’s in Management Information Systems from Pace University and a Bachelor of Engineering from PVPPCOE, Mumbai. I am a Certified SAFe® 6 Practitioner, an AWS Certified Solutions Architect, and a co-inventor on US Patent 12,254,347 — the HPC Data Pipeline that became the foundation for how Regeneron moves cryo-electron microscopy and sequencing data into the cloud.

25+ PB biotech data processed across platforms
2,000+ scientists & users served by the platform
35% operational cost reduction delivered
10+ yrs from individual contributor to tech lead
No. 02 Granted Patent
United States Patent US 12,254,347 B2 Granted · March 18, 2025
Read on Google Patents

HPC Data Pipeline.

A scalable cloud-based data processing and computing platform for large-volume scientific data — designed to take cryo-electron microscopy from instrument to 3D model in a fraction of the time of conventional sync workflows.

~4× throughput vs. baseline DataSync pipeline
1 TB / hr sustained raw cryo-EM transfer rate
11 co-inventors at Regeneron

Abstract

A multi-stage data transfer pipeline that introduces a dynamically generated filter into the prepare phase of cloud sync, limiting scans to only the files present at the staging location rather than recursing the full source and destination. Combined with on-demand FSx Lustre provisioning, multi-queue GPU and CPU scheduling, and a self-service dataset management utility, the system collapses the time from microscope to interpretable 3D structure.

No. 04 Stack

The tools in rotation.

Grouped by capability rather than by how they look on a résumé.

Cloud & Infra

  • AWS (EC2, S3, EKS, Lambda, SSM)
  • Google Cloud Platform
  • Terraform / IaC
  • Ansible & Ansible Tower
  • Docker / Kubernetes / ECS

Data & Analytics

  • Databricks (enterprise adoption)
  • Apache Spark / EMR
  • Apache Airflow / MWAA
  • Hadoop ecosystem
  • Dataiku, Superset

HPC & Science

  • AWS ParallelCluster / CfnCluster
  • Slurm & SGE
  • FSx Lustre
  • Seqera Platform (Nextflow)
  • Posit Workbench / Connect
  • Enterprise JupyterHub

Delivery & Practice

  • Certified SAFe® 6
  • Jenkins / BitBucket / Jira
  • CI/CD pipelines
  • GxP-compliant governance
  • MLOps / model deployment
No. 05 Experience

Where I’ve shipped.

  1. Aug 2025 — Present

    Tech Lead, HPC Engineering & Operations

    Regeneron Pharmaceuticals

    Lead the High Performance Computing team responsible for managing the clusters behind Regeneron’s scientific pipelines. Own advanced data science platforms including Posit Workbench and Enterprise JupyterHub, and partner closely with researchers on molecular dynamics, fluid dynamics, protein structure analysis, and gene expression workflows.

    • HPC
    • Posit
    • JupyterHub
    • Slurm
    • Scientific Computing
  2. Nov 2024 — Aug 2025

    Lead, Platform Engineering

    Regeneron Pharmaceuticals · Sleepy Hollow, NY

    Led feature engineering across the Regeneron Data Platform spanning Data Acquisition, Processing, Orchestration, Lakes, Fabric, and Metadata. Drove enterprise-wide adoption of the Databricks ecosystem and oversaw the product engineering practice for shared data platform services.

    • Data Platform
    • Databricks
    • Data Mesh
    • Product Engineering
  3. Dec 2021 — Dec 2024

    Senior Cloud Data Engineer

    Regeneron Pharmaceuticals · Sleepy Hollow, NY

    Led implementation of Data Access Controls and the Search Platform, strengthening data security and discoverability across the organization. Spearheaded the development of an AI-based mobile platform and managed both Data Platform Operations and Platform Engineering teams through delivery.

    • Data Access
    • Search
    • AI Platform
    • Team Lead
  4. Dec 2020 — Nov 2021

    Cloud DevOps Engineer II — HPC & BigData

    Regeneron Pharmaceuticals · Sleepy Hollow, NY

    Built and operated the HPC and BigData platforms behind scientific research: AWS ParallelCluster, Slurm, SGE, Apache Airflow on MWAA, Spark on EMR, Dataiku, and Superset. Release management on Jenkins, Docker, EKS, BitBucket, and AWS SSM. The period during which the HPC Data Pipeline patent work originated.

    • ParallelCluster
    • Airflow
    • EMR
    • EKS
    • Patent Work
  5. May 2018 — Nov 2020

    Cloud DevOps Engineer · Associate Cloud DevOps Engineer

    Regeneron Pharmaceuticals · Sleepy Hollow, NY

    Joined Regeneron’s cloud team. Built CI/CD, configuration management, and infrastructure-as-code patterns that became the foundation for later platform work, and earned the trust of researchers and central IT — the relationships that make scientific computing systems actually land.

    • AWS
    • Ansible
    • Jenkins
    • Docker
  6. Jun 2017 — Apr 2018

    Cloud Solutions Architect

    The Institute for Criminal Justice Training Reform

    Built an automated pipeline to deploy a Python application on Docker via GitHub, AWS ECS, and Jenkins. Configured Ubuntu on EC2 with MySQL and seed data, and moved deeply into Ansible playbooks and configuration management.

    • AWS ECS
    • Ansible
    • Jenkins
    • Python
  7. 2014 — 2016

    Earlier roles

    Green Catapult · Waghmare Designs · Trivia Softwares

    Full-stack development on Meteor and Node.js, MongoDB backends, business intelligence dashboards in Power BI and Datazen, real-time sentiment analysis on Twitter data, and a comprehensive information management system on Oracle 11g. The cross-stack foundation that the rest of the career builds on.

    • Node.js
    • MongoDB
    • Power BI
    • Oracle

Education & credentials

  • M.S. · Management Information Systems · Pace University
  • B.E. · Electrical, Electronics & Communications · PVPPCOE, Mumbai
  • USPTO · Co-inventor, US 12,254,347 B2 (HPC Data Pipeline)
  • Google Cloud · Certified Generative AI Leader (2026)
  • NVIDIA · Certified Associate — AI Infrastructure & Operations (2026)
  • Certified SAFe® 6 · Practitioner
  • AWS · Certified Solutions Architect(2016-2019)
  • Courses · Ansible Essentials · Linux Command Line · Intro to R · DevOps CI/CD on AWS
  • Honors · International Graduate Scholarship Recipient
No. 06 Selected Work

Things I’ve built.

Patented · Production

HPC Data Pipeline

On-premises instrument data is staged on NetApp, filtered by a Lambda fetch, moved via AWS DataSync into S3, then attached to on-demand FSx Lustre for ParallelCluster compute. Replacing a full source-and-destination scan with a dynamically generated filter brings dataset prep from roughly three hours per terabyte down to about thirty-four minutes. Co-invented and deployed at Regeneron.

  • AWS
  • Lambda
  • DataSync
  • FSx Lustre
  • Airflow
Patent on Google Patents ↗
Featured by Posit · Production

WatchtowR — Observability for Scientific Computing

A Shiny application on Posit Connect that unifies observability across seven Posit environments on AWS, pulling live metrics from the Posit APIs and AWS into a single interactive command center. It tracks license utilization, CPU and memory consumption, deployment health, and IDE adoption — cutting a 22-hour manual license audit to seconds and giving engineering leaders the evidence to make confident infrastructure decisions. Built by my Scientific Computation team and published as a Posit customer story.

  • Shiny
  • Posit Connect
  • R
  • AWS
  • Observability
  • MCP (next)
Read the Posit customer story ↗
Enterprise Platform

Regeneron Data Platform

Led platform engineering across Data Acquisition, Processing, Orchestration, Lakes, Fabric, and Metadata. Serves more than 2,000 scientists and business users with GxP-compliant governance, and contributed to organization-wide Databricks adoption.

  • Databricks
  • GxP
  • Data Mesh
  • 2000+ users
Scientific Computing

Posit & JupyterHub for Science

Advanced data science platforms — Posit Workbench, Posit Connect, and Enterprise JupyterHub — running on shared HPC infrastructure. Supports molecular dynamics, fluid dynamics, and protein structure analysis workflows used in research daily.

  • Posit
  • JupyterHub
  • HPC
  • R / Python
Security & Discovery

Data Access & Search Platform

Implemented Data Access Controls and the enterprise Search Platform — strengthening data security across the organization while making the right datasets discoverable for the people who need them.

  • Access Controls
  • Search
  • Governance
AI Platform

AI-based Mobile Development Platform

Spearheaded development of an AI-based platform driving innovation in mobile application capabilities for internal users — bridging data science models with real product surfaces.

  • AI
  • Mobile
  • MLOps
No. 07 Contact

Let’s build the Solutions that lets the interesting work happen.

Open to conversations on data platforms, AI infrastructure, biotech engineering, cloud architecture, drug discovery, leadership, or trading notes on HPC at scale.