Senior Data Engineer β Clinical Platforms (Databricks)
Remote
Location eligibility not specified
π Join Our Remote Data Products & Machine Learning Startup! π
At Muttdata, we build innovative Data Products and Machine Learning solutions that help companies solve complex business challenges. As a fast-growing, remote-first startup, we're passionate about technology, collaboration, and continuous learning.
We are looking for an experienced, ownership-driven Senior Data Engineer to join our team πΆπ. You'll lead the architecture, design, and implementation of a next-generation, in-house clinical trial software platform built directly on Databricks, bridging the gap between software development and large-scale data engineering.
This role works closely with frontend developers, software architects, clinical research teams, and Clinical QA and Validation teams. It requires solid hands-on experience with Databricks and a strong understanding of clinical data standards and regulated environments. Ownership, clear communication, and the ability to build robust, compliant, and scalable solutions are essential to succeed in this fast-paced, collaborative environment.
At Muttdata, we build innovative Data Products and Machine Learning solutions that help companies solve complex business challenges. As a fast-growing, remote-first startup, we're passionate about technology, collaboration, and continuous learning.
We are looking for an experienced, ownership-driven Senior Data Engineer to join our team πΆπ. You'll lead the architecture, design, and implementation of a next-generation, in-house clinical trial software platform built directly on Databricks, bridging the gap between software development and large-scale data engineering.
This role works closely with frontend developers, software architects, clinical research teams, and Clinical QA and Validation teams. It requires solid hands-on experience with Databricks and a strong understanding of clinical data standards and regulated environments. Ownership, clear communication, and the ability to build robust, compliant, and scalable solutions are essential to succeed in this fast-paced, collaborative environment.
π What We Do
- Leveraging our expertise, we build modern Machine Learning systems for demand planning and budget forecasting.
- Developing scalable data infrastructures, we enhance high-level decision-making, tailored to each client.
- Offering comprehensive Data Engineering and custom AI solutions, we optimize cloud-based systems.
- Using Generative AI, we help e-commerce platforms and retailers create higher-quality ads, faster.
- Building deep learning models, we enhance visual recognition and automation for various industries, improving product categorization, quality control, and information retrieval.
- Developing recommendation models, we personalize user experiences in e-commerce, streaming, and digital platforms, driving engagement and conversions.
π Our Partnerships
- Amazon Web Services
- Astronomer
- Databricks
π Our Values
- π We are Data Nerds
- π€ We are Open Team Players
- π We Take Ownership
- π We Have a Positive Mindset Β π Curious about what weβre up to? Check out our case studies and dive into our blog post to learn more about our culture and the exciting projects weβre working on! π
Responsibilities π€
- Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks (PySpark, Spark SQL, Delta Lake) to power custom clinical application backends.
- Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers for proprietary clinical trial software.
- Build performant, standards-compliant data structures to store EDC outputs, audit trails, device telemetry, and patient-reported outcomes, enabling rapid querying and downstream analytics.
- Partner with Clinical QA and Validation teams to ensure database structures, data pipelines, and clinical data repositories comply with GxP, 21 CFR Part 11, HIPAA, and GDPR.
- Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices into a unified Databricks Lakehouse architecture.
- Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times to support high-throughput, low-latency clinical platform workflows.
Required Skills π»
- 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
- Demonstrated experience building, extending, or maintaining custom software applications for clinical trials (e.g., custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms).
- Deep understanding of clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
- Strong experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
- Proficiency in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
- Experience working in cloud environments (AWS preferred, Azure or GCP).
- Advanced English to discuss technical requirements and solutions with clients in the United States
Nice to have π»
- Bachelor's or Master's degree in Computer Science, Data Engineering, Bioinformatics, or a related quantitative field.
- Experience with Databricks Workflows, Delta Live Tables (DLT), and Unity Catalog governance.
- Background working in a validated system environment (Computer System Validation / CSV).
π Perks
- Remote-first culture β work from anywhere! π
- AWS, DBT, Google Cloud, Azure & Databricks certifications fully covered
- In-Company English Lessons.
- Birthday off + an extra vacation week (Mutt Week! ποΈ)
- Referral bonuses β help us grow the team & get rewarded!
- Maslow: Monthly credits to spend in our benefits marketplace.
- βοΈποΈ Annual Mutters' Trip β an unforgettable getaway with the team!
- πΆ Monthly Childcare ReimbursementΒ β Because supporting families matters too
Source: the employer's careers page. Last checked 2026-10-02. Posted 2026-10-01.