All Work
B

BIOMEX Genomics Analytics Platform

Analytics Platform & Team LeadershipSep–Dec 2025

How can researchers run complex genomic analysis without depending on manual engineering support?

Role
Project Lead (Development Director)
Type
Team project, Dalhousie University
Status
Proof of concept
Stakeholder
Dalhousie University faculty researcher

Results

Cross-functional team led
10

Directors, senior and junior developers delivering to a faculty client.

Samples per run (target capacity)
60–90K

Architecture designed for real-world omics datasets of this scale.

Analysis pipeline delivered
DESeq2

End-to-end proof of concept: MA plot, volcano plot, and heatmap outputs, verified against expected results.

In Brief

Led a 10-person team rebuilding a desktop bioinformatics tool as a self-service web platform for a Dalhousie researcher, and delivered a working DESeq2 pipeline as a verified proof of concept.

Context

BIOMEX began as a legacy PyQt5 desktop application: powerful, but locked to one machine and dependent on manual engineering support to run analyses. A Dalhousie faculty researcher wanted it rebuilt as a self-service web platform so researchers could run genomic analyses themselves.

The semester’s goal, agreed with the client, was pragmatic: prove the approach works end to end by getting one complete analysis method running accurately in the new architecture, while building the frontend and documentation that future teams would extend.

This is a real team project with a real stakeholder. I led a 10-person team as Project Lead; the honest status is a validated proof of concept, not a finished product.

Data

Source
Researcher-supplied omics data: a raw counts matrix plus a sample metadata file (CSV)
Period
Per-analysis uploads (no fixed dataset)
Records
Designed for real-world datasets of roughly 60,000–90,000+ samples per run
Grain
Counts matrix (samples × genes) paired with per-sample metadata

Quality & Handling

  • File compatibility is verified in two stages: a fast client-side dimensional check (metadata rows vs counts columns), then a server-side value-level match of sample identifiers.
  • Large files are validated server-side to avoid overwhelming the browser.
  • The DESeq2 pipeline’s outputs were verified against expected results before acceptance.

Approach

  1. Upload counts + metadata
  2. Validate & match samples
  3. POST to R Plumber API
  4. Run DESeq2 (R)
  5. Encode plots (base64)
  6. Render results in React

The React frontend (hosted on Firebase) posts CSV contents to an R Plumber API on a faculty VM; the API runs DESeq2.R, then returns the generated plots as base64 images the app renders directly.

Validation

  • Verified against expected outputs

    The DESeq2 proof of concept produced MA, volcano, and heatmap outputs checked against the expected results from the reference implementation.

  • Client acceptance

    The faculty stakeholder confirmed the proof of concept and the project’s viability for continued development.

  • Isolated execution

    Each request reconstructs CSVs in an isolated server folder so concurrent users cannot overwrite each other’s data.

Limitations

  • A one-semester proof of concept: frontend ~80% complete, backend ~10%, with one analysis method (DESeq2) fully working.
  • The backend ran on a faculty VM exposed via ngrok for development — not production-hardened hosting.
  • Additional methods (edgeR, limma), preprocessing, and single-cell analysis were scoped for future teams, not delivered.

My Contribution

  • Owned: translating faculty requirements into sprint priorities and the overall application architecture.
  • Owned: coordinating delivery across frontend, backend, and bioinformatics contributors, keeping milestones on schedule.
  • Team: nine other developers built individual components; the client supplied domain expertise and the legacy PyQt5 reference implementation.
  • Inherited: DESeq2 statistical logic came from the Bioconductor ecosystem, not written from scratch.

Technical Notes

Architecture

React frontend on Firebase → R Plumber API (POST /run_deseq2, port 8000) on the Faculty of Computer Science VM → DESeq2.R. Plots are saved as PNGs, base64-encoded, and returned as JSON for the React app to render.

Delivery status & handoff

Frontend ~80% complete; backend ~10% (one working method). The closing document hands future teams a prioritized roadmap: complete the preprocessing pipeline (the blocking dependency), then add edgeR/limma, harden hosting (replace ngrok), and lock down CORS.

What I owned vs the team

I owned requirements translation into sprint priorities, the application architecture, and delivery coordination. Nine other developers built individual components; the client supplied domain expertise and the legacy reference; the DESeq2 statistical logic came from the Bioconductor ecosystem.

Sources

Data
Researcher-supplied omics data: a raw counts matrix plus a sample metadata file (CSV)
Period
Per-analysis uploads (no fixed dataset)
Volume
Designed for real-world datasets of roughly 60,000–90,000+ samples per run

Tables

Tables and inputs used
counts matrixGene expression counts; samples as columns, genes as rows.input
metadataPer-sample attributes (condition, reference level).input
DESeq2 outputsMA plot, volcano plot, and heatmap.output

Figures on this page

  • 10 Cross-functional team ledResume, 15_Bioinformatics_Closing_document.pdf
  • 60–90K Samples per run (target capacity)Resume, closing document
  • DESeq2 Analysis pipeline deliveredclosing document