Genomes, Prompts, and Shiny

Rethinking How We Explore Biological Data

Samuel Bharti

Computational Biologist · Doctoral Researcher · Intern, Shiny team at Posit

Biologists have a data problem. I build the tools.

You all know Shiny. So today is about what I point it at.

Their tools didn’t keep up. That gap is the job.

GENOMES

This is one tissue sample.

Every dot is a single molecule of RNA.

42 million of them, and the position is real.

The ten apps I built this summer

tahoe-explorerPlan a reanalysis before spending compute.
plotomics-liveTwenty-six figures, two engines.
variant-reviewerOne variant, without eighteen tabs.
genescout200 candidates to a ranked thirty.
gene-list-builderA disease panel from seven sources.
recount-explorerrecount3, without writing any code.
de-explorerWhich genes changed, and by how much.
signature-scoringWhich pathways are active where.
drug-perturbationDrugs that reverse a disease signature.
genome-explorerMutations in genomic context.

And five packages underneath them

Each came out of a problem the apps kept hitting. Solve it in the package, and nothing above it has to.

biobouncerwhether an identifier means anything
biohttphow we talk to a service
bioclientswhat each biological service returns
plotomicsone rendering core, three languages
biocohorthow a study stays organised

Twenty-six bespoke figures

Plotomics Live. Every one of these is a component you can call today, from R, Python or React.

Single-cell & spatial6
Gene expression4
Cancer genomics6
Genome & epigenome6
Structure & networks4
R ShinyPython ShinyShiny ReactGPUggplot2

The same page, two engines

Plotomics Live One toggle swaps the GPU React component for the ggplot2 image, over the same server-side computation.

R stays the analytical engine. The browser does the drawing it is good at.

Tahoe-100M

One public dataset. Free to anyone. And effectively nobody can open it.

379

drugs

50

human cell lines

100.6M

cells, one row each

You planned it before you paid for it

Picked a drug

Saw the cell lines that exist.

Saw the doses

And which ones are too thin to trust.

Left with a recipe

R on one tab, Python on the other.

Downloaded nothing

Every number computed where the file lives.

The app wrote this. Not an example, the export from the run you just watched.

library(duckdb); library(DBI)
con <- dbConnect(duckdb())
dbExecute(con, "INSTALL httpfs; LOAD httpfs;")

obs <- dbGetQuery(con, "
  SELECT * FROM read_parquet('hf://datasets/tahoebio/Tahoe-100M/metadata/obs_metadata.parquet')
  WHERE cell_name IN ('HS-578T', 'BT-474')
")

3,094,027 cells across 1,344 samples. The other tab is the same thing in Python.

The front door for a dataset that didn’t have one.

GeneScout

PROMPTS

Two hundred candidate genes,
reduced to 30.

UI built with Claude Design

One disease name, one defensible panel

Gene List Builder Breast cancer resolves to MONDO:0007254, and seven sources answer in parallel.

The weights are visible and adjustable. Every gene traces back to its source database.

variant-reviewer

A different shape of problem: scattered, not big.

That is a variant.
One letter, out of 3,000,000,000.

referenceACGTAC
this patientACATAC

One search box at the top.
Eighteen databases, underneath.

Five focused applications

Recount Explorer, DE Explorer, Signature Scoring, Drug Perturbation and Genome Explorer. One common workflow each, wrapped so somebody else can run it.

An analysis can end in a tool, not in a script the next person has to rerun.

One catalogue, 18,998 studies

Recount Explorer recount3, human and mouse. Browse the catalogue, check quality, run a PCA, export what you need, without writing any code.

The data was already public. Opening it should never have needed a script first.

Which genes changed, and by how much

DE Explorer Real TCGA lung data: adenocarcinoma against squamous-cell, plus normal tissue.

Move a threshold, every panel follows. No re-running anything.

Which pathways are active, and where

Signature Scoring Breast cancer, 50 Hallmark pathways, one subtype against another.

A single gene is noisy. A whole pathway is a signal.

Find a drug that reverses the disease

Drug Perturbation Connectivity scoring, in the spirit of CMap and LINCS.

If the disease turns genes up, find the compound that turns them down.

Mutations, put back in genomic context

Genome Explorer Recurrent TCGA-BRCA driver mutations on hg19, drawn with igv.js. Click a variant and the browser zooms to it.

A cohort-wide pattern and a single locus are the same view at two zoom levels.

biobouncer

The other thing every one of those apps calls first.

48

identifier systems

4

validation modes

pattern · cache · remote · existenceCRAN · PyPI · npm · Zenodo DOI

biohttp

Every external service fails in its own way. This is the part nobody should have to write twice.

7

concerns handled once

2

kinds of nothing

retry · throttle · cache · timeout · and three moreCRAN

The service said no

A clean empty answer. The gene is real, there is simply nothing recorded on it.

The service fell over

A timeout, a 500, a rate limit. Same silence on screen, completely different problem.

Knowing which one you got is the whole reason an app can lean on eighteen services at once.

bioclients

biohttp knows how to talk to a service. This one knows what each service is saying back.

29

biological services

2

layers, fetch and parse

Ensembl · UniProt · gnomAD · ClinVar · and 25 moreCRAN

Fetching and parsing are separate, so a stored response is enough to test a parser. No live API, no flaky test.

plotomics

The rendering core under Plotomics Live, and under anything else that wants these figures.

17

components

1

core, three languages

WebGL · Canvas · binary buffers, not JSONCRAN · PyPI · npm

Fifteen ship in all three languages. The two genome browsers are JavaScript and Python only.

biocohort

The least glamorous one, and the one my own research leans on hardest. Before an API or a figure matters, the study has to stay organised.

Subjects and samples

Who, and what came off them.

Assays

RNA, ATAC and whole genome, described in one manifest.

Corrections

A swapped label, fixed once, stays fixed.

Sample sheets

Generated from the manifest, not typed by hand.

one manifest per study · many assays · many organismsCRAN

These are cheap problems early and very expensive ones later.

Next steps

Write it up

Blog posts on the parts that generalise. No dates promised.

Keep them running

All ten are deployed on Connect Cloud now. Keeping them current is the part that does not end.

The data got big and scattered. Their tools didn’t follow. Shiny was a good place to fix that.

Every project, and where it lives

All fifteen are open source. The ten applications are also running, not only readable.

All ten run in the Shiny Showcase gallery, collected with their source in posit-dev/shiny-showcase-bioinformatics.

Plotomics Live is built with shinyreact. Barret Schloerke's posit::conf 2026 talk Beyond Bootstrap: Building Custom Shiny UI with React is the place to start.
Bonus: shinyreact-showcase, a gallery of shinyreact examples I built.

Thank you, Shiny team

I came in with biology problems. Almost none of my questions turned out to be biology questions. Thank you to Barret Schloerke and Carson Sievert, and to all of you.

QR code linking to www.samuelbharti.com
www.samuelbharti.com github.com/samuelbharti