bioclients: Clients for Biological Database Web Services
Source:R/bioclients-package.R
bioclients-package.RdLook up genes, variants and proteins from R, without writing a client for every biological web service. Each service gets one client that makes the request and returns a table. Parsing is a separate function that needs no network, so it can run on a saved response and be tested offline. Transport, retries, caching and error handling are left to the 'biohttp' package. Dependencies for single services are optional, so you do not install what you will not use. The services covered include 'Ensembl', described in Dyer et al. (2025) doi:10.1093/nar/gkae1071 , 'UniProt', in The UniProt Consortium (2025) doi:10.1093/nar/gkae1010 , 'gnomAD', in Chen et al. (2024) doi:10.1038/s41586-023-06045-0 , 'Open Targets', in Buniello et al. (2025) doi:10.1093/nar/gkae1128 , and the 'AlphaFold' Protein Structure Database, in Varadi et al. (2024) doi:10.1093/nar/gkad1011 . Each client's help page cites the service it calls.
Details
Every service module ships two halves, and the split is the point.
The client builds a request, calls into biohttp, and returns its
envelope. It knows URLs, parameters, and rate limits, and it touches the
network. Start at mygene_gene(), gnomad_constraint(),
gnomad_frequency(), or clinvar_classification().
The parser takes an already-parsed response body and returns a canonical
structure. It is pure, it never touches the network, and it is what the
offline tests exercise. See mygene_parse_hits(),
gnomad_parse_constraint(), and clinvar_parse_record().
A caller that already has a response body can use a parser on its own.
What comes back
Every client returns a biohttp envelope rather than raising, so a caller
branches on res$status and writes no tryCatch() of its own. Reach for
biohttp::body_or_null() when the reason for a failure is genuinely not
actionable.
Reaching a source that has nothing for a query is no_data, not an error.
It is an answer.
Batching
mygene_genes() and gnomad_constraints() ask about many genes in one or a
few requests rather than one request per gene. The round trip is where
essentially all the time in these calls goes, so this is the difference that
matters. Both return rows in input order, one row per input, so a caller can
zip results onto its inputs by position.
Author
Maintainer: Samuel Bharti samuelbharti.io@gmail.com (ORCID) [copyright holder]
Authors:
Samuel Bharti samuelbharti.io@gmail.com (ORCID) [copyright holder]
Other contributors:
Barret Schloerke barret@posit.co (ORCID) [thesis advisor]
Carson Sievert carson@posit.co (ORCID) [thesis advisor]
Posit Software, PBC (ROR) [copyright holder, funder]