The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

nhsbsa

R-CMD-check Codecov test coverage Lifecycle: experimental

nhsbsa is a low-level R client for the NHS Business Services Authority (NHSBSA) Open Data Portal, a CKAN data catalogue that publishes open datasets about NHS activity in England — prescribing, dental, pharmaceutical and contractor data among them. The package provides thin wrappers around the portal’s API actions and returns plain data — tibbles for tabular results and lists for metadata — leaving the interpretation of any particular dataset to the caller.

The package deliberately contains no knowledge of specific datasets, and wraps the useful read subset of the API’s actions. Function names and arguments mirror the CKAN API, so if you know the API you already know the package. See the portal’s own API page and the CKAN Action API reference for background. If you need an action the package does not yet wrap, please open an issue.

Status

nhsbsa is experimental and a work in progress. It was developed with Claude Code, modelled on the design of the author’s trud and gtexr API-client packages. Some functionality has yet to be exercised interactively against the live API, so please treat results with care. Bug reports, comments and suggestions are very welcome via the issue tracker.

Installation

Install the released version of nhsbsa from CRAN with:

install.packages("nhsbsa")

Or install the development version from GitHub with:

# install.packages("pak")
pak::pak("rmgpanw/nhsbsa")

Usage

library(nhsbsa)
library(dplyr)
#> 
#> Attaching package: 'dplyr'
#> The following objects are masked from 'package:stats':
#> 
#>     filter, lag
#> The following objects are masked from 'package:base':
#> 
#>     intersect, setdiff, setequal, union

Find datasets — list every id, or search:

datasets <- nhsbsa_package_list()
length(datasets)
#> [1] 2217

hits <- nhsbsa_package_search(q = "prescribing", rows = 5)
hits$count
#> [1] 661

Browsing the portal website and clicking a tag such as #Prescribing filters the dataset list (the page URL becomes /dataset/?tags=Prescribing). A filter query finds the datasets with that tag:

nhsbsa_package_search(fq = 'tags:"Prescribing"')$count
#> [1] 18

The API returns more datasets than the website shows for the tag, because the website hides the Freedom of Information disclosure log by default. Exclude that organisation to match the website’s count:

nhsbsa_package_search(
  fq = 'tags:"Prescribing" -organization:freedom-of-information-disclosure-log'
)$count
#> [1] 5

List a dataset’s resources (files), including each file’s download URL:

resources <- nhsbsa_list_resources("bnf-code-information-current-year")
resources |>
  select(name, format, url) |>
  slice_head(n = 6)
#> # A tibble: 6 × 3
#>   name                                     format url                           
#>   <chr>                                    <chr>  <chr>                         
#> 1 BNF_CODE_CURRENT_202503_VERSION_88       CSV    https://opendata.nhsbsa.net/d…
#> 2 BNF_CODE_CURRENT_202504_VERSION_88       CSV    https://opendata.nhsbsa.net/d…
#> 3 BNF_CODE_CURRENT_202505_VERSION_88       CSV    https://opendata.nhsbsa.net/d…
#> 4 BNF_CODE_CURRENT_202506_VERSION_88       CSV    https://opendata.nhsbsa.net/d…
#> 5 BNF_CODE_CURRENT_202507_VERSION_88       CSV    https://opendata.nhsbsa.net/d…
#> 6 BNF_CODE_CURRENT_202508_VERSION_88_FINAL CSV    https://opendata.nhsbsa.net/d…

Download one of them to disk. You choose the destination directory (it must already exist), and the file is saved there under its own name:

path <- nhsbsa_download_resource(
  "bnf-code-information-current-year",
  resource_id = resources$id[[1]],
  directory = tempdir()
)

Read rows from a tabular (datastore) resource without downloading the whole file. The datastore identifies a resource by its name (e.g. "EPD_202401"), and field names are case-sensitive. Use fields, sort and limit/offset to read:

nhsbsa_datastore_search(
  resource_id = "EPD_202401",
  fields = c("PCO_CODE", "BNF_CHEMICAL_SUBSTANCE", "ITEMS"),
  sort = "ITEMS desc",
  limit = 5
)
#> Warning: ! Retrieved 5 of 18080573 matching rows; 18080568 not returned.
#> ℹ Fetch the next page with `offset = 5` (reusing your other arguments),
#>   increasing `offset` until all rows are retrieved.
#> ℹ Raising `limit` returns more rows per request, up to the server-side maximum.
#> # A tibble: 5 × 3
#>   PCO_CODE BNF_CHEMICAL_SUBSTANCE ITEMS
#>   <chr>    <chr>                  <int>
#> 1 11J00    1404000H0               3584
#> 2 06H00    0212000B0               3571
#> 3 02Y00    0212000B0               3469
#> 4 12F00    1404000H0               3160
#> 5 11M00    1404000H0               3038

To filter by value or aggregate, use SQL. On this portal, datastore_search does not apply the CKAN filters/q parameters, so SQL is the way to filter:

nhsbsa_datastore_search_sql(
  resource_id = "EPD_202401",
  sql = "SELECT PCO_CODE, SUM(ITEMS) AS items
         FROM `EPD_202401`
         WHERE PCO_CODE = 'W2U3Z'
         GROUP BY PCO_CODE
         ORDER BY items DESC
         LIMIT 5"
)
#> # A tibble: 1 × 2
#>   PCO_CODE   items
#>   <chr>      <int>
#> 1 W2U3Z    3129964
nhsbsa_datastore_search_sql(
  resource_id = "EPD_202401",
  sql = "SELECT PCO_CODE, SUM(ITEMS) AS items
         FROM `EPD_202401`
         GROUP BY PCO_CODE
         ORDER BY items DESC
         LIMIT 5"
)
#> # A tibble: 5 × 2
#>   PCO_CODE   items
#>   <chr>      <int>
#> 1 91Q00    3151968
#> 2 W2U3Z    3129964
#> 3 A3A8R    3124609
#> 4 D9Y0V    2697474
#> 5 15N00    2413359

See vignette("nhsbsa") for an overview of the portal, how the package maps onto the website, and the different ways to query data.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.