The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

rcxl

CRAN status R-CMD-check

Read tabular data from xlsx files with a native C parser.

rcxl decodes each worksheet directly into R vectors as it is scanned, with no intermediate document model. This keeps memory use close to the size of the data itself. The miniz and libdeflate decompressors are bundled, so the package has no R dependencies beyond base R.

Installation

From CRAN:

install.packages("rcxl")

The development version from GitHub:

# install.packages("remotes")
remotes::install_github("vlshields/rcxl")

Usage

read_xlsx() returns a plain data.frame, taking the first row as the header:

library(rcxl)

path <- system.file("extdata", "flights.xlsx", package = "rcxl")
df <- read_xlsx(path)
dim(df)
#> [1] 100  10

Column types are guessed from the cells in the read area; dates come back as POSIXct in UTC.

xlsx_sheets() lists sheets, and sheet takes a name or an index:

multi <- system.file("extdata", "multisheet.xlsx", package = "rcxl")
xlsx_sheets(multi)
#> [1] "Alpha" "Beta"  "Gamma"

read_xlsx(multi, sheet = "Beta")
read_xlsx(multi, sheet = 2)

read_xlsx_all() reads every sheet from a single workbook parse and returns a named list of data.frames:

sheets <- read_xlsx_all(multi)
names(sheets)
#> [1] "Alpha" "Beta"  "Gamma"

Reads can be windowed with an A1 range (column-only and row-only forms work too), or with skip / n_max:

read_xlsx(path, range = "B3:D10")
read_xlsx(path, range = "B:C")
read_xlsx(path, skip = 5, n_max = 10)

col_names accepts TRUE, FALSE (auto names, header row becomes data), or a character vector sized to the read area. col_types takes "text", "numeric", "date", "logical", "guess", "skip", or "list", recycled if length 1:

read_xlsx(path, col_names = FALSE)
read_xlsx(path, col_types = "text")                    # everything as character
read_xlsx(path, col_types = c("skip", rep("guess", 9)))  # drop the first column

Other arguments: na blanks matching cells, trim_ws trims surrounding whitespace, and name_repair is one of "unique" (default), "minimal", or "check_unique". See ?read_xlsx for details.

Set the RCXL_THREADS environment variable to override the worker count. Sheets under 4MB are read serially regardless.

License

MIT. Bundled miniz and libdeflate retain their own copyright and license terms; see inst/COPYRIGHTS.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.