The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

selectr

License (3-Clause BSD) GitHub Actions CRAN version codecov Downloads per month

selectr is a package which makes working with HTML and XML documents easier. It does this by performing translation of CSS selectors into XPath expressions so that you can query XML and xml2 documents easily.

library(selectr)
xpath <- css_to_xpath("#selectr")
xpath
#> [1] "descendant-or-self::*[@id = 'selectr']"

Installation

Install the release version from CRAN

install.packages("selectr")

Install the development version from GitHub

# install.packages("remotes")
remotes::install_github("sjp/selectr")

Overview

The key functions in selectr are:

Documents read with htmlParse() (XML) or read_html() (xml2) are auto-detected and queried with the HTML translator, so :checked, :disabled, :link and case-insensitive names all work without passing translator = "html" yourself. Queries also chain: querySelectorAll() accepts a node set as well as a document, so querySelectorAll(querySelectorAll(doc, "table"), "tr") runs the second selector from each node the first matched. :scope, :is(), :where(), :has() and :nth-child(An+B of S) are all supported. See ?selectors for the full table of what selectr supports and what each translates to, and ?css_to_xpath for the reasoning behind its more surprising entries.

Examples

Scraping an HTML document

library(selectr)
library(xml2)

html <- paste0(
  "<html><body>",
  "<table id='products'>",
  "<tr><td class='name'>Widget</td><td class='price'>9.99</td></tr>",
  "<tr><td class='name'>Gadget</td><td class='price'>19.99</td></tr>",
  "</table>",
  "<input type='checkbox' checked>",
  "</body></html>")
doc <- read_html(html)

# :has() picks out the table, chaining then walks its rows
rows <- querySelectorAll(doc, "table:has(.price) tr")
querySelectorAll(rows, "td.name")
#> {xml_nodeset (2)}
#> [1] <td class="name">Widget</td>
#> [2] <td class="name">Gadget</td>

# The html translator (auto-detected from read_html()) implements :checked
querySelector(doc, "input:checked")
#> {html_node}
#> <input type="checkbox" checked="checked">

Querying a namespaced XML document

library(selectr)
library(xml2)

# A document with both SVG and MathML content
svgdoc <- read_xml(system.file("demos/svg-mathml.svg", package = "selectr"))
querySelectorAllNS(svgdoc, "svg|script, math|mo",
                   c(svg = "http://www.w3.org/2000/svg",
                     math = "http://www.w3.org/1998/Math/MathML"))

Parsing a large namespaced document with xml2 builds its namespace map on every call by default; pass ns = character(0) to querySelector()/querySelectorAll() to skip that lookup when the document is known to be un-namespaced.

Structured, position-annotated errors

Every error css_to_xpath() and the querySelector*() functions raise inherits selectr_error, so a caller can catch the whole family or a specific class such as selectr_parse_error, which also carries the 1-based character position the parser gave up at:

tryCatch(
  css_to_xpath("div >"),
  selectr_parse_error = function(e) cat(conditionMessage(e), "\n")
)
#> Expected selector, got <EOF at 6>
#>   |
#>   | div >
#>   |      ^

See ?css_to_xpath (section “Errors”) for the full condition hierarchy.

Development

The Makefile wraps the usual tasks:

A devcontainer is provided with R and the package’s dependencies preinstalled.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.