The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

zuhtml

R-CMD-check coverage

zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.

html_read() reads a file, a URL or any R connection. zuhtml has no HTTP client of its own: a URL goes through base R’s url(), and a fetcher that needs headers or authentication hands it the body. zuhtml does not run JavaScript or sanitize HTML.

Installation

install.packages("zuhtml")

The development version, from GitHub:

# install.packages("pak")
pak::pak("pedrobtz/zuhtml")

Example

library(zuhtml)

doc <- html_parse('
  <div class=product><h2>Sencha</h2><span class=price>3.50</span>
    <a href="sencha.html">details</a></div>
  <div class=product><h2>Genmaicha</h2>
    <a href="genmaicha.html">details</a></div>',
  base_url = "https://example.org/shop/"
)

cards <- html_elements(doc, ".product")
data.frame(
  name  = html_text_clean(html_element(cards, "h2")),
  price = html_text_clean(html_element(cards, ".price")),
  url   = html_url(html_element(cards, "a"))
)
#>        name price                                     url
#> 1    Sencha  3.50    https://example.org/shop/sencha.html
#> 2 Genmaicha  <NA> https://example.org/shop/genmaicha.html

html_element() returns one result per card, with a missing value where a card has no price, so the columns stay aligned.

html_read() reads a file, a connection or a URL. A URL becomes the document’s base URL, so relative links resolve against the page:

doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"

rows <- html_elements(doc, "table tr")
views <- data.frame(
  topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
  url   = html_url(html_element(rows, "a"))
)
head(views, 3)
#>                  topic                                                        url
#> 1    Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science      https://cran.r-project.org/web/views/Agriculture.html
#> 3    Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.html

The getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.

Licence

zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and LICENSE.note explains how the two apply.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.