The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.
html_problems() lists what was repaired."0012" stays "0012".<meta> tags, JSON-LD,
microdata) and forms, and convert any node to Markdown with
html_markdown().encoding, or the page’s
<meta charset>.html_read() reads a file, a URL or any R connection.
zuhtml has no HTTP client of its own: a URL goes through base R’s
url(), and a fetcher that needs headers or authentication
hands it the body. zuhtml does not run JavaScript or sanitize HTML.
install.packages("zuhtml")The development version, from GitHub:
# install.packages("pak")
pak::pak("pedrobtz/zuhtml")library(zuhtml)
doc <- html_parse('
<div class=product><h2>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a></div>
<div class=product><h2>Genmaicha</h2>
<a href="genmaicha.html">details</a></div>',
base_url = "https://example.org/shop/"
)
cards <- html_elements(doc, ".product")
data.frame(
name = html_text_clean(html_element(cards, "h2")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.htmlhtml_element() returns one result per card, with a
missing value where a card has no price, so the columns stay
aligned.
html_read() reads a file, a connection or a URL. A URL
becomes the document’s base URL, so relative links resolve against the
page:
doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"
rows <- html_elements(doc, "table tr")
views <- data.frame(
topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
url = html_url(html_element(rows, "a"))
)
head(views, 3)
#> topic url
#> 1 Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science https://cran.r-project.org/web/views/Agriculture.html
#> 3 Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.htmlThe getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.
zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and
LICENSE.note
explains how the two apply.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.