The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
zuhtml turns real-world HTML into ordinary R values: character vectors, lists and data frames. It parses the way a browser does, so malformed markup is repaired rather than rejected. You give it a string, raw bytes, a file, a URL or a connection.
A small catalogue page, as it might have been saved from a site. Some
markup is sloppy on purpose: unclosed <li>s, unquoted
attributes, a stray end tag, and a product card without a price.
page <- '
<!DOCTYPE html>
<title>Tea shop</title>
<nav><ul><li><a href="/">Home</a><li><a href="sale/">Sale</a></ul></nav>
<p>Free shipping over 30 EUR</span>
<div class=product>
<h2 class=name>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a>
</div>
<div class=product>
<h2 class=name>Genmaicha</h2>
<a href="genmaicha.html">details</a>
</div>
<table>
<thead><tr><th>Size<th>Grams</thead>
<tr><td>Small<td>0100
<tr><td>Large<td>0250
</table>'
doc <- html_parse(page, base_url = "https://example.org/shop/")
doc
#> <zuhtml_document>
#> root: <html> with <head>, <body>
#> nodes: 59
#> input: 472 bytes (UTF-8)
#> problems: 1html_read() does the same for a file.
base_url is where the page came from; relative links are
resolved against it.
html_elements() finds every element that matches a CSS
selector. html_element() finds the first match below
each input node, and keeps a missing node where there is none.
That is what keeps extracted columns aligned when some records lack a
field:
cards <- html_elements(doc, ".product")
cards
#> <zuhtml_nodeset[2]>
#> [1] <div class="product">
#> [2] <div class="product">
products <- data.frame(
name = html_text_clean(html_element(cards, ".name")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
products
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.htmlGenmaicha has no price, so it gets NA rather than
shifting the column.
html_text_clean() gives text as a reader wants it;
html_text() gives it exactly as parsed.
html_attr() reads attributes, and
html_serialize() writes nodes back as HTML.
Links, lists and tables have their own extractors:
html_links(doc, absolute = TRUE)
#> text href url
#> 1 Home / https://example.org/
#> 2 Sale sale/ https://example.org/shop/sale/
#> 3 details sencha.html https://example.org/shop/sencha.html
#> 4 details genmaicha.html https://example.org/shop/genmaicha.html
lapply(html_elements(doc, "nav ul"), html_list)
#> [[1]]
#> [1] "Home" "Sale"
html_tables(doc)
#> [[1]]
#> Size Grams
#> 1 Small 0100
#> 2 Large 0250Table columns are character: "0100" keeps its leading
zero. Convert types yourself when you know them, for example with
type.convert().
Real pages nearly always have markup errors, which the parser
repairs. html_problems() lists them:
vignette("selectors"): the supported CSS subset.vignette("tables-and-lists"): how tables and lists are
read.vignette("limits-and-encoding"): resource limits,
encodings, and what zuhtml does not do.These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.